GTC SJ 2026: The AI Native Digital Health Stack A Developer's Guide to 2026
1002 segments
All right, fantastic. Good morning
everybody. So I am Brad Generu. I'm the
global lead for healthcare alliances
here invidia. Uh and uh we're holding uh
digital health developer day. Uh so for
the next two hours uh we have a a great
set of speakers lined up, a lot of
content. We're going to be talking about
all the uh SDKs and technologies that
help enable uh all the developments that
you're building to transform the world
of healthcare. Uh our first hour is
we're going to talk a little bit about
the technologies that Nvidia brings. Um
I'm going to talk a little bit about uh
the the healthc care perspective of the
technologies. Uh and then I've got
experts uh lined up. We have Arun and
Audi who will cover um the reasoning
side, the language, the the speech side
as well as uh inference at scale. So um
so you know strap in we're going to go
through all these technology. It's going
to be great. Um I'm covering my
colleague Ragoff Bonnie. If you you saw
him on the program, he was unable to
make it this morning. So you get me.
So let's talk about uh healthcare uh and
some of the challenges that that we're
facing today. Uh it's been, you know,
pretty incredible watching uh the
different breakthroughs that that's
happened. But all these things are so
necessary because when we think about uh
what's happening in healthcare where we
have an aging patient population, we
have populations around the world that
do not have access to health care. Um we
have uh burnout happening on the
physician side, on the nursing side,
that's unprecedented. Uh if you think
about uh the the number of hours that
our clinicians have to work to catch up
on everything that's happening during
the day, it's pretty incredible. And
we're seeing that, you know, uh most
hospitals are struggling to even make it
from a financial perspective. We need to
do more uh with less. Uh and that's just
a reality of it. And I mean the great
news is technology has been uh racing
forward across all industries. uh and
what we're seeing with the breakthroughs
on uh large language models, healthc
care has been a a benefactor uh in a way
that technologies before had just have
not uh and uh we're seeing these these
transformations happen from the core
from the EMR from the systems all the
way forward to meeting the patients
where they're at uh in uh in their
journeys. If we just take a look at
where we've been uh and and where we're
going um back in the early 20 actually I
mean AI as a concept has been around for
more than 60 years. We celebrated 60
years last year for what artificial
intelligence is, but it was really
breakthroughs that happened around uh
the early 2010s, the late 2000 zeros of
the decade when we uh started to look at
uh how we could unlock compute to uh
meet challenges around things like
computer vision. It started with can we
identify a dog or a cat in in a picture
but all the way through to can we
extrapolate that and say can we identify
a stroke or a cancer in a picture in a
medical image. We we had this rise of
computer vision.
Then with transformer models, we
[clears throat] saw the rise of large
language models uh and transforming the
way that we uh look at language and this
is where generative AI got its start
where we we had the rise of chat GPT. We
then saw a breakthrough in uh what you
know it started with tool calling this
idea of agentic AI. So being able to uh
call upon our services uh trigger
workflows, interact with those workflows
uh transform the way that we do things
across all industries, but we saw this
particularly in healthcare. We started
to look at what kind of workflows can we
interact with our clinicians and meet
them where they're at, provide them
their insights, but also help them in
the work that they're doing. The next
wave, and you heard about this all
throughout GTC this year, is physical
AI. uh and uh if we look at what does it
mean for a hospital to become a robot
onto itself, can we use computer vision
and large language models and tool
calling to transform the entire
experience? Uh and we're seeing this
happen before our eyes.
[clears throat]
You've probably heard this uh AI is a
five layer cake. Um we look at the
technologies from you know right from
what powers our data centers to the
chips whether that's the uh GPU the CPU
the DPU uh to the infrastructure that uh
connects those together with our data
our storage our networking our
communication uh to the foundational
models and to the applications that run
on top of that uh and that's no
different than what we see in healthcare
uh and how we stack our applications
leveraging these components uh into
building uh the applications of the
future. When we think about across the
entire landscape um in healthcare and
life sciences and beyond uh there are
many different areas that we could focus
in on. Um I cover digital health uh and
so we've got our health agents uh we've
got our inference at scale but we also
have uh pretty phenomenal technologies
in the genomics space in the drug
discovery space in the medical imaging
space [clears throat] there's many
different technologies unlock it and if
you have a look at just some of the
technologies like Monai like parabrics
like bioneo uh these are the
technologies that help unlock everything
that's happening in the healthcare and
life sciences space
in digital health. Uh this whole
movement around agent care allows us to
unlock patient engagement. We've got uh
great partners who are meeting the
patients where they're at to connect
with them. Uh on Tuesday we had this
great uh panel uh of speakers that uh
included Maven Health uh and Sword
Health uh sorry Maven Clinic uh that
covered things like women's care,
fertility care, uh mental health and
physiootherapy. be able to connect with
the patients themselves and enable them
and unlock their journeys so that they
can get the most out of their health
care systems. Uh later today we'll uh
have uh RAD AI uh as well as a bridge uh
that will help us talk about uh
improving the clinician experience when
we think about radiologists uh or uh ER
doc, surgeons uh those who are
documenting the uh patient physician
conversation
um or documenting what we see in medical
images. Technologies that help uh
improve the clinician experience will
help them reduce burnout. We could also
uh optimize what's happening in the back
office in our health records uh
departments being able to help on the
coding side on the billing side and we
can extend this all the way to through
to uh clinical trials and working with
our uh pharma partners to be able to
match our patients uh with dire rare
conditions with the clinical trials that
just might save their life.
Building these systems are complex and
this is where Nvidia comes in to help
the entire developer ecosystem with
technologies that get them there faster
that get them there from an enterprise
perspective. Uh being able to build on
top of this. So this is using uh SDKs
and APIs that uh help reduce the
architectural complexity to make it
scalable, reusable, uh always driving
the the performance, you know, in many
cases 5x or 10x what's available out in
the open-source space. uh but also
taking in this you know uh being safe
leveraging guard rails to help unlock
this. So we we have these libraries that
meet our developers where they need them
best so that they could do their life's
work.
Um, I won't go through the technologies
because we've got, like I said, a great
set of speakers lined up. But some of
the things that we use is, uh, Neotron,
uh, Nemo, uh, Dynamo and Triton, uh, for
driving the, uh, the language, the
reasoning, uh, and the inference at
scale. Uh, we also have, uh, voice and
speech with Neotron speech. Um but we uh
up on our website you have uh you know
all the information the documentation to
help drive that knowledge to be able to
take these things and connect them into
the workflows that you're building uh to
augment them and drive them at scale.
When we think about how agents work it
really comes down to three steps which
is perceiving the world around us. So
this is using signal processing. This is
using uh you know uh keyboard inputs for
for chat experiences. It's using voice
uh to be able to capture what someone is
saying. It's using video feeds to
capture what's happening in the room. Um
it's using um language to you know catch
sentiment and so much more that we can
get from perceiving the world around us.
uh we then use our agents to reason over
that to take all those inputs to take
the knowledge that it has the database
that it has connectivity to and really
think about what is it trying to solve
for come up with a plan and then drive
action from that whether that's calling
upon uh other APIs or uh trigger
workflows or connect back with a
response to the user whether that's the
the patient the physician the nurse uh
or other
these agents can be strung together so
you might have you your uh pharmacist,
your physician, your nurse, uh your
health records, uh your unit
coordinators, uh your informatics teams,
all connected using uh agents that work
together harmoniously.
The way that we stack these uh make it
really straightforward for us to be able
to um uh leverage uh this kind of cross
departmental interdisciplinary
workflows, which is really really
important. We can hang these together uh
in uh what are called blueprints. So in
in this blueprint which is a little bit
hard to see on the screen but
essentially this is an endtoend uh
conversational uh workflow that uh could
start from our um uh toy Jensen
physician uh to doing the uh speech to
text uh to connecting with agents with
knowledge across different um areas such
as uh writing doctor's notes or driving
intake workflows or being able to uh
navigate creating um appointments for
our patients connecting into the
databases um with um uh Nemo Retriever
uh connecting with uh our guard rails
applications connecting with our
knowledge bases uh and things like drug
formularies and appointment schedules
driving the responses back, creating
that uh text to speech and then
providing that response back out to our
user in under a second uh is really the
ability for us to drive these
conversations. If you want to learn more
about what we're doing in the digital
health space, if you want to have uh you
know one pager access to all the
different tools, all the different
success stories, the programs that we
offer, uh you could join our digital
health developers program. So
developer.envidia.com/digitalalthdevelopers.
Uh great place to get started. Um but uh
you know with the experts here, uh drop
us questions. Uh we're always happy to
have conversations. Uh and uh with that
uh I will pass it off to Arun who will
walk us through reasoning LLMs. Arun,
come on up. Thank you so much everybody.
Here you go.
Hello everyone.
Great to be here. Um I'm a group product
manager at Nvidia responsible for LLM
inference. Uh I'll go over into more
detail on specifically two topics. what
are the models we have when it comes to
LLMs and what is the inference stack
that we have and how can you take
advantage of it.
All right. So the first thing to note is
that this is version three of Neotron.
Uh we released version one, version two
and now we are version three. And you
may have heard Jensen say this. We will
keep releasing Neotron as long we shall
as we shall live. What does this mean
for you? If you are looking to build on
top of a model that is going to keep
improving over time both from an
accuracy as well as a performance point
of view, Emotron is a great model to
build on. It is completely open and it
is we are committed to keep building on
it and keep improving it. And the way
it's set up today is for version three,
we have three different model sizes,
nano, super, and ultra. Nano is a 30B
model intended for very targeted tasks.
If you're looking to get the model
working for a very specific thing, you
could use the nano super is intended 12B
is intended for multi- aent systems and
the ultra 500B is intended for
essentially a replacement for frontier
models. Something that can think, make a
plan and orchestrate a whole agentic
system.
The key innovations in the Neotron 3
model family are this now uses the
hybrid architecture which significantly
improves efficiency as well as
significantly increases the the the
context length has been increased by
about 7x from the previous generation to
now. So it has a 1 million context
length which means that for pretty much
most workflows you could just provide
the entire context of to the model and
it should be able to take it and provide
you the response that you need. In
addition to the model, Emotron, you can
think of it as an ecosystem that comes
with the data and the tooling that you
would need if you wanted to customize
and build your own version of Neimotron.
When it comes to data sets, we have the
full suite of things. I'll go into it in
more detail. As well as we have tooling
called Nemo Gym that allows you to
create your own reinforcement learning
environments to tune the model. And
we've also published the full research
on what is under the hood for each of
these models. Here's
a couple of slides just taking a quick
look at how does Neotron compare to the
rest of the open models in the
ecosystem. The yax the x-axis here is
the intelligence index is a combination
of benchmarks from artificial
intelligence. Uh it's an average of
various different benchmarks showing how
intelligent a model is. And the y- axis
here is the openness index. How open is
the model? How easy is it going to be
for you to take the model and fine-tune
it if you want it? How easy for you to
trace it back and understand what is the
data that's been used to train this
model and you can see that Neotron 3
super is on the upper right is a great
thing and uh there in terms of
intelligence especially it is much
better than most other models available
out there.
Here's another view. Here it is
intelligence versus efficiency. The
y-axis here is intelligence and the
x-axis is efficiency. And you can see
that the top two models when it comes to
intelligence are the Neotron 3 Super 12B
and the QN3 122B 3.5122B.
But the Neotron uh 3 Super is much more
efficient. So if you're looking to take
advantage of the compute that you have
and generate more tokens, Neotron 3
Super is going to be a much better
choice at at this point.
Again a quick summary of what all you
get with Neotron 3. It's a hybrid MOE
architecture about 4x improvement from
the previous architecture in terms of
the overall efficiency.
We also included by default a technique
called multi-token prediction so that
during inference time it predicts three
three times the number of tokens and is
able to automatically select the right
token that is going to help improve the
overall latency and throughput context
length improved by about 7x and uh the
using the our Nemo RL gym tooling we are
able to also improve the intelligence
with reinforcement learning by about 2x.
So this is upcoming. The Neotron 3 Ultra
is expected to come up in about 2 three
months and we're already seeing that the
accuracy results are it's the for its
model class size. It is the best
available open model out there. So we're
incredibly excited for it. In about two
to three months, you should see the
Neotron 3 Ultra released and you should
be able to use it for applications that
require a lot more complex thinking and
is a replacement for potentially for
Frontier models.
a quick slide on everything that we've
released with data. So, so far we spoke
about what is available as a model and
what you can use out of the box but you
can also create your own version of
Neotron and we have various partners
that that have been able to go this
route uh take the data that we have add
in their own custom data and create
versions of Neotron that really work for
their enterprise applications.
Just a quick look at uh the various
different kinds of data that we've
released includes uh pre-training data,
post- training data including
reinforcement learning, supervised
fine-tuning data as well as data for
safety and for creating uh or how we
created Neimatron to work for various
different demographics and personas
across geographies.
So you have the model architecture, the
model is available and you have the
tooling to improve the accuracy. There
is the reinforcement learning tooling as
well as the Nemo framework if you wanted
to train the model. There are techniques
that are available to improve the
efficiency of the model. This includes
things like neural architecture search
that you can use to distill the model.
You can train a model, you can distill
it to make it smaller to fit your very
specific use case.
We also have a feature called thinking
budget that allows you at runtime to
restrict the number of thinking tokens
that need to be generated. Let's say you
are an application and you have a very
strict latency constraint. You need the
response to come out within say less
than a second. Uh in those cases you can
restrict the thinking budget to say
1,000 tokens 2,000 tokens so that as
soon as the model hits the token budget
it is going to produce start producing a
response so that you are always meeting
your latency budget.
>> [snorts]
>> And also we also provide various
different quantization techniques and
tooling for that so that you can take a
model that is in BF16 and you can create
FP8 or NVFP4 versions of it that's going
to improve the overall latency and
throughput. And lastly recipes for doing
each of these step by step. How do you
train the model? How do you fine-tune a
model? How do you make it more efficient
with quantization? Everything is
available as a recipe that you can take
and customize and build things yourself.
That was a quick look at what is
available from a model and customization
point of view. Next, we'll go into LM
inference and uh what does Nvidia
provide to help you serve models. This
can be this is can be used for Neotron.
It can also be used for any open model
out there.
So I'll start with NIM here Nvidia
inference microservices and what we do
with with NIM is simple. So a new model
gets released a open model either from
Nvidia or from the community any open
model. We take the model and we take all
possible optimizations both from the
community as well as from Nvidia and we
build a container that includes all of
those optimizations. We validate that
from a functional point of view from a
performance point of view that this
works great and we make sure that it is
secure. We validate that there are no
vulnerabilities, no critical or high
vulnerabilities. We package it up all
together and provide you a container. We
also provide other Kubernetes
support including helm shots, kerf
deployment uh recipes, nim operator etc.
So that you have everything that you
need to take this and deploy it wherever
you want. The great thing is that
because it's packaged as a docker
container with kubernetes native
support, you can really deploy it
anywhere. It can be on prem, it can be
on AWS or it can be on any cloud that
you wanted to deploy on. And you can you
can also move from one place to another
if you decide to choose a different
infrastructure later.
And here's what the the advantage that
NIM provides and what are we doing
internally. So there are a whole set of
things that you would need to do when
you have to do inference. So you there
is a model that's available. You have to
choose what is the right backend to use,
what are the right techniques to use for
optimization. Do the validation. Make
sure that you are able to track metrics
properly. Make sure that Kubernetes
deployment is set up properly and then
also make sure that you are able to keep
doing this on a monthly basis. Just to
look at the complexity here, let's say
you are validating 10 different models.
There are 10 models. There are at least
three to four different inference
backends. There are a variety of
inference techniques. There are a
variety of hardware. So if you just
multiply them, there is a lot of
different experimentation that you would
need to do to make sure on a on a
month-to-month basis to make sure that
your inference deployment is set up to
get the best possible efficiency. And
that is all the complexity that the
Nvidia NIM factory is taking care of and
making sure that it's delivering you a
NIM that you can just take and deploy.
Just a quick look at what is coming. So
um we've so far we've had the NIM 1.0
stack and at this GTC we announced that
we are releasing the NIM 2.0 stack. This
is now a single backend container that
is going to more transparently expose
the underlying inference backend
capabilities. So, we've had many
partners ask us, hey, this feature is
available in VLM. Will this be available
in NIM? Now, it will be because you're
going to transparently be able to see
what version of VLM is included in the
NIM and all of the features that are
supported there are going to be
supported in NIM. Same for other
backends like TRTLM, SGLANG, etc. So,
the the bottom line here is that the
features per security patches everything
you're going to be able to get faster.
As soon as something is available in
open source, it's going to get packaged
up and you're going to get in a secure
manner as a num container. That was the
first major announcement. The second
announcement is we're also now starting
to support distributed inference.
Uh previously NIM was a single container
but now going forward we also have
what's called as NIMD and I'll go into
this in the next slide also supports
various different distributed
inferencing capabilities that allow you
to take advantage when you're deploying
at of uh optimizations possible when
you're deploying at scale.
So what we are seeing with our partners
is that last year people were deploying
LLMs for various different applications
and this year and people are really
scaling it up to hundreds of GPUs,
hundreds of nodes and when you do that
there is a lot of opportunity for
further optimization. These are some
examples of things that you could do. So
for example, let's say there is an
incoming request coming in and it is
going to a particular GPU. By default,
when the next request comes in, it is
going to go to a different GPU. Even if
there is a lot of commonality between
request one and request two. But if you
are able to intelligently route request
two to the same GPU that previously
processed request one, you don't need to
recomputee the KV cache. So that is
going to significantly save you the
compute in terms of uh and reduce your
overall latency and throughput. That
that's one example. The another example
is disagregated serving where you can
split up the LLM workflow into a
pre-fill part and the generation part so
that you process the context in one GPU
and do the decoding in a different GPU.
This allows you to more flexibly control
the latency and throughput. There is
also the KV block management where if
you are in a situation where there
you're running a really longunning agent
and there is a lot of recurring
computations that are coming in you are
able to save the computer KV cache into
other layers of memory so that you're
able to reuse it instead of keeping it
computing it over and over again.
So a variety of techniques we see that
for each of these techniques it results
in orders of magnitude improvement with
disagregated serving we've seen cases
where you get up to 3 to 4x with KV
routing about 2 to 3x and KV block
management another 3 to 4x so it all
compounds together and we're really
excited that all of these capabilities
are now also going to come to NIM in a
package manner so you can take it and
and deploy it.
So this is what the NIMD looks like. You
take all of these distributed inference
capabilities and make it enterprise
ready. So provide enterprise support.
Make sure that there are no CVS and make
sure that it is working on the hardware
and models that you care about.
This is a quick example that that that
was released this week for Neimotron 3
super that I was mentioning earlier.
Just enabling the KVR routing
immediately improves the time to force
token by 2x. So just round robin routing
you get x TDFT and with uh the KV
routing it reduces by about 2x.
That was a quick overview of the LM
models and inferencing. I'll now hand it
over to Adi.
Thank you.
[applause]
>> Thank you everyone and thank you Brad.
Uh so I hope all of you uh watched the
keynote and in the keynote there were a
lot of voice based demos. One of them
was the in cabin when they were driving
and the car was speaking to them. So
it's also a start for edge deployment of
speech. Um we also saw the open models
Neimotron voice chat. I'll be speaking
more about that. And also we had the
machinery and Olaf Disney's Olaf. He was
speaking for the first time. Uh last
year uh the bot was beeping. So it it's
really an exciting year and we really
see that speech is maturing uh and and
it's now being unlocked by multiple
verticals. And today I'll speak more
about uh healthcare.
So I wanted to start with a with a the
diagram of when you're applying voice
today, what's happening behind the
scenes. Uh it's usually a system or what
we call a pipeline. Uh Brad showed that
in the healthcare blueprint and also
Arun just presented three types of LLMs.
A small one, a middle one and a big one.
But when you're deploying a voice agent
today in that cascaded pilot as you call
it, you can actually choose whatever LLM
you want. You want a small very
efficient one because you don't have a
lot of uh money to spend on every step
and you want it to be quick, use a
smaller one. You want to have a huge one
and ultra and do really sophisticated
reasoning when the voice is speaking,
you can use a different one. So um the
cascaded option provides to you a very
um broad flexibility and customization
options but the way it works is it's
it's really complicated. You need to
handle audio and voice either with web
RTC or websocket. So you have to um
shift the audio around the different
components and then send it back. It's
also about latency and of course when
you need to process it in real time it's
all becoming very hard because you also
have the network factor
all the data that needs to be stored and
in the LLM and to have context and to
have history and then you also have the
interruption voice activation detection
when is the user speaking when is the
bot is speaking am I just thinking for a
minute and I don't want the bot to
interrupt so it's pretty complex systems
uh here I don't even have a rag and I
don't have uh guard rails and safeties
but it can be a pretty complex um uh
process. I still think that at least for
the next few years those deployments and
NVIDIA provides you with an ASR LLM TTS
but you can choose any model that you
like. There's really good TTS models out
there that providing even more uh
flexibility and naturalness and it it
really gives you the the option to pick
and choose. Uh but then when you have
that all together that's ready for
production. Uh we also providing there's
a link on the top for a blueprint. It's
very simple. It's a dev example and you
can take that code and deploy it. Um,
and it also has orchestration of all of
the voice and it's about one or two
hours from taking that example and
deploying and speaking to your own
agent. But this is cascaded. Uh,
cascaded also have a lot of problems or
issues because as you transfer voice or
context between those steps, you're
losing some things. You're losing the
dynamics. You're losing the the voice
itself. In healthcare, it's it's has an
extra meaning because if I have issues
with my voice and I'm going to a doctor
and I have a sore throat or I'm coughing
or I'm I'm even I don't know, I'm
forgetting things and I'm uming and you
do want to have all of those things
actually being heard by the voice agent
and that's where um full duplex is
actually coming to play.
>> Let's see. What do you want to talk
about?
>> I'd love to hear a weather joke from
you. Please share. Here's the joke. Why
did the sun go to therapy? Because it
had too many clouds.
[clears throat]
>> That's a classic. I like it.
>> I thought you So, how's your day going?
>> It's going well. Thanks for asking.
>> I'll explain after.
>> Um, having a little bit of a sore throat
right now.
>> Oh, no. That's no fun. Is it scratchy or
more of a burning sensation?
Yeah, it's a little bit of sore throat
and a slight burning sensation.
>> I see. That could be from overuse or
maybe even a mild infection. Have you
been talking a lot lately?
>> No, I haven't. I've been really quiet.
I've been staying in bed all day.
>> That might be it. Sometimes lying down
can cause mucus to pull in the throat.
Try propping up your head with a few
pillows.
>> Okay. And how do you know this?
I'm a doctor, so I'm trained to listen
and provide medical advice.
>> Okay, I understand.
>> All right. So, what you just saw, I hope
it sounded more natural. Um, what you
just saw is actually two of our full
duplex speech-to-pech models. Uh, one of
them was announced at the uh keynote as
as an early access and the other one is
a research uh model that we released a
couple of months ago. was called
Personoplex and it's a fine-tuned
version of uh Moshi from QTI Labs. And
on the the one with the um circle, the
one that was like, "Who are you?" and
had this rusty voice. This is personal.
Uh it's super natural. And and then we
just added a few mixed voices. Those are
not real voices. And the one uh on the
left on your left, uh the one that says
I'm a doctor, that's that's the Neotron
voice chat. And we we prompted them. We
just said, "Right, it's just role
playing. It's not really a doctor." Um,
and we just, one of them was a patient,
and we said, "You're just a patient."
And the other one was a doctor. Um,
those two those models are actually much
more natural in the way that they
converse. Um, I hope you're able to
understand that both of them are models.
They're not humans. But the nice thing
about those models is even as they
mature, um, you will be able to even use
that for practice. So if you have a
practitioner or if you want to train
somebody to get those calls from
patients then one side can be that
person who doesn't even know how to
describe their illness or if if it's
like person like me whose English is
their second language and if you'll tell
me to describe a scratch or a burning
sensation I also heard that from a
different culture you'll say I have ants
or I have spiders on on my hand and then
well go get a pest uh expert but no it's
actually a burning sensation or a cold
sensation. So those are really
interesting things that you can play
with. And then the other side might even
be that cascaded production use case
that you have. So there's a lot of
applications where now the voice is more
natural. Uh I don't know if you heard it
was like ah it had this relief um and it
was more emotional. So specifically in
healthcare uh that can be super helpful
and those are things you cannot get uh
in the in the cascaded options but well
maybe you can it's super hard. Um so
those models by the way this was created
with with cursor so I do encourage all
of you to try vibe coding it's super
cool and what I asked you to do is just
take a statical image uh a static image
and show us what those models will do.
So we need them to do tool calling. So
it means when they're asking what's for
breakfast and that one sec is is the
model answering instantly on the back
end it's actually performing a tool
calling and maybe there's a smarter LLM
or a backbone that's doing all those
calculations or thinking and then it
comes back with the answer. It can be
rag it can be a different system um and
then it also handles interruption. So
the the the model will know when you're
barging in or when you're just taking
your time. So those are really cool and
hard problems to solve and we're trying
to um squeeze them into a single model
so it will be easier to deploy. You can
take it to the edge you can just have a
much more natural um conversation. So
we're starting with an EA because
because it's hard and those models in
order for them to be fast we can just
fit a small nano uh LLM backbone. So if
you expect it to be super smart as your
cascaded not today. Um but you do have
the option today then to choose the both
of them and use the both of them uh as
you like. So as as for the voice chat
it's an early access um the the team
here the SA and the devil will help you
um get your hands on one um and it's
still an evaluation because we do want
to see exactly what's the first use case
that you would like to apply such a
model as we're maturing it. Um and they
do give provide a balance. So this is
again artificial analysis um recent
speech-to-pech uh leaderboard that was
just updated very similar to what Arun
was showing just up and right we're just
changing the model names uh but we are
trying to deliver something that
provides value and the value here is is
um a good balance between the
intelligence that you see on the x-axis
and then the emotions or the
conversational
um naturalenness dynamics of of of the
call and it's really hard to measure
that um a good benchmark today is full
duplex bench which provides really those
paws and the interruptions and that's on
the on the um yaxis and just getting
both of them at the same model is really
hard. So it it's really a balance. You
can be very smart uh but then you have a
large LLM or you can be very good in
dynamics like personoplex but it will
not be that intelligence.
Yeah. So with that last um slide about
how you can tune those models, both the
speech-to-pech models and both the
cascaded to fit healthcare. If there are
specific medical terms, if there's um a
doctor that's speaking in a in a
specific way and you want to capture
that, that's super hard in healthcare
and a repeated problem with speech. Um
so there's three layers. You can take an
LLM and you can start adding uh three
layers of fine-tuning. One of them is
just term boosting that's being
supported at inference. You don't have
to train or fine-tune that. The second
level is a language model or the
language layer. So you can fine-tune and
just add text. So whenever you're adding
text, the probability of a next word
that is closer to a medical terminology
will pop up rather than just a regular
word. So just add text and the model
will predict more the next word based on
on the um on the specific uh jargon or
terminology. And the last one is also
acoustic. So if you're working in a very
um um noisy environment or there's fire
mics or different background noises, you
can always take that extra step. Um it's
not super hard. You do need to have
those skills of acoustic fine-tuning,
but that's also applicable. uh and then
once you're deploying it to uh to to
production or to inference it really
depends on your use case as I said some
of them require a bigger backbone of LM
some of them require something that's
more lightweight uh so do pick and
choose now you have the option and it's
there uh and then of course as it's
happening there's rag and tool call um
and and you do need to understand
exactly how are you reaching out with
the wider system in order to fetch data
or to retrieve it or where all the
questions uh were provided by the user
or does the bot needs to collect more
questions before you move it forward in
the next step. Um
yeah, so that's it. Uh the last slide is
actually just from from today just to
show you that the ASR leaderboard is
really shifting fast. Uh if you want to
deep dive into the advancement on ASR uh
go to hugging face leaderboard, go to uh
artificial analysis and see the
advancement of proprietary models.
[clears throat]
Bless you and open models. Um and just
recently the leaderboard of the open
source was really exciting. It's like
every model um is changing but you can
see that we're trying to support the the
ecosystem and just released more and
more open ASR models that then are
working on the both the cascaded and are
both supporting our speechtoech models.
So thank you so much back to you.
[applause]
Great.
All right.
So just a minute or two to wrap up. I
mean what you heard today is um the
technologies that that you need to
transform healthcare so that you can do
your life's work uh is available now uh
some in early access some for for
download. Um what's so important is that
we are restoring what it means uh to
deliver health care for our clinicians.
If you think about what's happened over
the last 30 years where uh we've asked
doctors to key more stuff into
computers, where we've asked nurses to
make more phone calls and and and send
faxes uh and not be doing the nursing
and not be doing uh the delivery of of
care. Uh and now we're able to give that
back so that doctors can be doctors,
nurses can be nurses, clerks could be
clerks. Uh and we could transform
healthc care for all. Uh thank you all
for the phenomenal work that you are
doing to transform uh this healthcare
ecosystem. Uh and I will turn it back
now to Cedric to close out our session.
Thank you so much everybody.
Ask follow-up questions or revisit key timestamps.
This presentation, titled 'Digital Health Developer Day,' provides an overview of how NVIDIA's AI technologies, including large language models (LLMs) and physical AI, are being leveraged to transform the healthcare sector. Speakers discuss the challenges in healthcare—such as clinician burnout and aging populations—and explain how NVIDIA's SDKs, NIM (NVIDIA Inference Microservices), Neotron 3, and advanced voice-agent models are enabling solutions. The session covers both cascaded and full-duplex speech models, emphasizing flexibility, natural conversational dynamics, and tools for fine-tuning models to medical terminology.
Videos recently processed by our community