HomeVideos

GTC SJ 2026: The AI Native Digital Health Stack A Developer's Guide to 2026

Now Playing

GTC SJ 2026: The AI Native Digital Health Stack A Developer's Guide to 2026

Transcript

1002 segments

0:04

All right, fantastic. Good morning

0:06

everybody. So I am Brad Generu. I'm the

0:08

global lead for healthcare alliances

0:10

here invidia. Uh and uh we're holding uh

0:13

digital health developer day. Uh so for

0:15

the next two hours uh we have a a great

0:18

set of speakers lined up, a lot of

0:20

content. We're going to be talking about

0:21

all the uh SDKs and technologies that

0:24

help enable uh all the developments that

0:27

you're building to transform the world

0:29

of healthcare. Uh our first hour is

0:33

we're going to talk a little bit about

0:34

the technologies that Nvidia brings. Um

0:36

I'm going to talk a little bit about uh

0:38

the the healthc care perspective of the

0:41

technologies. Uh and then I've got

0:43

experts uh lined up. We have Arun and

0:46

Audi who will cover um the reasoning

0:49

side, the language, the the speech side

0:52

as well as uh inference at scale. So um

0:55

so you know strap in we're going to go

0:57

through all these technology. It's going

0:58

to be great. Um I'm covering my

1:01

colleague Ragoff Bonnie. If you you saw

1:02

him on the program, he was unable to

1:04

make it this morning. So you get me.

1:08

So let's talk about uh healthcare uh and

1:13

some of the challenges that that we're

1:15

facing today. Uh it's been, you know,

1:17

pretty incredible watching uh the

1:20

different breakthroughs that that's

1:21

happened. But all these things are so

1:23

necessary because when we think about uh

1:25

what's happening in healthcare where we

1:27

have an aging patient population, we

1:29

have populations around the world that

1:31

do not have access to health care. Um we

1:34

have uh burnout happening on the

1:37

physician side, on the nursing side,

1:39

that's unprecedented. Uh if you think

1:42

about uh the the number of hours that

1:44

our clinicians have to work to catch up

1:47

on everything that's happening during

1:49

the day, it's pretty incredible. And

1:50

we're seeing that, you know, uh most

1:53

hospitals are struggling to even make it

1:55

from a financial perspective. We need to

1:57

do more uh with less. Uh and that's just

2:00

a reality of it. And I mean the great

2:03

news is technology has been uh racing

2:06

forward across all industries. uh and

2:09

what we're seeing with the breakthroughs

2:11

on uh large language models, healthc

2:14

care has been a a benefactor uh in a way

2:17

that technologies before had just have

2:19

not uh and uh we're seeing these these

2:22

transformations happen from the core

2:24

from the EMR from the systems all the

2:27

way forward to meeting the patients

2:29

where they're at uh in uh in their

2:31

journeys. If we just take a look at

2:34

where we've been uh and and where we're

2:36

going um back in the early 20 actually I

2:40

mean AI as a concept has been around for

2:42

more than 60 years. We celebrated 60

2:44

years last year for what artificial

2:46

intelligence is, but it was really

2:48

breakthroughs that happened around uh

2:50

the early 2010s, the late 2000 zeros of

2:54

the decade when we uh started to look at

2:57

uh how we could unlock compute to uh

3:01

meet challenges around things like

3:02

computer vision. It started with can we

3:04

identify a dog or a cat in in a picture

3:08

but all the way through to can we

3:09

extrapolate that and say can we identify

3:12

a stroke or a cancer in a picture in a

3:15

medical image. We we had this rise of

3:17

computer vision.

3:19

Then with transformer models, we

3:22

[clears throat] saw the rise of large

3:24

language models uh and transforming the

3:26

way that we uh look at language and this

3:29

is where generative AI got its start

3:31

where we we had the rise of chat GPT. We

3:35

then saw a breakthrough in uh what you

3:38

know it started with tool calling this

3:39

idea of agentic AI. So being able to uh

3:43

call upon our services uh trigger

3:46

workflows, interact with those workflows

3:49

uh transform the way that we do things

3:51

across all industries, but we saw this

3:53

particularly in healthcare. We started

3:55

to look at what kind of workflows can we

3:57

interact with our clinicians and meet

3:59

them where they're at, provide them

4:00

their insights, but also help them in

4:03

the work that they're doing. The next

4:05

wave, and you heard about this all

4:07

throughout GTC this year, is physical

4:09

AI. uh and uh if we look at what does it

4:13

mean for a hospital to become a robot

4:16

onto itself, can we use computer vision

4:18

and large language models and tool

4:20

calling to transform the entire

4:22

experience? Uh and we're seeing this

4:24

happen before our eyes.

4:27

[clears throat]

4:28

You've probably heard this uh AI is a

4:30

five layer cake. Um we look at the

4:33

technologies from you know right from

4:36

what powers our data centers to the

4:38

chips whether that's the uh GPU the CPU

4:41

the DPU uh to the infrastructure that uh

4:45

connects those together with our data

4:48

our storage our networking our

4:50

communication uh to the foundational

4:52

models and to the applications that run

4:55

on top of that uh and that's no

4:56

different than what we see in healthcare

4:59

uh and how we stack our applications

5:01

leveraging these components uh into

5:04

building uh the applications of the

5:06

future. When we think about across the

5:09

entire landscape um in healthcare and

5:12

life sciences and beyond uh there are

5:14

many different areas that we could focus

5:16

in on. Um I cover digital health uh and

5:19

so we've got our health agents uh we've

5:22

got our inference at scale but we also

5:24

have uh pretty phenomenal technologies

5:26

in the genomics space in the drug

5:29

discovery space in the medical imaging

5:32

space [clears throat] there's many

5:33

different technologies unlock it and if

5:35

you have a look at just some of the

5:36

technologies like Monai like parabrics

5:39

like bioneo uh these are the

5:41

technologies that help unlock everything

5:42

that's happening in the healthcare and

5:44

life sciences space

5:46

in digital health. Uh this whole

5:49

movement around agent care allows us to

5:52

unlock patient engagement. We've got uh

5:54

great partners who are meeting the

5:56

patients where they're at to connect

5:58

with them. Uh on Tuesday we had this

6:00

great uh panel uh of speakers that uh

6:04

included Maven Health uh and Sword

6:06

Health uh sorry Maven Clinic uh that

6:09

covered things like women's care,

6:10

fertility care, uh mental health and

6:13

physiootherapy. be able to connect with

6:15

the patients themselves and enable them

6:17

and unlock their journeys so that they

6:20

can get the most out of their health

6:21

care systems. Uh later today we'll uh

6:25

have uh RAD AI uh as well as a bridge uh

6:28

that will help us talk about uh

6:30

improving the clinician experience when

6:32

we think about radiologists uh or uh ER

6:36

doc, surgeons uh those who are

6:38

documenting the uh patient physician

6:41

conversation

6:42

um or documenting what we see in medical

6:45

images. Technologies that help uh

6:47

improve the clinician experience will

6:49

help them reduce burnout. We could also

6:52

uh optimize what's happening in the back

6:54

office in our health records uh

6:56

departments being able to help on the

6:58

coding side on the billing side and we

7:00

can extend this all the way to through

7:01

to uh clinical trials and working with

7:04

our uh pharma partners to be able to

7:07

match our patients uh with dire rare

7:10

conditions with the clinical trials that

7:12

just might save their life.

7:15

Building these systems are complex and

7:17

this is where Nvidia comes in to help

7:19

the entire developer ecosystem with

7:21

technologies that get them there faster

7:23

that get them there from an enterprise

7:25

perspective. Uh being able to build on

7:27

top of this. So this is using uh SDKs

7:30

and APIs that uh help reduce the

7:32

architectural complexity to make it

7:34

scalable, reusable, uh always driving

7:37

the the performance, you know, in many

7:40

cases 5x or 10x what's available out in

7:43

the open-source space. uh but also

7:45

taking in this you know uh being safe

7:48

leveraging guard rails to help unlock

7:49

this. So we we have these libraries that

7:52

meet our developers where they need them

7:54

best so that they could do their life's

7:56

work.

7:58

Um, I won't go through the technologies

8:01

because we've got, like I said, a great

8:02

set of speakers lined up. But some of

8:04

the things that we use is, uh, Neotron,

8:07

uh, Nemo, uh, Dynamo and Triton, uh, for

8:11

driving the, uh, the language, the

8:13

reasoning, uh, and the inference at

8:15

scale. Uh, we also have, uh, voice and

8:18

speech with Neotron speech. Um but we uh

8:23

up on our website you have uh you know

8:25

all the information the documentation to

8:28

help drive that knowledge to be able to

8:30

take these things and connect them into

8:32

the workflows that you're building uh to

8:34

augment them and drive them at scale.

8:38

When we think about how agents work it

8:40

really comes down to three steps which

8:42

is perceiving the world around us. So

8:44

this is using signal processing. This is

8:46

using uh you know uh keyboard inputs for

8:49

for chat experiences. It's using voice

8:53

uh to be able to capture what someone is

8:55

saying. It's using video feeds to

8:57

capture what's happening in the room. Um

8:59

it's using um language to you know catch

9:02

sentiment and so much more that we can

9:04

get from perceiving the world around us.

9:06

uh we then use our agents to reason over

9:08

that to take all those inputs to take

9:10

the knowledge that it has the database

9:12

that it has connectivity to and really

9:15

think about what is it trying to solve

9:16

for come up with a plan and then drive

9:19

action from that whether that's calling

9:21

upon uh other APIs or uh trigger

9:24

workflows or connect back with a

9:26

response to the user whether that's the

9:28

the patient the physician the nurse uh

9:30

or other

9:32

these agents can be strung together so

9:34

you might have you your uh pharmacist,

9:37

your physician, your nurse, uh your

9:40

health records, uh your unit

9:41

coordinators, uh your informatics teams,

9:44

all connected using uh agents that work

9:46

together harmoniously.

9:49

The way that we stack these uh make it

9:51

really straightforward for us to be able

9:53

to um uh leverage uh this kind of cross

9:57

departmental interdisciplinary

9:59

workflows, which is really really

10:00

important. We can hang these together uh

10:03

in uh what are called blueprints. So in

10:06

in this blueprint which is a little bit

10:08

hard to see on the screen but

10:10

essentially this is an endtoend uh

10:12

conversational uh workflow that uh could

10:15

start from our um uh toy Jensen

10:19

physician uh to doing the uh speech to

10:22

text uh to connecting with agents with

10:24

knowledge across different um areas such

10:27

as uh writing doctor's notes or driving

10:30

intake workflows or being able to uh

10:33

navigate creating um appointments for

10:35

our patients connecting into the

10:37

databases um with um uh Nemo Retriever

10:41

uh connecting with uh our guard rails

10:44

applications connecting with our

10:46

knowledge bases uh and things like drug

10:48

formularies and appointment schedules

10:50

driving the responses back, creating

10:52

that uh text to speech and then

10:55

providing that response back out to our

10:57

user in under a second uh is really the

11:01

ability for us to drive these

11:02

conversations. If you want to learn more

11:05

about what we're doing in the digital

11:06

health space, if you want to have uh you

11:08

know one pager access to all the

11:10

different tools, all the different

11:12

success stories, the programs that we

11:14

offer, uh you could join our digital

11:16

health developers program. So

11:17

developer.envidia.com/digitalalthdevelopers.

11:20

Uh great place to get started. Um but uh

11:23

you know with the experts here, uh drop

11:25

us questions. Uh we're always happy to

11:28

have conversations. Uh and uh with that

11:31

uh I will pass it off to Arun who will

11:34

walk us through reasoning LLMs. Arun,

11:37

come on up. Thank you so much everybody.

11:40

Here you go.

11:44

Hello everyone.

11:47

Great to be here. Um I'm a group product

11:50

manager at Nvidia responsible for LLM

11:52

inference. Uh I'll go over into more

11:55

detail on specifically two topics. what

11:58

are the models we have when it comes to

12:00

LLMs and what is the inference stack

12:02

that we have and how can you take

12:03

advantage of it.

12:10

All right. So the first thing to note is

12:12

that this is version three of Neotron.

12:15

Uh we released version one, version two

12:17

and now we are version three. And you

12:18

may have heard Jensen say this. We will

12:21

keep releasing Neotron as long we shall

12:23

as we shall live. What does this mean

12:25

for you? If you are looking to build on

12:28

top of a model that is going to keep

12:31

improving over time both from an

12:33

accuracy as well as a performance point

12:35

of view, Emotron is a great model to

12:37

build on. It is completely open and it

12:39

is we are committed to keep building on

12:41

it and keep improving it. And the way

12:44

it's set up today is for version three,

12:47

we have three different model sizes,

12:50

nano, super, and ultra. Nano is a 30B

12:53

model intended for very targeted tasks.

12:55

If you're looking to get the model

12:57

working for a very specific thing, you

12:59

could use the nano super is intended 12B

13:02

is intended for multi- aent systems and

13:05

the ultra 500B is intended for

13:07

essentially a replacement for frontier

13:10

models. Something that can think, make a

13:12

plan and orchestrate a whole agentic

13:14

system.

13:16

The key innovations in the Neotron 3

13:19

model family are this now uses the

13:22

hybrid architecture which significantly

13:24

improves efficiency as well as

13:26

significantly increases the the the

13:29

context length has been increased by

13:30

about 7x from the previous generation to

13:32

now. So it has a 1 million context

13:34

length which means that for pretty much

13:36

most workflows you could just provide

13:38

the entire context of to the model and

13:41

it should be able to take it and provide

13:42

you the response that you need. In

13:45

addition to the model, Emotron, you can

13:47

think of it as an ecosystem that comes

13:49

with the data and the tooling that you

13:51

would need if you wanted to customize

13:53

and build your own version of Neimotron.

13:56

When it comes to data sets, we have the

13:58

full suite of things. I'll go into it in

14:00

more detail. As well as we have tooling

14:02

called Nemo Gym that allows you to

14:04

create your own reinforcement learning

14:06

environments to tune the model. And

14:09

we've also published the full research

14:11

on what is under the hood for each of

14:13

these models. Here's

14:18

a couple of slides just taking a quick

14:19

look at how does Neotron compare to the

14:22

rest of the open models in the

14:24

ecosystem. The yax the x-axis here is

14:27

the intelligence index is a combination

14:29

of benchmarks from artificial

14:30

intelligence. Uh it's an average of

14:33

various different benchmarks showing how

14:35

intelligent a model is. And the y- axis

14:37

here is the openness index. How open is

14:39

the model? How easy is it going to be

14:41

for you to take the model and fine-tune

14:43

it if you want it? How easy for you to

14:45

trace it back and understand what is the

14:47

data that's been used to train this

14:49

model and you can see that Neotron 3

14:52

super is on the upper right is a great

14:55

thing and uh there in terms of

14:57

intelligence especially it is much

14:59

better than most other models available

15:01

out there.

15:04

Here's another view. Here it is

15:06

intelligence versus efficiency. The

15:09

y-axis here is intelligence and the

15:11

x-axis is efficiency. And you can see

15:13

that the top two models when it comes to

15:16

intelligence are the Neotron 3 Super 12B

15:18

and the QN3 122B 3.5122B.

15:22

But the Neotron uh 3 Super is much more

15:25

efficient. So if you're looking to take

15:26

advantage of the compute that you have

15:28

and generate more tokens, Neotron 3

15:30

Super is going to be a much better

15:31

choice at at this point.

15:37

Again a quick summary of what all you

15:39

get with Neotron 3. It's a hybrid MOE

15:42

architecture about 4x improvement from

15:45

the previous architecture in terms of

15:47

the overall efficiency.

15:49

We also included by default a technique

15:52

called multi-token prediction so that

15:53

during inference time it predicts three

15:56

three times the number of tokens and is

15:58

able to automatically select the right

16:00

token that is going to help improve the

16:01

overall latency and throughput context

16:04

length improved by about 7x and uh the

16:07

using the our Nemo RL gym tooling we are

16:11

able to also improve the intelligence

16:13

with reinforcement learning by about 2x.

16:18

So this is upcoming. The Neotron 3 Ultra

16:22

is expected to come up in about 2 three

16:23

months and we're already seeing that the

16:25

accuracy results are it's the for its

16:29

model class size. It is the best

16:30

available open model out there. So we're

16:32

incredibly excited for it. In about two

16:34

to three months, you should see the

16:35

Neotron 3 Ultra released and you should

16:38

be able to use it for applications that

16:40

require a lot more complex thinking and

16:42

is a replacement for potentially for

16:44

Frontier models.

16:47

a quick slide on everything that we've

16:49

released with data. So, so far we spoke

16:51

about what is available as a model and

16:53

what you can use out of the box but you

16:56

can also create your own version of

16:58

Neotron and we have various partners

17:00

that that have been able to go this

17:01

route uh take the data that we have add

17:04

in their own custom data and create

17:07

versions of Neotron that really work for

17:09

their enterprise applications.

17:11

Just a quick look at uh the various

17:13

different kinds of data that we've

17:14

released includes uh pre-training data,

17:17

post- training data including

17:19

reinforcement learning, supervised

17:21

fine-tuning data as well as data for

17:23

safety and for creating uh or how we

17:26

created Neimatron to work for various

17:28

different demographics and personas

17:29

across geographies.

17:35

So you have the model architecture, the

17:37

model is available and you have the

17:39

tooling to improve the accuracy. There

17:42

is the reinforcement learning tooling as

17:44

well as the Nemo framework if you wanted

17:46

to train the model. There are techniques

17:48

that are available to improve the

17:50

efficiency of the model. This includes

17:52

things like neural architecture search

17:53

that you can use to distill the model.

17:55

You can train a model, you can distill

17:57

it to make it smaller to fit your very

17:58

specific use case.

18:01

We also have a feature called thinking

18:02

budget that allows you at runtime to

18:04

restrict the number of thinking tokens

18:06

that need to be generated. Let's say you

18:08

are an application and you have a very

18:10

strict latency constraint. You need the

18:12

response to come out within say less

18:14

than a second. Uh in those cases you can

18:16

restrict the thinking budget to say

18:18

1,000 tokens 2,000 tokens so that as

18:21

soon as the model hits the token budget

18:23

it is going to produce start producing a

18:25

response so that you are always meeting

18:26

your latency budget.

18:29

>> [snorts]

18:31

>> And also we also provide various

18:33

different quantization techniques and

18:34

tooling for that so that you can take a

18:36

model that is in BF16 and you can create

18:38

FP8 or NVFP4 versions of it that's going

18:40

to improve the overall latency and

18:42

throughput. And lastly recipes for doing

18:45

each of these step by step. How do you

18:46

train the model? How do you fine-tune a

18:48

model? How do you make it more efficient

18:49

with quantization? Everything is

18:50

available as a recipe that you can take

18:52

and customize and build things yourself.

18:56

That was a quick look at what is

18:58

available from a model and customization

19:00

point of view. Next, we'll go into LM

19:02

inference and uh what does Nvidia

19:04

provide to help you serve models. This

19:06

can be this is can be used for Neotron.

19:08

It can also be used for any open model

19:10

out there.

19:13

So I'll start with NIM here Nvidia

19:15

inference microservices and what we do

19:18

with with NIM is simple. So a new model

19:20

gets released a open model either from

19:22

Nvidia or from the community any open

19:24

model. We take the model and we take all

19:27

possible optimizations both from the

19:29

community as well as from Nvidia and we

19:32

build a container that includes all of

19:34

those optimizations. We validate that

19:36

from a functional point of view from a

19:37

performance point of view that this

19:39

works great and we make sure that it is

19:41

secure. We validate that there are no

19:43

vulnerabilities, no critical or high

19:45

vulnerabilities. We package it up all

19:47

together and provide you a container. We

19:49

also provide other Kubernetes

19:52

support including helm shots, kerf

19:54

deployment uh recipes, nim operator etc.

19:57

So that you have everything that you

19:59

need to take this and deploy it wherever

20:01

you want. The great thing is that

20:02

because it's packaged as a docker

20:04

container with kubernetes native

20:06

support, you can really deploy it

20:07

anywhere. It can be on prem, it can be

20:09

on AWS or it can be on any cloud that

20:11

you wanted to deploy on. And you can you

20:13

can also move from one place to another

20:14

if you decide to choose a different

20:16

infrastructure later.

20:19

And here's what the the advantage that

20:21

NIM provides and what are we doing

20:23

internally. So there are a whole set of

20:26

things that you would need to do when

20:27

you have to do inference. So you there

20:30

is a model that's available. You have to

20:32

choose what is the right backend to use,

20:34

what are the right techniques to use for

20:35

optimization. Do the validation. Make

20:38

sure that you are able to track metrics

20:40

properly. Make sure that Kubernetes

20:41

deployment is set up properly and then

20:43

also make sure that you are able to keep

20:45

doing this on a monthly basis. Just to

20:47

look at the complexity here, let's say

20:48

you are validating 10 different models.

20:50

There are 10 models. There are at least

20:51

three to four different inference

20:52

backends. There are a variety of

20:54

inference techniques. There are a

20:55

variety of hardware. So if you just

20:57

multiply them, there is a lot of

20:58

different experimentation that you would

21:00

need to do to make sure on a on a

21:01

month-to-month basis to make sure that

21:03

your inference deployment is set up to

21:05

get the best possible efficiency. And

21:07

that is all the complexity that the

21:09

Nvidia NIM factory is taking care of and

21:12

making sure that it's delivering you a

21:13

NIM that you can just take and deploy.

21:17

Just a quick look at what is coming. So

21:20

um we've so far we've had the NIM 1.0

21:23

stack and at this GTC we announced that

21:26

we are releasing the NIM 2.0 stack. This

21:28

is now a single backend container that

21:30

is going to more transparently expose

21:33

the underlying inference backend

21:35

capabilities. So, we've had many

21:36

partners ask us, hey, this feature is

21:37

available in VLM. Will this be available

21:39

in NIM? Now, it will be because you're

21:41

going to transparently be able to see

21:43

what version of VLM is included in the

21:45

NIM and all of the features that are

21:46

supported there are going to be

21:47

supported in NIM. Same for other

21:49

backends like TRTLM, SGLANG, etc. So,

21:52

the the bottom line here is that the

21:54

features per security patches everything

21:56

you're going to be able to get faster.

21:58

As soon as something is available in

21:59

open source, it's going to get packaged

22:00

up and you're going to get in a secure

22:02

manner as a num container. That was the

22:05

first major announcement. The second

22:06

announcement is we're also now starting

22:08

to support distributed inference.

22:11

Uh previously NIM was a single container

22:13

but now going forward we also have

22:16

what's called as NIMD and I'll go into

22:20

this in the next slide also supports

22:22

various different distributed

22:23

inferencing capabilities that allow you

22:25

to take advantage when you're deploying

22:27

at of uh optimizations possible when

22:30

you're deploying at scale.

22:33

So what we are seeing with our partners

22:36

is that last year people were deploying

22:39

LLMs for various different applications

22:41

and this year and people are really

22:43

scaling it up to hundreds of GPUs,

22:45

hundreds of nodes and when you do that

22:47

there is a lot of opportunity for

22:48

further optimization. These are some

22:50

examples of things that you could do. So

22:52

for example, let's say there is an

22:54

incoming request coming in and it is

22:56

going to a particular GPU. By default,

22:58

when the next request comes in, it is

23:00

going to go to a different GPU. Even if

23:01

there is a lot of commonality between

23:03

request one and request two. But if you

23:05

are able to intelligently route request

23:07

two to the same GPU that previously

23:09

processed request one, you don't need to

23:11

recomputee the KV cache. So that is

23:13

going to significantly save you the

23:16

compute in terms of uh and reduce your

23:19

overall latency and throughput. That

23:20

that's one example. The another example

23:22

is disagregated serving where you can

23:24

split up the LLM workflow into a

23:27

pre-fill part and the generation part so

23:29

that you process the context in one GPU

23:31

and do the decoding in a different GPU.

23:33

This allows you to more flexibly control

23:35

the latency and throughput. There is

23:38

also the KV block management where if

23:40

you are in a situation where there

23:42

you're running a really longunning agent

23:44

and there is a lot of recurring

23:46

computations that are coming in you are

23:48

able to save the computer KV cache into

23:51

other layers of memory so that you're

23:53

able to reuse it instead of keeping it

23:56

computing it over and over again.

23:58

So a variety of techniques we see that

24:00

for each of these techniques it results

24:02

in orders of magnitude improvement with

24:04

disagregated serving we've seen cases

24:05

where you get up to 3 to 4x with KV

24:08

routing about 2 to 3x and KV block

24:10

management another 3 to 4x so it all

24:12

compounds together and we're really

24:13

excited that all of these capabilities

24:15

are now also going to come to NIM in a

24:17

package manner so you can take it and

24:19

and deploy it.

24:21

So this is what the NIMD looks like. You

24:23

take all of these distributed inference

24:24

capabilities and make it enterprise

24:27

ready. So provide enterprise support.

24:29

Make sure that there are no CVS and make

24:31

sure that it is working on the hardware

24:32

and models that you care about.

24:36

This is a quick example that that that

24:38

was released this week for Neimotron 3

24:40

super that I was mentioning earlier.

24:41

Just enabling the KVR routing

24:43

immediately improves the time to force

24:45

token by 2x. So just round robin routing

24:48

you get x TDFT and with uh the KV

24:52

routing it reduces by about 2x.

24:56

That was a quick overview of the LM

24:57

models and inferencing. I'll now hand it

24:59

over to Adi.

25:03

Thank you.

25:05

[applause]

25:08

>> Thank you everyone and thank you Brad.

25:12

Uh so I hope all of you uh watched the

25:14

keynote and in the keynote there were a

25:17

lot of voice based demos. One of them

25:20

was the in cabin when they were driving

25:22

and the car was speaking to them. So

25:25

it's also a start for edge deployment of

25:27

speech. Um we also saw the open models

25:31

Neimotron voice chat. I'll be speaking

25:33

more about that. And also we had the

25:35

machinery and Olaf Disney's Olaf. He was

25:38

speaking for the first time. Uh last

25:40

year uh the bot was beeping. So it it's

25:43

really an exciting year and we really

25:45

see that speech is maturing uh and and

25:48

it's now being unlocked by multiple

25:50

verticals. And today I'll speak more

25:52

about uh healthcare.

25:55

So I wanted to start with a with a the

25:58

diagram of when you're applying voice

26:00

today, what's happening behind the

26:03

scenes. Uh it's usually a system or what

26:05

we call a pipeline. Uh Brad showed that

26:08

in the healthcare blueprint and also

26:10

Arun just presented three types of LLMs.

26:14

A small one, a middle one and a big one.

26:17

But when you're deploying a voice agent

26:19

today in that cascaded pilot as you call

26:21

it, you can actually choose whatever LLM

26:24

you want. You want a small very

26:26

efficient one because you don't have a

26:28

lot of uh money to spend on every step

26:31

and you want it to be quick, use a

26:33

smaller one. You want to have a huge one

26:35

and ultra and do really sophisticated

26:38

reasoning when the voice is speaking,

26:40

you can use a different one. So um the

26:42

cascaded option provides to you a very

26:45

um broad flexibility and customization

26:49

options but the way it works is it's

26:51

it's really complicated. You need to

26:52

handle audio and voice either with web

26:55

RTC or websocket. So you have to um

26:58

shift the audio around the different

27:00

components and then send it back. It's

27:02

also about latency and of course when

27:04

you need to process it in real time it's

27:07

all becoming very hard because you also

27:08

have the network factor

27:11

all the data that needs to be stored and

27:14

in the LLM and to have context and to

27:16

have history and then you also have the

27:19

interruption voice activation detection

27:21

when is the user speaking when is the

27:23

bot is speaking am I just thinking for a

27:27

minute and I don't want the bot to

27:28

interrupt so it's pretty complex systems

27:31

uh here I don't even have a rag and I

27:34

don't have uh guard rails and safeties

27:36

but it can be a pretty complex um uh

27:40

process. I still think that at least for

27:43

the next few years those deployments and

27:46

NVIDIA provides you with an ASR LLM TTS

27:49

but you can choose any model that you

27:51

like. There's really good TTS models out

27:54

there that providing even more uh

27:57

flexibility and naturalness and it it

27:59

really gives you the the option to pick

28:01

and choose. Uh but then when you have

28:04

that all together that's ready for

28:06

production. Uh we also providing there's

28:08

a link on the top for a blueprint. It's

28:10

very simple. It's a dev example and you

28:13

can take that code and deploy it. Um,

28:16

and it also has orchestration of all of

28:18

the voice and it's about one or two

28:21

hours from taking that example and

28:23

deploying and speaking to your own

28:24

agent. But this is cascaded. Uh,

28:27

cascaded also have a lot of problems or

28:30

issues because as you transfer voice or

28:33

context between those steps, you're

28:35

losing some things. You're losing the

28:37

dynamics. You're losing the the voice

28:39

itself. In healthcare, it's it's has an

28:42

extra meaning because if I have issues

28:44

with my voice and I'm going to a doctor

28:45

and I have a sore throat or I'm coughing

28:49

or I'm I'm even I don't know, I'm

28:51

forgetting things and I'm uming and you

28:54

do want to have all of those things

28:56

actually being heard by the voice agent

28:58

and that's where um full duplex is

29:01

actually coming to play.

29:02

>> Let's see. What do you want to talk

29:04

about?

29:05

>> I'd love to hear a weather joke from

29:07

you. Please share. Here's the joke. Why

29:09

did the sun go to therapy? Because it

29:11

had too many clouds.

29:14

[clears throat]

29:14

>> That's a classic. I like it.

29:16

>> I thought you So, how's your day going?

29:20

>> It's going well. Thanks for asking.

29:24

>> I'll explain after.

29:25

>> Um, having a little bit of a sore throat

29:28

right now.

29:30

>> Oh, no. That's no fun. Is it scratchy or

29:32

more of a burning sensation?

29:34

Yeah, it's a little bit of sore throat

29:36

and a slight burning sensation.

29:39

>> I see. That could be from overuse or

29:42

maybe even a mild infection. Have you

29:45

been talking a lot lately?

29:46

>> No, I haven't. I've been really quiet.

29:48

I've been staying in bed all day.

29:51

>> That might be it. Sometimes lying down

29:54

can cause mucus to pull in the throat.

29:56

Try propping up your head with a few

29:58

pillows.

29:59

>> Okay. And how do you know this?

30:02

I'm a doctor, so I'm trained to listen

30:05

and provide medical advice.

30:07

>> Okay, I understand.

30:10

>> All right. So, what you just saw, I hope

30:12

it sounded more natural. Um, what you

30:15

just saw is actually two of our full

30:17

duplex speech-to-pech models. Uh, one of

30:20

them was announced at the uh keynote as

30:22

as an early access and the other one is

30:24

a research uh model that we released a

30:27

couple of months ago. was called

30:28

Personoplex and it's a fine-tuned

30:30

version of uh Moshi from QTI Labs. And

30:34

on the the one with the um circle, the

30:37

one that was like, "Who are you?" and

30:39

had this rusty voice. This is personal.

30:41

Uh it's super natural. And and then we

30:43

just added a few mixed voices. Those are

30:45

not real voices. And the one uh on the

30:49

left on your left, uh the one that says

30:51

I'm a doctor, that's that's the Neotron

30:54

voice chat. And we we prompted them. We

30:57

just said, "Right, it's just role

30:58

playing. It's not really a doctor." Um,

31:01

and we just, one of them was a patient,

31:03

and we said, "You're just a patient."

31:05

And the other one was a doctor. Um,

31:08

those two those models are actually much

31:10

more natural in the way that they

31:12

converse. Um, I hope you're able to

31:14

understand that both of them are models.

31:16

They're not humans. But the nice thing

31:19

about those models is even as they

31:20

mature, um, you will be able to even use

31:23

that for practice. So if you have a

31:25

practitioner or if you want to train

31:27

somebody to get those calls from

31:29

patients then one side can be that

31:32

person who doesn't even know how to

31:34

describe their illness or if if it's

31:36

like person like me whose English is

31:38

their second language and if you'll tell

31:40

me to describe a scratch or a burning

31:43

sensation I also heard that from a

31:45

different culture you'll say I have ants

31:47

or I have spiders on on my hand and then

31:50

well go get a pest uh expert but no it's

31:54

actually a burning sensation or a cold

31:55

sensation. So those are really

31:57

interesting things that you can play

32:00

with. And then the other side might even

32:02

be that cascaded production use case

32:04

that you have. So there's a lot of

32:06

applications where now the voice is more

32:08

natural. Uh I don't know if you heard it

32:10

was like ah it had this relief um and it

32:13

was more emotional. So specifically in

32:15

healthcare uh that can be super helpful

32:17

and those are things you cannot get uh

32:20

in the in the cascaded options but well

32:22

maybe you can it's super hard. Um so

32:26

those models by the way this was created

32:28

with with cursor so I do encourage all

32:30

of you to try vibe coding it's super

32:32

cool and what I asked you to do is just

32:34

take a statical image uh a static image

32:37

and show us what those models will do.

32:39

So we need them to do tool calling. So

32:41

it means when they're asking what's for

32:43

breakfast and that one sec is is the

32:46

model answering instantly on the back

32:48

end it's actually performing a tool

32:50

calling and maybe there's a smarter LLM

32:53

or a backbone that's doing all those

32:55

calculations or thinking and then it

32:57

comes back with the answer. It can be

32:58

rag it can be a different system um and

33:01

then it also handles interruption. So

33:03

the the the model will know when you're

33:06

barging in or when you're just taking

33:08

your time. So those are really cool and

33:11

hard problems to solve and we're trying

33:13

to um squeeze them into a single model

33:16

so it will be easier to deploy. You can

33:18

take it to the edge you can just have a

33:20

much more natural um conversation. So

33:24

we're starting with an EA because

33:25

because it's hard and those models in

33:27

order for them to be fast we can just

33:29

fit a small nano uh LLM backbone. So if

33:32

you expect it to be super smart as your

33:34

cascaded not today. Um but you do have

33:38

the option today then to choose the both

33:40

of them and use the both of them uh as

33:42

you like. So as as for the voice chat

33:45

it's an early access um the the team

33:48

here the SA and the devil will help you

33:50

um get your hands on one um and it's

33:53

still an evaluation because we do want

33:55

to see exactly what's the first use case

33:58

that you would like to apply such a

33:59

model as we're maturing it. Um and they

34:03

do give provide a balance. So this is

34:06

again artificial analysis um recent

34:08

speech-to-pech uh leaderboard that was

34:11

just updated very similar to what Arun

34:13

was showing just up and right we're just

34:14

changing the model names uh but we are

34:17

trying to deliver something that

34:18

provides value and the value here is is

34:21

um a good balance between the

34:23

intelligence that you see on the x-axis

34:26

and then the emotions or the

34:28

conversational

34:29

um naturalenness dynamics of of of the

34:32

call and it's really hard to measure

34:34

that um a good benchmark today is full

34:38

duplex bench which provides really those

34:40

paws and the interruptions and that's on

34:43

the on the um yaxis and just getting

34:46

both of them at the same model is really

34:48

hard. So it it's really a balance. You

34:50

can be very smart uh but then you have a

34:53

large LLM or you can be very good in

34:54

dynamics like personoplex but it will

34:57

not be that intelligence.

35:00

Yeah. So with that last um slide about

35:04

how you can tune those models, both the

35:07

speech-to-pech models and both the

35:08

cascaded to fit healthcare. If there are

35:11

specific medical terms, if there's um a

35:15

doctor that's speaking in a in a

35:16

specific way and you want to capture

35:18

that, that's super hard in healthcare

35:20

and a repeated problem with speech. Um

35:23

so there's three layers. You can take an

35:25

LLM and you can start adding uh three

35:28

layers of fine-tuning. One of them is

35:29

just term boosting that's being

35:32

supported at inference. You don't have

35:33

to train or fine-tune that. The second

35:36

level is a language model or the

35:38

language layer. So you can fine-tune and

35:40

just add text. So whenever you're adding

35:42

text, the probability of a next word

35:46

that is closer to a medical terminology

35:48

will pop up rather than just a regular

35:50

word. So just add text and the model

35:53

will predict more the next word based on

35:56

on the um on the specific uh jargon or

35:59

terminology. And the last one is also

36:01

acoustic. So if you're working in a very

36:04

um um noisy environment or there's fire

36:07

mics or different background noises, you

36:09

can always take that extra step. Um it's

36:12

not super hard. You do need to have

36:14

those skills of acoustic fine-tuning,

36:16

but that's also applicable. uh and then

36:18

once you're deploying it to uh to to

36:20

production or to inference it really

36:22

depends on your use case as I said some

36:24

of them require a bigger backbone of LM

36:27

some of them require something that's

36:28

more lightweight uh so do pick and

36:30

choose now you have the option and it's

36:32

there uh and then of course as it's

36:35

happening there's rag and tool call um

36:37

and and you do need to understand

36:39

exactly how are you reaching out with

36:41

the wider system in order to fetch data

36:43

or to retrieve it or where all the

36:45

questions uh were provided by the user

36:47

or does the bot needs to collect more

36:49

questions before you move it forward in

36:51

the next step. Um

36:54

yeah, so that's it. Uh the last slide is

36:56

actually just from from today just to

36:58

show you that the ASR leaderboard is

37:00

really shifting fast. Uh if you want to

37:02

deep dive into the advancement on ASR uh

37:06

go to hugging face leaderboard, go to uh

37:08

artificial analysis and see the

37:10

advancement of proprietary models.

37:12

[clears throat]

37:12

Bless you and open models. Um and just

37:15

recently the leaderboard of the open

37:17

source was really exciting. It's like

37:18

every model um is changing but you can

37:21

see that we're trying to support the the

37:23

ecosystem and just released more and

37:25

more open ASR models that then are

37:27

working on the both the cascaded and are

37:30

both supporting our speechtoech models.

37:32

So thank you so much back to you.

37:35

[applause]

37:38

Great.

37:40

All right.

37:42

So just a minute or two to wrap up. I

37:44

mean what you heard today is um the

37:48

technologies that that you need to

37:50

transform healthcare so that you can do

37:52

your life's work uh is available now uh

37:55

some in early access some for for

37:57

download. Um what's so important is that

38:01

we are restoring what it means uh to

38:05

deliver health care for our clinicians.

38:07

If you think about what's happened over

38:09

the last 30 years where uh we've asked

38:11

doctors to key more stuff into

38:13

computers, where we've asked nurses to

38:15

make more phone calls and and and send

38:17

faxes uh and not be doing the nursing

38:20

and not be doing uh the delivery of of

38:23

care. Uh and now we're able to give that

38:26

back so that doctors can be doctors,

38:28

nurses can be nurses, clerks could be

38:30

clerks. Uh and we could transform

38:31

healthc care for all. Uh thank you all

38:34

for the phenomenal work that you are

38:36

doing to transform uh this healthcare

38:38

ecosystem. Uh and I will turn it back

38:41

now to Cedric to close out our session.

38:44

Thank you so much everybody.

Interactive Summary

This presentation, titled 'Digital Health Developer Day,' provides an overview of how NVIDIA's AI technologies, including large language models (LLMs) and physical AI, are being leveraged to transform the healthcare sector. Speakers discuss the challenges in healthcare—such as clinician burnout and aging populations—and explain how NVIDIA's SDKs, NIM (NVIDIA Inference Microservices), Neotron 3, and advanced voice-agent models are enabling solutions. The session covers both cascaded and full-duplex speech models, emphasizing flexibility, natural conversational dynamics, and tools for fine-tuning models to medical terminology.

Suggested questions

4 ready-made prompts