HomeVideos

GTC SJ 2026: Healthcare Reimagined - Bridging Digital Intelligence and Physical Autonomy

Now Playing

GTC SJ 2026: Healthcare Reimagined - Bridging Digital Intelligence and Physical Autonomy

Transcript

989 segments

0:06

Good morning everyone.

0:08

Welcome to the last day of GTC. Now that

0:12

I know how to use this clicker, uh let's

0:15

get started.

0:16

Healthcare is fundamentally shaped by

0:18

physics. Right? So I want to anchor us

0:21

as I talk about all these innovations to

0:23

say this is this is not a net new

0:25

chapter. It's building on the continuum

0:27

that we've been on the for last two

0:30

decades at NVIDIA. But the industry

0:32

itself, right? A century ago with the

0:36

invention of X-ray, CT and ultrasound,

0:39

medicine got its superpower, the ability

0:42

to look inside the human body without

0:44

cutting it.

0:46

Then came image reconstruction, turning

0:48

raw signals into images, and accelerated

0:51

computing found its killer application.

0:55

Then we moved into imageguided therapies

0:58

where imaging started assisting in the

1:00

interventions and guiding the clinicians

1:02

and the surgeons. Deep learning had a

1:05

big bang and we moved from an age of

1:10

hardwaredefined rule-based sensor

1:12

processing to AI that can now start

1:15

learning physics. AI moved beyond the

1:18

scanner to ambient rooms where it could

1:20

understand, hear, listen, observe,

1:23

assist, and that's been the age of

1:25

agentic AI. And we're today at an

1:28

inflection point. Right? All of this

1:31

work leads us to the moment of physical

1:33

AI where AI can see, perceive, but act

1:38

in the physical world. And we'll see

1:40

that across embodiment. And it's all

1:42

built on the same foundation. We'll see

1:45

that across surgical robotics, some of

1:47

the hardest tests of physical AI. We'll

1:50

see it in clinical cobot embodiment,

1:52

care companions, supply chain cobots,

1:56

assistive humanoids.

1:58

And this is really needed for

2:00

healthcare. Think about it. We have

2:03

160,000 hospitals, 72,000 procedures, 8

2:08

million medical devices.

2:11

And while all of this data, all of this

2:13

sensor, all the devices can be

2:14

digitized,

2:16

care is actually delivered physically.

2:20

So if we imagine a world where there

2:22

will be all this assistance that can

2:24

act, the permutation and combination of

2:27

environments,

2:28

the O setups, the different devices, the

2:31

kind of procedure types warrant that

2:34

there needs to be investment in

2:36

simulation platforms that can allow for

2:39

these robots to learn in new

2:41

environments,

2:43

allow them to acquire new skills,

2:46

practice them before they can ever be

2:48

deployed in the real world. So I just

2:51

wanted to anchor us on the need in

2:52

healthcare for physical AI.

2:56

And why do we hear so much about

2:59

physical AI here at GTC is because there

3:01

have been three fundamental tectonic

3:04

shifts that are laying the foundation

3:06

for it. We in the healthcare team build

3:08

upon everything else at NVIDIA. we we

3:12

get the privilege of standing on the

3:14

shoulders of giants at Nvidia who are

3:16

building these foundations and build the

3:19

healthcare derivatives of it. So the

3:22

first pillar is world foundation models

3:25

right world models like cosmos can

3:29

understand and generate complex medical

3:32

environments that's what the team has

3:34

done with post-training cosmos for

3:36

healthcare surgical scenes and clinical

3:38

workflows

3:40

that's the first piece so you have world

3:42

models the second pillar is robotic

3:45

policy models like Groot and Groot H

3:49

that were announced at GTC that can now

3:51

enable these robots to have policies

3:54

that are generalizable across skills and

3:57

can translate from perception and

3:59

reasoning into action.

4:02

And the third pillar is the physics

4:04

simulation.

4:06

We have very high fidelity simulation

4:09

capabilities from classical simulation

4:11

that is fundamentally an accelerated

4:13

computing problem to now neural

4:15

simulation and generative physics. And

4:18

these proxy simulation engines are

4:21

coming very close to real world

4:22

performance. This is what's happening

4:25

across industries that's lending itself

4:28

to healthcare to have its physical AI

4:30

moment.

4:32

And that's why for physical AI

4:36

so far we've always spoken about AI and

4:38

the data challenge and the ability

4:40

whoever has the real world data mode.

4:43

Physical AI changes that because think

4:46

about it like in autonomous driving we

4:49

had all these numbers of hundreds and

4:51

thousands and millions of miles you had

4:52

to drive right but there is no way you

4:56

can drive through every terrain every

4:58

corner every edge case. So the way it is

5:01

actually there there has to be a data

5:03

flywheel investment to bring physical AI

5:05

to life. It has to start with real world

5:07

data because you have to ground these

5:09

models in reality in the physics of the

5:11

world we live in and you get these

5:13

pre-trained models but then there's a

5:15

continuous loop of post-training

5:18

and immense generation of synthetic data

5:21

and like I said you could be generating

5:23

this data with classical simulation

5:25

techniques which are well grounded in

5:28

the physics of the world these models

5:30

have to interact in but lately and more

5:33

so like my colleagues earlier shared

5:35

there are neural simul ation techniques

5:36

and reinforcement learning techniques.

5:39

But this is the flywheel that will allow

5:41

a data gathering strategy to have these

5:45

models in a continuous loop to learn. So

5:48

it's not just about access to AI ready

5:50

data. It's not just volume. It's also

5:53

that quality and coverage. And can you

5:56

you there's no surgeon, no hospital, no

5:59

one hospital would have seen all the

6:01

cases. So can you start leaning on these

6:03

simulation engines to generate synthetic

6:06

data to bring that diversity which

6:09

translates to robustness and

6:11

transparency for these AI models

6:14

and the first investment we had to make

6:16

was in this real world data. So a huge

6:20

shout out to the founding members of

6:23

open age who've been hard at it for the

6:25

last 18 months. Axel Kger, Dr. Nasir

6:28

Nawab, Sean Hoover, Madi Aizian that

6:32

started as a as a conversation has

6:35

turned into one of the largest release

6:37

of healthcare robotics data. Over 35

6:40

plus partners, the likes of CMR

6:42

surgical, moon surgical across 11

6:45

embodiment and 750 plus hours of

6:48

healthcare robotics data. This is the

6:50

first investment we had to make. And now

6:52

that all this data existed, we could

6:54

actually take the vision language action

6:57

models like the Grootbased models which

7:00

have essentially right there are two

7:02

component. There's a cosmos reasoning

7:03

engine which looks at the vision and

7:06

language and understand the environment

7:08

and intent to generate semantic tokens

7:12

that are then ingested by the diffusion

7:14

transformer with the robot states to

7:17

create actionbased policies or action

7:20

tokens.

7:21

And these models if you just take root

7:23

out of the box does not generalize well

7:25

for surgical robotics but open H made it

7:28

possible right so you have that

7:30

foundation and that's what we launched

7:32

at GTC this year group H is fully open

7:36

source it's available to get started on

7:38

GitHub to be download the weights can be

7:41

downloaded on hugging face so that's

7:43

great you have real world data you've

7:45

trained the action policy model that's

7:48

getting us much better performance than

7:50

zero short of any of these generalist

7:52

VAS. But it's not going to be enough.

7:55

And it's not going to be enough because

7:57

of the post-training synthetic data

7:59

generation loop that I shared about.

8:01

You're going to have to train this model

8:03

a lot more. And that's why the third

8:06

pillar of our investment has been in

8:08

Invidia Cosmos.

8:10

The world models, the world states in

8:12

which these robots have to interact,

8:15

they have to be as close as possible to

8:18

reality. So the first derivative we are

8:22

releasing for Cosmos uh H is for

8:25

transfer. It allows for controllable

8:28

synthetic data generation. You could try

8:31

uh you could start in Isaac same

8:33

omniverse to bring

8:36

real digital twin environments of the

8:38

human anatomy that's grounded in the

8:41

physics. So it's it's not generative AI

8:43

that can start hallucinating and then

8:45

you can be applying surgical uh transfer

8:47

to it to to create variations of that

8:51

world. So that is huge.

8:54

Surgical predict allows you to now start

8:58

generalizing and predicting different

9:01

action states. So it's many skills that

9:04

can be in one environment. Again, it's a

9:07

derivative of the foundational Cosmos

9:10

predict model from Nvidia and we've

9:12

essentially brought it all the surgical

9:15

data sets and the surgical domain

9:17

understanding so that it can actually

9:19

predict these different skills, practice

9:21

it in simulation and that simulation is

9:23

is very close to to to the real world.

9:27

And finally, the surgical simulator. Um

9:30

this is what uh the talk earlier

9:32

explained as well but this this is a

9:35

this is a foundation model of a kind

9:38

where we are now moving from rulebased

9:41

simulation to learnable simulation where

9:44

you can evaluate these policies and so

9:46

on and for the aesthetics of the slide

9:49

we've always kept it to benchtop but I

9:51

was hoping the audience here will will

9:54

would like to see some gory surgical

9:56

videos so I'll take that chance

9:58

Right from the left, what you're seeing

10:01

are segmentation depth masks bought in

10:03

from Isaac Sim and you apply Cosmos

10:05

transfer. In the right hand side is a

10:07

generative it's a generative AI powered

10:10

laparoscopic surgical video feed.

10:14

On the right hand side you see an image

10:17

and a prompt to say use the robotic

10:19

forcep to puncture a needle into the

10:21

soft tissue. And I have to admit I'm an

10:23

engineer. I have no surgical background,

10:26

but the video on the right does show

10:28

something looking like taking the

10:31

robotic foresip and puncturing into the

10:33

tissue.

10:35

And the action controlled the Cosmos

10:37

World Foundation model allows for you to

10:40

bring these embodiment and build

10:42

simulators

10:44

that can now start evaluating

10:47

different tasks across different

10:49

policies in different embodiment.

10:52

And that actually lays the foundation

10:54

for how we see

10:57

simulators for even some of the hardest

11:00

problems like surgical robotics being

11:01

solved

11:05

and all of these world foundation models

11:07

the data set we bring it and we anchor

11:11

it into our platform called Isaac for

11:14

healthcare. So I'm essentially making

11:16

sure you all know how is all of this

11:18

being released to developers

11:20

and Isaac for healthcare. In Nvidia

11:22

Isaac is the robotic platform. Isaac for

11:26

healthcare are domain specific

11:27

extensions on Isaac. So what do we go

11:31

do? We go find what are the domain

11:33

specific hard challenges. It includes

11:36

sensor simulation libraries, simulation

11:39

and synthetic data generation pipelines.

11:41

Some of them leveraging the foundation

11:43

models I just shared and models and

11:45

policies like Groot and others. We also

11:49

publish in it end to end workflows,

11:52

blueprints, developer recipes that

11:54

allows anybody to get kickarted

11:57

and it's a way of translating all the

12:00

technology foundations that Nvidia has

12:02

done in the three pillars that I spoke

12:04

about world models, action models and

12:07

simulation technologies and translate

12:09

that to use cases for the healthcare

12:11

industry.

12:14

So double clicking on some of these

12:16

domain specific elements, we have uh

12:19

sensor simulation libraries for

12:22

ultrasound, right? That's allowing you

12:25

to understand end toend simulation

12:27

pipelines from tissue acoustic

12:30

properties to probe configuration to be

12:32

able to generate real realistic scans

12:35

using NVIDIA's optics.

12:38

We also at this GTC have released X-ray

12:42

sensor simulation that allows you to

12:45

understand the end-to-end physics of

12:46

that pipeline from CT derived volumes to

12:49

ray propagation to detector response

12:53

and this allows you now to generate

12:56

synthetic sensor data at scale. So it's

13:01

a fantastic release. Isaac for

13:03

healthcare v0.5 sensor simulation

13:06

libraries available today. Uh in terms

13:09

of the synthetic data generation

13:10

pipelines

13:12

we publish these best practices. Some of

13:15

you might be using Isaac and others but

13:17

a lot of them are for industrial AMRs

13:20

and so on. And so we are really trying

13:22

to close the gap for the healthcare

13:24

developers. You can be building digital

13:26

twin

13:28

by bringing your own patients because

13:29

the world a lot of healthcare lives in

13:32

is the human anatomy and the human

13:34

physiology and so you can be leveraging

13:37

some of the models from manai to

13:39

generate a synthetic MRI data. It can be

13:43

grounded in a real world CT scan. You

13:45

generate a 3D mesh out of it. bring it

13:48

into Isaac's sim and from there you can

13:50

apply cosmos h and you have this engine

13:53

to be able to now understand the human

13:55

anatomy and create all sorts of uh world

13:59

states with that understanding

14:02

we have a pipeline to bring your digital

14:05

twins these are the physical

14:07

environments the robots have to operate

14:09

in and it's the same pipeline you you

14:12

create your digital twin you bring in

14:14

your digital twin assets you can use

14:16

neural reconstru instruction to multiply

14:18

that allow for those interactions and

14:21

those environments to be permutate

14:23

permutated and combined in Isaac's sim

14:27

cosmos transfer and you have the same

14:29

pipeline to essentially create the

14:31

foundation for this simulated world for

14:34

the robots to really learn to be robots.

14:37

We have the surgical video generator

14:39

that leverages the Cosmos H predict

14:42

model that we just shared and you can be

14:44

generating uh your synthetic data out of

14:46

that. And we also have a pipeline that

14:49

leverages the robotic generative physics

14:52

simulator. And all these pipelines are

14:56

available to you today to to just use

14:58

them out of the box and see how these

15:00

technologies connect together. But we

15:03

also publish all the post-training

15:05

recipes, the pre-training recipes, the

15:08

data sets we used for all these models.

15:10

So you could be in the room and

15:12

thinking, but this is all on

15:13

laparoscopic procedures and that's not

15:15

my domain. It's fine. You could you

15:17

could be taking all of that best

15:19

practices and adapting these models to

15:21

your use cases.

15:24

And this is something how the the

15:26

simplest view that we could get to for

15:28

the the physical AI robotics workflow,

15:32

right? So you start with that

15:33

pre-trained Groot H model that we spoke

15:36

about that's been trained with the open

15:38

data set we've released. So you start

15:41

with the pre-trained open Groot H model.

15:44

And now you need to multiply the data

15:46

and the environment variables and the

15:48

world variables Groot has seen. And so

15:51

you could do that in with real world

15:54

data. That's fantastic. You could be

15:57

doing that in classical simulation by

16:00

bringing those anatomical models from

16:02

your CT and MR scans into Isaac SIM and

16:06

applying Cosmos to multiply that or you

16:09

could be doing it with other generative

16:11

AI techniques. What you get out of that

16:13

is a post-trained Groot H model that now

16:16

understands has seen more variability.

16:19

But you still it still needs to practice

16:21

new skills. Maybe it is pre-trained on

16:23

suturing and you needed to get to

16:25

subtask autonomy of needle pickup. And

16:28

so you Isaac lab gives gives that

16:30

environment for that continuous

16:32

reinforcement learning imitation

16:34

learning to acquire these new skills.

16:36

And then Cosmos H simulator that

16:39

learnable simulator can actually become

16:41

the test bed for all these policies to

16:44

be evaluated.

16:46

And we shared in the paper that was

16:48

published uh by Sean and the team that

16:52

the policies like Grood Pi Zero and all

16:55

other VAS when evaluated in this Cosmos

16:59

based simulator

17:01

mimics the ranking of real world

17:03

performance.

17:05

That's really accelerating the

17:06

development time. Imagine having to

17:09

validate each policy on the physical

17:12

embodiment in the lab. That's days and

17:14

weeks. And here you have a simulator and

17:17

you can get that answer in a few hours

17:20

once it's at real-time performance even

17:22

faster. And then you have the simulation

17:24

engine and you know exactly why the

17:26

policy is not performing well and you

17:27

can go multiply more data. So this is

17:30

all about accelerating the researcher

17:32

and developer cycle.

17:35

Uh we've had the fantastic opportunity

17:37

to working with some of the pioneers in

17:39

the industry who are taking these

17:41

technologies seeing their value. Um and

17:43

it's likes of CMR surgical

17:46

uh J&J MedTech for their Monarch

17:48

platform has built on Isaac for

17:50

healthcare for their digital twin and

17:52

are investing with Cosmos based based

17:55

models for their data generation and AI

17:57

model training.

17:59

Moon Surgical, Medbot, Exar Labs, Lem

18:01

Surgical. Let me just highlight a few.

18:04

Like I said, CMR Surgical has really

18:07

pioneered that space. They were they

18:10

could see the value of the foundations

18:12

we are laying. They have come out and

18:14

contributed to the open age data set.

18:17

They are being uh leveraging cosmos age

18:20

the cosmos simulator groupage based

18:22

policy to start in their R&D space

18:25

around subtask automation for their

18:27

surgical procedures. They're also

18:30

leveraging the holoscan and IGX

18:31

platform.

18:33

Johnson and Johnson uh MedTech is using

18:36

Cosmos and Isaac SIM for their synthetic

18:39

data generation pipeline. And it's been

18:42

quite amazing in the last year. It's

18:44

just been 12 months. This is how fast

18:46

this field is moving to see world models

18:49

like Cosmos that understood roads and

18:51

signs and and our physical world, how

18:54

well they adapt and can be post-rained

18:57

to our anatomical and physiological

19:01

world. and our partners are seeing some

19:03

great success with that too.

19:06

Moon surgical is using Cosmos and Isaac

19:09

SIM and Isaac for healthcare uh for

19:12

system level autonomy. So they're still

19:14

in the environmental space but they are

19:15

able to import all the kinematics and

19:17

the robot states to really have that get

19:20

that system level autonomy for the robot

19:22

to get prepared before ever touching the

19:24

patient.

19:26

And so everything I shared so far was

19:28

around surgical robotics. We believe

19:31

that is the hardest test of physical AI

19:33

and that's why we wanted to apply our

19:35

technology foundations to it. But it is

19:38

the same foundations that go beyond

19:40

surgical robotics too. And that's why at

19:43

this GTC we've really expanded Isaac for

19:45

healthcare beyond clinical devices and

19:48

surgical robotics into hospital

19:50

automation. Project Rio was that. I hope

19:53

you've all had an opportunity to to see

19:56

the demo. But it's essentially a recipe,

19:59

a blueprint that developers can get

20:02

started today to say how do I assemble

20:05

my digital twin? How do I create these

20:07

assets? Uh we have fantastic partners,

20:10

the likes of like wheel, softs serve,

20:12

who are masters at doing that. But once

20:15

you've assembled that digital twin, you

20:17

start your data flywheel, the data

20:19

collection efforts, whether through

20:22

teleyop uh through capabilities like

20:24

mimic gen, cosmosbased models, but

20:26

you're essentially in that second bucket

20:28

multiplying the data, the variations

20:31

that the robot has seen and then finally

20:33

getting to policy training and the

20:35

testing in the loop until you're

20:37

confident that this this policy can be

20:40

deployed and I can take it to the real

20:41

world. it it's the same workflow. Uh

20:45

we've had the privilege Peritas AI has

20:48

shown this embraced this as we've built

20:51

it along and they're showing it at the

20:53

show floor. uh we have an amazing video

20:56

with them but they've fully embraced

20:57

from Isaac for healthcare cosmos age

21:00

this synthetic data generation to

21:02

accelerate their development of bringing

21:04

their uh Perry that's the name of their

21:07

humanoid the Perry robot to to the O and

21:10

they've been doing their experiments and

21:12

evaluations with Advent Health uh so

21:15

this is this is not sitting in a lab I

21:17

think all these customers we're sharing

21:19

about this is out of that and we're

21:21

really getting we're hoping and we We

21:24

absolutely believe 2026 is the year that

21:26

we're going to see the deployment

21:28

challenges and the age of deployment

21:30

coming uh for these for these robots.

21:34

And so, okay, you get the data set, you

21:37

can generate all this new data, you can

21:40

make the robots learn new skills, but

21:42

it's all in simulation.

21:44

But the rubber really hits the road once

21:46

you're in deployment.

21:48

And one big release we are having is

21:51

with Holliscan 4.0. O it is the ultimate

21:55

the ultimate deployment stack for

21:57

physical AI. I would say it's it's

21:59

physical AI native and there are three

22:02

key features that make it so. It now

22:04

supports EtherCAT that allows you to

22:07

have these control to directly the uh to

22:11

control the motor of the robot directly.

22:13

It allows interoperability with the

22:15

likes of Ross 2NG streamer. So you're

22:18

not it's it's not a completely new

22:20

stack. It's very interoperable with your

22:23

existing development stacks hopefully

22:25

and it allows for GPU resident graphs

22:27

that are fully accelerated on the GPU

22:30

with no interventions with the CPU

22:32

allowing for that speed of light lowest

22:35

latency in that deployment time. So

22:39

hollowcan 4.0 was released at GTC. Folks

22:42

in the room who are roboticists, who are

22:44

thinking of their next design platform,

22:47

please please go check out Hollowoscan.

22:50

And the ultimate combination is

22:51

hollowoscan running on an IGX Thor. This

22:55

is our industrial-grade edge AI platform

22:58

for realtime safety certified compute

23:01

for robotics devices, factory

23:03

automation.

23:05

IGX thaw really matters because it has

23:08

eight times more compute and I just want

23:12

you all to take that in. Everything that

23:15

I spoke earlier about these VLMs and

23:17

VLAS and Cosmos based models, all of

23:21

that AI is going to land on an edge

23:23

computer. And so edge device design

23:26

choices are being made now. And you want

23:28

to have your AI platforms and your

23:31

devices futureproofed for this age of

23:34

AI. It's no more in the age of CNN's

23:37

right. So IGX Thor and Holoscan 4.0 is a

23:41

fantastic platform which is absolutely

23:43

physically I native

23:46

and this essentially makes the same

23:48

point right what we're doing with

23:49

holoscan is you have this speed of light

23:52

latency and then sometimes you hear but

23:54

I don't need it to go that fast. That's

23:56

not the point. The point is how much

23:58

more can you do in that latency budget?

24:00

Now

24:02

you could do a simple CNN and a

24:04

detection model. If you're here at GTC,

24:07

you're seeing where the field is going.

24:08

And the models of the future are not

24:10

detection CNN's doing singular tasks.

24:12

They are the WLMs. They are the WAS. So

24:15

you want to future proof those

24:16

platforms. And Holcan is built for that.

24:21

All right. This is everything I shared

24:23

with you. This GTC has been about our

24:26

healthcare physical AI platform coming

24:28

out. We had we have released the open

24:31

age data set amazing foundation models

24:34

all the simulation best practices

24:36

recipes uh packaged in Isaac for

24:39

healthcare

24:40

but there's one critical piece and that

24:43

is that as you're thinking about in

24:45

robots how they have sensors for

24:47

perception of environment robots that

24:50

have to operate inside the human body

24:52

their perception is medical devices our

24:56

good old CT and MR and ultrasound sound

24:59

sensors. And so I just want to quickly

25:01

touch on some of the new models we're

25:03

releasing there as well. These are our

25:06

medical AI models, but they also come

25:08

together in service of physical AI

25:10

because a large part of domain specific

25:13

physical AI is inside the human body.

25:17

And so for Nvidia's medical imaging open

25:19

models all built on manai hopefully

25:21

you've heard of that. Um we have segment

25:24

family, we have a generate family,

25:26

reason and raw to insights. This GDC we

25:29

are releasing uh some key key updates to

25:32

that

25:34

introducing the Nvidia raw to insights

25:37

for ultra sound.

25:40

This is this is the first time you're

25:43

hearing about how we believe and where

25:46

we believe AI is headed. This is AI

25:49

based physics sensor processing. So

25:52

directly from the sensor to recon you

25:54

have physics grounded AI from the raw RF

25:57

signal. It can re in real time adapt uh

26:00

do patient adaptive imaging.

26:03

It's production ready and it was all

26:05

built uh in collaboration with Seaman's

26:09

research team with Altera and Nvidia's

26:13

IGX [snorts] and the HSB really allows

26:16

for that intervention of the raw RF data

26:18

to be able to train this model. So

26:21

you're the key takeaway is that you're

26:23

able to AI is getting to a point where

26:26

it can start understanding and deriving

26:29

insights from raw physics

26:32

and maybe there is a lot more insight in

26:34

that sensor processing that that that

26:36

gets lost by the time it's in the

26:38

imaging domain. So please take a note

26:41

GDC 2026 you heard the first time about

26:44

raw to insights and we'll be sharing

26:46

much more about that in the in the

26:48

coming year. We're also introducing raw

26:51

to insights for MRI. Uh this came out of

26:54

our challenge in Mikai with the CMRX

26:59

recon challenge. The foundation model

27:02

delivers consistent highquality

27:04

reconstruction across scanners,

27:06

protocols, and sites. It's shown its

27:08

robustness. It's number one on that

27:10

leaderboard. And the point again is that

27:13

for all the medical device developers,

27:16

researchers and as you guys are thinking

27:18

of where's innovation in that recon

27:20

pipeline, we really think this is the

27:22

moment just like in 2017

27:25

or 20056 where CUDA was discovered for

27:29

recon. This is the moment that AI is

27:31

really going to come to the raw physics

27:33

sensor domain.

27:36

Uh we're also introducing uh NV generate

27:39

for CT and MR. It's a state-of-the-art

27:42

3D diffusion model for CT and MR

27:44

scanners. Uh, and this one's amazed us

27:47

uh, a by its interest by all the

27:49

customers that we speak to, but also the

27:51

kind of use cases. So as an example the

27:55

uh when we have to bring anatomy into

27:57

digital twins there's a lot more CT data

27:59

out there and we generate allows you to

28:02

uh take a CT image and actually generate

28:06

a higher quality more resolution MR

28:08

image which is still grounded in the

28:10

anatomy of uh of the use case and all

28:12

these generating models are not meant

28:14

for any sort of clinical decision-m

28:16

that's not where we think the use cases

28:18

are. This is all in service of

28:20

diversifying the data sets as you train

28:22

these uh new models and modalities.

28:26

Uh Nvidia's medical imaging open models

28:30

are integrated uh into Hopper's AI

28:33

foundry. So if you've heard of Hopper,

28:35

they are here. They have an AI foundry

28:37

platform where they host their

28:40

foundation models like Cury and others

28:43

and Nvidia's open models are integrated

28:45

in that. It's a great place to get

28:47

started. I see Robert in the audience.

28:50

So in case you want to learn more about

28:51

Hopper. [snorts] Um and we're also

28:54

seeing integration with the likes of

28:56

Phillips where where they're taking all

28:58

these foundation models like segment

29:00

generate reason

29:02

these open foundations and post-training

29:05

it with their scanners their specific

29:07

scanners uh details and settings to to

29:11

really get the value uh for for their

29:14

use cases. So huge shout out to Phillips

29:16

as well.

29:18

Lastly, I shared a lot uh and then

29:22

there's just two places where you can

29:24

get started to to soak it all in to get

29:26

started. Uh it's the GitHub NVIDIA met

29:30

that's where we are hosting all the open

29:32

models.

29:33

They are all the source code is all

29:35

under Apache 2.0. We've released the

29:37

pre-training scripts, post-training

29:39

scripts. We're really wanting to move

29:41

the ecosystem forward. So we we are

29:44

sharing in real time as we are learning

29:47

that's up there. Hugging face is where

29:49

we are hosting all the weights. Openh is

29:52

also available on hugging face to

29:54

download today. And Isaac for healthcare

29:56

GitHub is where you can see how all

29:59

these tools and technologies and models

30:01

and data sets can actually be applied to

30:04

to post train for your use cases for

30:06

your applications whether in surgical

30:09

robotics um in clinical medical devices

30:12

uh or for hospital automation any kind

30:15

of embodiment.

30:17

With that thank you all for your time

30:20

today and I'll take questions.

30:24

>> [applause]

30:28

>> Thank you Pera for the talk. Um for

30:30

people who have questions, please line

30:32

up at the microphones that are here in

30:33

the aisle and we'll go first person over

30:36

here.

30:37

>> Okay. My question is about autonomous

30:39

surgery. When do you think this is going

30:41

to happen and how much governance hurdle

30:44

you think the world has to go through?

30:47

It will depends on the countries.

30:50

>> Excellent question.

30:52

Surgical autonomy we believe is the

30:55

absolute northstar in surgical robotics

30:59

for for it to come. So that's we believe

31:01

is is long way out. What we are seeing

31:03

is subtask automation, subtask autonomy

31:07

and I think our partners like Medbot if

31:10

you go to SRS and other conferences they

31:13

have very grand visions about it. even

31:15

at their talk at GTC um Medbot is one of

31:18

the pioneers who's taking some really

31:20

innovative leaps in terms of the

31:22

regulatory stuff. It does come down to

31:24

which countries are we talking about and

31:26

it really the the the word autonomy has

31:29

a has a large gamut and so it really

31:32

then comes down to what is the subtask

31:34

automation and where where those

31:36

guardrails lie.

31:40

>> Thanks for the talk. Um, with these

31:44

approaches where you're passing the raw

31:45

data to the models, do you have some

31:48

kind of a hypothesis about what might be

31:50

missed in the non raw approach or any

31:53

kind of early hints of that?

31:58

>> I would say we still in early research

32:00

to even say that as we pass the raw

32:02

data, can it even start understanding

32:04

the the acoustic behavior and the

32:06

behavior outside of the imaging domain?

32:09

the the domain is getting accelerated

32:11

sooner and I think we will be at a point

32:13

where we'll be able to show what's the

32:15

before and after. So like the seammen's

32:16

ultrasound uh data set or and the and

32:19

the model that we're sharing as as it

32:21

gets applied further like what is the

32:24

what was missed in the before approach

32:26

that now that this AI can understand all

32:29

the sensor physics it's able to augment

32:32

or or predict um I think that's still

32:35

TBD but I do think this would be one of

32:37

the the most interesting topics in some

32:40

of the academic conferences I highly

32:42

recommend you connect with Sean Hoover

32:44

uh who's sitting right there and he's

32:46

leading a lot of this work.

32:48

>> Okay, great. Thanks so much.

32:49

>> Sure.

32:50

>> Okay. Also related to this question uh

32:54

the model that you just show uh going

32:58

from the raw data to imaging domain uh

33:01

you have several different kind of uh

33:03

physical sensor right are they are based

33:05

on the same architecture I mean the

33:06

model itself or it's a different model

33:10

to do the different thing uh different

33:12

raw input data

33:15

>> currently what we are releasing in the

33:17

two model variants we have so for

33:18

ultrasound we are releasing a model

33:21

architecture which is um trained with

33:24

the open age RF data that was

33:26

contributed by semens.

33:28

>> We don't believe that the model

33:30

architecture will be changing based on

33:32

the physics sensor modality but as we

33:35

continue to build upon it we'll we'll

33:37

share more um the the raw to insights

33:40

for MRI is a completely different

33:42

architecture and actually solving quite

33:43

a significantly different problem. it's

33:45

actually getting into the more recon

33:48

approach like an AI based recon approach

33:51

um as compared to just a CUDA based

33:54

recon approach.

33:55

>> Okay, good.

33:56

>> Sure.

33:58

>> Hey, good morning. I have two questions

33:59

on the Cosmos age the data generation.

34:02

So number one so uh how long can can it

34:06

generate the data like for seconds for

34:08

minutes or hours? And the second

34:10

question is is the user able to

34:12

configure what type of data it can

34:15

generate. For example, can can you

34:17

generate a data for a successful task or

34:20

for a failed task?

34:22

>> Uh I followed your first question. Can

34:24

you repeat your second question?

34:26

>> Sure. Sorry. Uh the second question is

34:28

can the model generate a task uh uh the

34:30

data for a successful task or for a

34:33

failed task or should is a user able to

34:36

configure that?

34:38

>> Got it. Yeah. Okay. Okay, so for your

34:39

first question, how many hours of Cosmos

34:42

data can it generate like as much

34:43

compute as you can throw at it, right?

34:46

So it's not limited by it's only these

34:48

many seconds of clip or or so on so

34:50

forth. Um so that hopefully answers the

34:54

first part of your question and your

34:56

second part it can. So uh your second

34:58

part of the question is is it able to

35:00

only uh generate cases of successful

35:03

tasks? That's not uh the case. Indeed,

35:06

Cosmos has surprised us uh with the its

35:09

ability to learn physics

35:11

>> even just by the the videos and so on.

35:15

So, it it does need that constant

35:17

prompting and post-training loop, right?

35:19

So, it's not like you generate all the

35:21

data and you train Cosmos and you're

35:23

done. You generate data, you train

35:25

Cosmos, you prompt it for the corner

35:27

cases you're looking for and is it quite

35:29

getting it, not getting it? And you

35:31

might have to try different approaches

35:32

for it. But it absolutely has shown that

35:34

it can learn the physics and actually

35:36

generate cases that it hasn't seen. And

35:38

that's that's really the point of

35:40

Cosmos, right? It is meant for those

35:42

corner cases. And so as rich uh as a

35:47

input you can give to it. We've seen it

35:49

performs well the best. It performs well

35:51

in all these different variants, but it

35:53

really performs really good when it's

35:55

when you have uh the anchored digital

35:59

assets from CT andMR that I shared

36:02

earlier with you. Um that really grounds

36:04

it in the physics of the of the world

36:07

that it's trying to learn about.

36:09

>> Okay. Thank you.

36:10

>> Sure.

36:12

>> Hi, this is Suresh from United Health

36:14

Group. Um thank you for the talk and

36:16

it's exciting to see you know how much

36:19

physical AI is progressing. Uh obviously

36:21

I'm part of the division in in pharmacy.

36:24

Uh so I'm interested to leverage some of

36:26

the world models uh like you know

36:28

applicability and I definitely connect

36:30

with you and I also let know our

36:33

colleagues on the healthcare side like

36:34

know where you know we can leverage

36:36

that. So do you know who I can contact

36:38

or can I connect with you directly?

36:40

>> Yeah I mean you can connect with me. We

36:42

have amazing NVIDIA team here. So we can

36:44

maybe chat afterwards and I'll get you

36:46

connected with the team.

36:47

>> Okay. Sure.

36:48

>> Thank you.

36:49

>> Hi, thank you for the talk. Um I have

36:52

three questions. The first question is

36:55

um how do you like when you go from a

36:58

simulation to deployment? Something is

37:00

going to break. So in your experience,

37:03

what is the first thing that breaks and

37:05

that simulation is still not able to

37:07

capture completely?

37:11

What is the first thing that breaks? I

37:14

think sim to real is not a solved

37:17

problem, right? Uh we are also making

37:20

those endeavors and the first things

37:22

some of the first things that break are

37:23

the the the

37:26

even just the because your digital twins

37:29

are not physics consistent and accurate.

37:33

Um the it it's it's system level issues

37:36

that we end up seeing with the port

37:38

connectivity issues. So it's not a

37:39

seamless switch today. I think that's

37:42

what we're trying to work out with uh I

37:44

don't know if you attended the earlier

37:46

talk where Madi and Sean shared with the

37:48

hollow scan. So if you are building your

37:50

pipelines with hollow scan then you can

37:54

test that deployment in simulation and

37:57

if you take that same stack our goal is

38:01

to minimize that gap of sim to real

38:03

deployment.

38:05

too many it's very hard to say what is

38:06

the one thing that breaks but too many

38:08

things break and the reason they break

38:10

is that it's it's a completely different

38:11

stack right so there still has a lot of

38:14

bespoke work to be done as you translate

38:16

all that simulation learning um but I

38:20

would say that's much more on the on the

38:22

system and the network and the and and

38:25

the system connector side of the thing

38:28

>> thank you my second question is that if

38:32

you have two systems one system is

38:34

trained completely on simulation, tested

38:36

in phases and then deployed. And then

38:39

you have the second system which is more

38:41

um traditional and conventional and you

38:43

test it and test you build it and test

38:45

it in phases with real-time deployment.

38:48

At the end, which system still performs

38:50

better or what's the difference between

38:53

the completely simulated system versus

38:55

the tested in realtime system?

38:59

I don't think the conversation will be

39:02

so much about systems that are fully

39:04

trained in simulation and systems that

39:06

are trained fully with classical

39:08

approaches. So as an example like I mean

39:11

maybe we can anchor the problem a little

39:12

bit. So if you had to say uh

39:16

I believe today surgical robotics for

39:19

their simulators these companies pay

39:20

multi-million dollars for that very

39:22

accurate uh simulators right but we have

39:27

to start now adding the variables it's

39:29

mult it's it's a cost variable it's a

39:32

development variable it's potentially a

39:34

multi-year approach so just to create a

39:38

com using classical simulation and

39:40

computing technologies to create uh uh

39:45

accurate simulator for a procedure

39:48

much more costly many years of

39:51

deployment development uh and then in

39:54

deployment I'd say I considering that is

39:57

the only system that's used in practice

39:59

the fidelity of the system is good right

40:02

now if you come to what we shared for

40:04

instance with the cosmos 8 surgical

40:06

simulator

40:08

the development and the cost times are

40:10

reduced by a magnitude ude and uh we our

40:14

early work is showing that the the

40:16

behavior of the policy ranking is is

40:19

quite close to real world performance of

40:21

those policies. So we don't believe a

40:24

world where all the classical simulators

40:27

will get replaced by this. We believe

40:29

that the people and the pioneers who've

40:31

been investing the likes of surgical

40:33

sciences who are building these

40:34

classical simulators will will converge

40:37

and there will be an hybrid approach and

40:39

there will be some aspects of that tool

40:41

tissue interaction or the [laughter]

40:45

some aspects of that physics that you

40:47

will need classical techniques and there

40:49

will be some aspects of that accelerated

40:51

development where a close proxy from a

40:53

simulator is good enough. Right? But our

40:56

hope is that that significantly reduces

40:58

a the cost and the development time it

41:01

takes to build these highfidelity

41:04

surgical simulators. Uh which allows for

41:07

the technology to to grow faster and

41:09

expand faster into multiple procedures

41:11

and so on.

41:13

>> Got it. So a hybrid approach, right? I

41:15

notice.

Interactive Summary

This presentation explores NVIDIA's advancements in 'Physical AI' for the healthcare sector, emphasizing the transition from traditional rule-based robotics to embodied AI capable of perceiving and acting in physical environments. The speaker outlines three fundamental pillars: world foundation models (Cosmos), robotic policy models (Groot/Groot H), and high-fidelity physics simulation. By leveraging real-world robotics data and synthetic data generation, developers can accelerate the training and testing of surgical robots and hospital automation tools. The session also introduces new medical imaging models, such as raw-to-insights processing for ultrasound and MRI, and highlights tools like Isaac for Healthcare and Holoscan 4.0 as essential infrastructure for future-proofing edge AI deployments in clinical and surgical settings.

Suggested questions

4 ready-made prompts