HomeVideos

GTC SJ 2026: Physical AI for Healthcare Robotics - Simulation-First Design & Accelerated Development

Now Playing

GTC SJ 2026: Physical AI for Healthcare Robotics - Simulation-First Design & Accelerated Development

Transcript

1116 segments

0:04

Good morning. We are going to have two

0:06

backto-back sessions for healthcare

0:08

robotics. You have seen many great

0:10

presentations at GTC. A lot of them

0:13

about physical AI. Now we are going to

0:15

take physical AI and tell you what the

0:18

leading research and companies are doing

0:21

bringing it to the most difficult

0:22

challenge and the most difficult test of

0:24

physical AI which is healthcare. So I'll

0:28

kick off by just repeating some of the

0:30

slides that you have seen right the

0:32

opportunity is massive the need is

0:34

massive

0:36

every person in this room either is a

0:38

patient or related to a patient the

0:40

health care system has a crisis of

0:43

demand and supply and robotics in our

0:47

hypothesis can help immensely. So as you

0:50

listen to the speakers across different

0:52

use cases of uh healthcare, you're going

0:55

to see how physical AI, how these

0:58

horizontal technologies are helping

1:00

already bring in a lot of innovative

1:03

solutions to make things faster, more

1:06

efficient and more accurate.

1:08

But for physical AI, everybody knows the

1:12

main challenge is data, right? And this

1:14

slide is by design supposed to be

1:17

boring. Four boxes. In an ideal case

1:21

scenario, you're going to have the data

1:23

that is curated, annotated. You know

1:25

what it what you're looking at. It is

1:27

searchable. You find the distribution of

1:30

the data that you have. You understand

1:32

what you don't have enough of or you

1:34

don't have at all. And then you go and

1:37

bridge that data gap, those longtail

1:40

data, and use synthetic data generation

1:43

and simulation to bridge the gap.

1:45

Ultimately, you go use the data to train

1:47

a model and deploy and everybody's

1:49

happy. As simple as it looks, every

1:52

single one of these boxes needs many

1:56

many PhD students to work manually to

2:01

make it really happen for a proof of

2:03

life. The data curation annotation is

2:06

very time consuming and inconsistent

2:09

between two even residency departments.

2:11

The the outcome is totally different.

2:14

When it comes to the synthetic data

2:15

generation and simulation for many years

2:18

researchers tried to solve the

2:21

challenges of healthcare robotics, when

2:23

it comes to surgical robots, clinical

2:25

robots, there are gaps and ultimately

2:28

when it comes to training the models,

2:30

right? We are dealing with a dynamic

2:32

environment in healthcare. Whether

2:34

you're looking at the anatomy or

2:35

hospital halls, everything changes in

2:38

every second. There is no prescribed

2:41

path. So you need generalizable models.

2:43

Right? So across this presentation

2:47

from the first presenter to the last,

2:49

you're going to see the new innovations

2:51

across each of these boxes and then

2:54

you're going to hopefully connect the

2:56

dots how these technologies are going to

2:57

help get us where we want to go.

2:59

basically three fundamental shifts that

3:02

we see in other industries that we are

3:04

bringing into healthcare robotics. The

3:05

word foundation models that create

3:08

physically accurate data without solving

3:11

the physics, right? That could be our

3:14

kind of a solution to address the

3:17

fidelity of simulation, right?

3:20

Generative physics. The second bucket

3:23

are those foundation models that you

3:26

have seen many of them coming out these

3:29

days from many of our partners and

3:31

Nvidia itself models like Groot Alpamo

3:35

that reasons when you go and drive and

3:39

it tells you why I'm doing what I'm

3:40

doing which is very very important when

3:42

you are thinking about healthcare

3:44

applications right explanability and the

3:47

last part is what is being happening

3:49

today with accelerating simulation and

3:52

increasing its fid fidelity. These three

3:54

pillars and the shifts in technology is

3:57

unlocking physical AI in other domains

4:00

and we think is going to do the same for

4:02

physical AI in healthcare robotics and

4:05

our task is to bring those and make them

4:07

real. So today when you look at it you

4:11

are going to deal with a lot of models

4:14

that are going to be open sourced right

4:16

and we are on the leading uh kind of

4:19

edge for these type of models but your

4:23

development starts by taking those

4:26

models and adding your own data. If you

4:29

have real world data, great. If you

4:31

don't, you have two other path. Generate

4:33

data like models like Cosmos that we

4:36

talked about to generate data and you're

4:38

going to hear from our partners that are

4:40

doing that or simulate data, solve the

4:43

partial differential equations, the

4:45

traditional way, first principal way.

4:47

Now you have a good healthy mix of three

4:49

sources of data, post- train those open

4:51

models, deploy, test. You can test in

4:55

silicone which is going to be much

4:57

faster in simulation or inward

4:59

foundation model and the flywheel keeps

5:02

going on. It is all about the data

5:04

flywheel and you see many of the leading

5:08

uh companies on this page. The same

5:11

pipeline applies for nonclinical robots,

5:14

hospital automation. And with that,

5:18

without further ado, I want to introduce

5:21

Filipo to start from the first couple of

5:24

boxes of that boring slide and then we

5:27

go from there.

5:28

>> Beautiful.

5:31

[applause]

5:32

>> Well, thank you. Thank you Mustafa for

5:33

having me here. Uh, my name is Filipo

5:35

Fukori. I am a surgeon. I am the program

5:37

director for the robotic fellowship at

5:39

Lenoxil and Northwell Health, the

5:40

largest healthcare provider in New

5:42

Northeast New States. So what we do is

5:44

we take board certified surgeons and we

5:46

tweak their skills uh and we get them

5:48

ready to go out there and and practice

5:50

in the real world. I also lead a

5:52

laboratory that does research in um

5:56

predictive uh analytics for surgery. So

5:59

looks into the performance of a surgeon

6:01

as a whole and comes up with ways and

6:03

models in order to improve the

6:04

performance of the surgeon. Um we have

6:07

seen an incredible uh innovation in the

6:09

surgical environment over the past 30

6:11

years. You know for thousands of years

6:13

people used to do surgery uh with uh in

6:16

the open fashion and then 30 years ago

6:18

we started doing surgery with like small

6:20

canulas and laparoscopy and then the

6:22

real revolution came with the robot. Now

6:24

the robot uh has more than 30% of the

6:27

share market of a minimally invasive

6:28

surgical space. Uh and the robot allows

6:31

you to collect all of the data. We're

6:33

really only leveraging a minority of the

6:35

data that is coming out of the robot and

6:37

we're all really working hard in order

6:39

to understand the data that comes out of

6:41

the robot and to make it uh actionable

6:44

so that uh we can bring out to the field

6:46

better surgeons. Uh but there are a lot

6:48

of challenges dealing with uh healthcare

6:51

data. First and foremost, it's really

6:53

hard to get data out of hospitals with

6:55

because of uh medical legal problems.

6:57

then it's very hard to curate the data.

6:59

Just surgeons and all of those clinical

7:01

people that need to give you the

7:02

annotations are just really hard to pin

7:04

down. Uh and thanks to video language

7:07

models these days like we can embed the

7:09

video language models in our pipelines

7:10

and we can actually really improve and

7:13

cut down dramatically the num the time

7:15

that is required to annotate these data

7:17

sets. And then you know whenever

7:18

[clears throat] all data is not just the

7:20

same right like when you have a video of

7:22

a surgical procedure uh the high

7:24

impactful moments are going to be two or

7:26

3% when you're going to have a bleed

7:28

when you're going to have an inadvertent

7:29

injury of a certain structure and also

7:32

we have a huge lack of ontologies and

7:35

taxonomies in surgery. We just never

7:37

asked oursel that problem and so we're

7:39

working in order to fix that problem as

7:41

well. Um and we are also working really

7:44

hard in order to enable uh going from uh

7:48

you know real real data into simulation

7:50

data and together with uh Cosmos uh the

7:53

Cosmos product uh we are going to be

7:56

able to create better simulations that

7:58

uh really take from the real world and

8:00

allow us to uh train uh better

8:03

generation of uh surgeons and then

8:05

potentially enable autonomous sections

8:07

and so on. Because the final and the

8:09

ultimate like goal that we have here is

8:11

to for a model to have full surgical

8:13

scene understanding. So for a model to

8:15

understand exactly what is going on uh

8:17

in the screen and Nvidia is helping us

8:19

out by doing that and we're taking a

8:21

page off the playbook that they use an

8:23

autonomous vehicle. So we're using uh

8:25

the same kind of um pipeline uh and with

8:29

some of their products including Cosmos

8:31

Search and Surgical Robotics. If you're

8:33

a developer, if you're working on

8:34

creating models around surgical

8:36

performance, you can search like, you

8:38

know, a specific, you know, event or you

8:40

can search a specific procedure uh and

8:42

you can use that data in order to

8:44

leverage uh and and you know, create

8:46

better models that are uh specifically

8:48

tailored to that event as well. And as

8:50

was uh telling you, we've been doing a

8:51

lot of work around surgical ontologies.

8:53

I mean, there's no way around this. Like

8:55

in order to create surgical ontologies

8:57

for procedures, you need to get uh a lot

9:00

of surgeons under the same roof.

9:02

Surgeons that are both smart and savvy

9:04

about ontologies and the ultimate use

9:07

case for computer vision model and

9:09

kinematics and autonomous surgery. So

9:11

this is what we did a couple months ago.

9:13

We got a lot of our uh members from

9:15

sages together and we created an

9:17

ontology that looks into the surgical

9:19

gestures and surgical gestures are key

9:22

uh because basically all the surgical

9:24

gesture is what you know really powers

9:26

video language action models and so on.

9:29

Uh and you know a month later we were

9:31

able to apply uh the surgical gesture

9:34

framework uh into the open age uh data

9:37

set and this was a great honor. Um and

9:39

uh obviously you know whenever you're

9:41

annotating these surgical procedures. So

9:43

for example in the top left uh videos um

9:46

you're going to see uh two things.

9:48

You're going to see certain anatomical

9:49

structure being annotated because like

9:51

those are the places in the human body

9:53

where you have large vessels and if you

9:54

get down there you could be creating

9:56

like even some lethal injuries. Uh and

9:59

also like the target anatomy uh for

10:02

repairing a hernia for example. So you

10:04

see the ligament over there that it's

10:05

being annotated and that's kind of the

10:07

landmark for us to perform a high

10:09

quality surgery and put in a mesh so

10:12

that the patient doesn't have a

10:14

recurrence of his own hernia.

10:16

Uh and you know talking about simulation

10:18

I'm a huge fan of simulation. This is

10:20

like a state-of-the-art product. It's a

10:22

sim now too from intuitive surgical and

10:24

this is me at Italian tech week a couple

10:26

of months ago. I enjoy like talking to

10:29

audiences of engineers uh and I was

10:31

simulating uh you know in real time live

10:34

uh and uh this is a great product but

10:36

it's still like a simulation where every

10:39

pixel uh is rendered pixel by pixel and

10:42

we really do think that we can do better

10:44

than that and together with uh Sean

10:46

Hoover uh and Nigel and uh Kai we're

10:49

really trying to bring to life uh the

10:52

Cosmos um realtime simulated

10:54

environment. So this is a different

10:56

approach to simulation. Instead of

10:58

rendering everything pixel by pixel, the

11:00

idea uh is that we basically uh create a

11:03

ground truth given by the first frame uh

11:06

of the surgical operation and then we

11:08

use the kinematic data in order to

11:10

simulate the interaction with the tissue

11:12

and we've created a whole suite of

11:14

models including gshian models to

11:16

simulate the interactions between the

11:17

tissue and the tool. And uh we're

11:20

working hard to bring this uh to pro uh

11:22

to production and uh you know so that

11:24

all of our surgical trainees uh can have

11:27

a better environment where to train and

11:28

learn because like believe it or not

11:30

like simulation is is hopefully

11:33

underutilized in surgical training. So

11:35

the people that are actually operating

11:37

on uh on people they sit down at a

11:39

console with me and they do actually

11:41

learn how to operate on real humans. And

11:43

this is not fair. It's not ethical. And

11:45

that's why we're trying to really work

11:47

hard to make this better. And uh last

11:49

but not least, I just wanted to show you

11:52

uh this uh video of an autonomous robot

11:55

that was pre-trained with open age and

11:56

was trained regroup. Uh I'm not going to

11:58

talk about the technicalities that it's

11:59

going to be for axle, but in the left

12:01

upper left corner you see a surgical uh

12:04

intern. So a first year surgical reset

12:06

suturing. Suturis is actually one of the

12:08

most complex things that you can do with

12:09

a robot. And you see how much the

12:11

surgical intern is hesitating. And you

12:13

know this was amazing work uh done by

12:16

Axel and the Nvidia team uh and using

12:18

video language action models which you

12:20

know reinforces the fact that you know

12:21

gesture ontologies and gesture words are

12:23

really really uh important uh the model

12:26

and the the robot was able to throw a

12:29

stitch and and tie a knot. And with that

12:31

I wanted to thank all of my team for the

12:33

incredible work that they're doing in

12:35

advancing uh you know our surgical

12:37

knowledge. Thank you. [applause]

12:47

Uh thank you Mustafa for inviting me. Uh

12:49

really uh thrilled uh to be here today

12:52

and uh tell you about our work on

12:54

autonomous robotic surgery. Um wanted to

12:57

start with quick disclosure. Um uh in

13:00

addition to being an associate professor

13:02

at Johns Hopkins, we recently founded as

13:05

semor surgical where I'm the chief robot

13:07

officer. Um

13:12

Um so uh uh yeah uh here um uh you know

13:17

we are showing this uh rising uh case

13:19

load uh so this upcoming healthcare

13:21

crisis uh where uh we have projected uh

13:25

uh in the next uh uh 10 years a doubling

13:28

of the case load. So we as engineers

13:30

really need to provide you know

13:32

technical uh you know help to cope with

13:35

this rising case load. Um uh so uh

13:39

manual and teleyoperated robotic

13:41

assisted uh uh procedures of course

13:44

heavily depend on the experience and the

13:46

skill of the operating surgeon. Uh our

13:49

vision is to augment critical parts of

13:52

this uh manual operated uh surgery with

13:55

increasing levels of autonomy. Uh this

13:58

could reduce complications uh

14:00

democratize access to expert surgery and

14:03

help with that raising uh rising case

14:05

load.

14:07

Uh traditionally uh we have um you know

14:10

solved this uh problem uh with a

14:13

modelbased approach. Um so we have uh

14:16

you know uh published a paper in science

14:19

robotic 2022.

14:21

Um where uh we uh you know uh uh first

14:26

track tissue in 3D um then planned we

14:29

had to like develop our own custom

14:31

suturing tool uh custom uh camera

14:34

system. Um and while this is

14:36

predictable, we definitely hit a ceiling

14:38

on performance about 60% as uh you know

14:41

stitch-to-stitch success rate. Um and um

14:45

uh and it it it really doesn't scale to

14:47

new procedures. So what is super

14:49

exciting about these learning based

14:51

methods and in particular imitation

14:52

learning based methods that we can build

14:54

one framework and then add more data and

14:57

learn you know new techniques and new

14:59

skills uh you know on the fly and get

15:02

better and better. Um so we started this

15:05

uh we use uh for the hardware the Da

15:07

Vinci research kit. Uh we added a uh

15:10

wrist camera and uh started with um

15:13

tackling uh some fundamental surgical

15:16

task like lifting tissue, needle pickup

15:19

and handover and not tying. Uh high

15:21

level uh really um you know uh

15:24

relatively simple. We get uh expert uh

15:27

robot demonstrations. uh we capture the

15:30

video and the associated kinematics and

15:32

then learn uh and train uh you know uh

15:35

transformer models uh to uh you know

15:39

produce the right action giving the

15:40

right video input. And with just you

15:43

know a few hundred uh expert

15:45

demonstrations we were able to show that

15:47

we can uh you know perform some of these

15:50

you know expert uh you know demonst uh

15:52

expert um you know procedures um you

15:56

know with high fidelity. So on the left

15:58

you see not tying on uh the middle uh

16:01

you know 100% success rate uh needle

16:03

pickup and handover and then lifting

16:06

tissue the simplest task you know 100%

16:08

success rate

16:10

um but uh you know what is super nice uh

16:14

this is not sequentially programmed so

16:16

it's really robust uh to disturbances so

16:18

here we knock the suture thread out of

16:20

the grasp and it automatically repairs

16:23

itself and you know completes the task

16:27

But this is all suture pad and this is

16:29

really small little surgical task right.

16:31

So how do we get it to the next level

16:33

and perform you know procedure level you

16:36

know phases of uh you know uh robotic

16:39

surgery. Uh we've been digging into uh

16:42

the colcystectomy taking out the

16:43

gallbladder. Uh we published this in

16:45

science robotic uh last summer. We're

16:47

focusing on this middle phase of

16:49

clipping and cutting the bile duct and

16:52

cystic artery. Um uh we uh obtained

16:56

about 30 uh gallbladders uh in house and

17:00

uh repeated uh this clipping and cutting

17:03

uh you know 20 uh times on each uh

17:06

sample to get about 16 hours of uh

17:09

gallbladder uh surgery data. Um and um

17:13

then uh we trained this hierarchical

17:16

policy. Uh so now we divided the whole

17:18

face of surgery into different subtasks

17:21

um and first have an uh high level uh

17:24

language policy that looks at the

17:26

history of the video and then finds out

17:28

what is the phase of the surgery that

17:30

we're currently doing and if there are

17:31

any corrections needed and that feeds

17:34

into a language conditioned um you know

17:36

low-level policy that to then execute

17:39

and uh you know find the associated

17:41

kinematics. It's very nice to also

17:43

intervene as a surgeon because the robot

17:45

tells you what it wants to do next. Uh

17:47

so you know really nice way to intervene

17:50

and uh here uh some uh an example video

17:54

then in our study fully autonomous did

17:57

not require any intervention. You see

17:59

some you know adjustments when it

18:01

initially misses on uh the bile duct and

18:04

here the clipier comes in uh places

18:07

three clips then we change into a

18:09

scissor and autonomously uh cut the bile

18:12

duct and then the same on the right with

18:15

the cystic artery. So really nice

18:17

demonstration that this hierarchical

18:19

framework can achieve you know high um

18:22

you know fidelity autonomous soft tissue

18:25

surgery um in this uh xvivo porcen model

18:30

um and here's a compilation of you know

18:33

eight uh consecutive uh you know study

18:36

uh videos uh in our Xvivo model.

18:40

Um we since uh deployed uh this

18:43

framework the SRT framework on different

18:46

robotic systems for different

18:48

application. Uh top left you see

18:50

endoscopic uh guidance um and then uh

18:53

for endoluminal guidance uh under X-ray

18:57

uh you see uh the suture bot that we

18:59

published in Nurops for end to- end

19:01

suturing suturing. uh on the bottom uh

19:04

uh left is uh for trachea uh tumor

19:08

resection and bottom right is an example

19:10

for partial nefrectomy so cutting out a

19:13

portion of a kidney cancer. So all these

19:16

examples are in-house data set uh and

19:18

policies that we trained with in-house

19:20

you know bespoke you know policies. So

19:23

really to take this to the next level,

19:25

super excited to have collaborated with

19:27

Nvidia and so many partners to create a

19:29

much larger data set so we can you know

19:32

create much more robust uh you know uh

19:34

policies and uh uh uh uh really tackle

19:37

you know cross embodiment task and and

19:39

more uh you know uh more applications.

19:43

Um so we collected uh over 150,000 uh

19:47

you know uh trajectories with

19:48

kinematics. That's the largest um

19:51

surgical robotic data set uh published

19:54

uh put it out on on on on Monday and

19:56

already have a thousand downloads of the

19:59

you know one terabyte of data. So I

20:01

think it really creates a lot of uh

20:03

excitement in our community. Uh what can

20:05

we do with it? Uh we can train uh Groot

20:08

to uh perform suturing. So this is

20:10

trained on uh the uh full uh uh open age

20:14

data set and then fine-tune on

20:15

suturebot. And here's an example of end

20:18

toend suturing. Uh so super exciting

20:20

early results. Uh hopefully we'll follow

20:22

this up with a uh paper in the next

20:25

months uh with a more detailed study. Uh

20:28

but already uh in this initial you know

20:30

few weeks of testing uh we showed that

20:32

we can outperform the state-of-the-art

20:34

in suturing with this crude age model.

20:38

Um and what is also super exciting that

20:40

it does require a lot less uh suture

20:43

data to get proficient. So on only 33%

20:47

of the data set we already achieve uh

20:49

really high performance.

20:52

Another collaboration is with uh Nvidia

20:55

on Cosmos um where uh we uh trained a

20:59

world model trained on open age and you

21:01

can see always the video pairs. Uh the

21:04

left is the original video, the right

21:06

one is the AI generated one. uh and uh

21:09

so super exciting that we can you know

21:11

make this so realistic and that helps us

21:14

in policy testing.

21:16

Um so uh where do I see the the future

21:19

of autonomous surgery? It's going to

21:21

take a little while before procedure

21:23

level autonomy uh you know becomes uh

21:25

you know a product uh but I think these

21:27

autonomous functions in laparoscopic

21:30

surgery will come over the next few

21:32

years. Uh super exciting that Moon

21:34

Surgical already has uh the you know uh

21:37

camera control uh you know other uh

21:39

lowerhanging fruits that are a little

21:41

less risky would be suction holding and

21:43

tissue traction and port placement. Um

21:48

uh yeah with that uh let me thank uh my

21:51

uh team and all the great collaborators

21:54

and uh thank you so much for your

21:56

attention.

21:57

[applause]

22:05

All right. Hello everyone. Um, Mustafa

22:07

Jay, thanks for having me here to talk a

22:09

little bit about how we're looking to

22:10

close the data gap to enable physical AI

22:12

on Maestro. And so, um, thanks Axel also

22:15

for the shout out on our our scope pilot

22:17

feature. So, for those of you who aren't

22:19

familiar with Maestro, want to quickly

22:20

introduce it. It's a two armed cartbased

22:23

system that collects to off-the-shelf

22:24

laparoscopic instrumentation. And we

22:26

believe that this is enabling a new

22:28

category of minimally invasive surgery.

22:29

One that allows for more procedures with

22:31

fewer resources. Uh enhances procedure

22:33

consistency for increased patient safety

22:35

and and or staff quality of life. And

22:37

crucially for for this conference and

22:39

audience, we're equipping the O with a

22:41

closed loop feedback mechanism to

22:42

deliver AI based continuous

22:44

improvements. And we're available in the

22:46

US and the EU and we have a broad

22:47

indication for holding and positioning

22:49

uh laparoscopic instruments. So the

22:52

system itself at its at its heart is a

22:54

collaborative surgical robot. So the

22:56

user is able to grab or hold of his or

22:58

her tools and move them anywhere in the

22:59

workspace. Um when they let go of the

23:01

tools, they'll be locked in place. Um

23:03

and so the system sort of int infers the

23:05

intent of of the of the user. In this

23:07

particular part of the clip here, you

23:09

can see a setup joint being

23:10

repositioned. And this is key to the

23:12

sort of later part of my talk where I'm

23:14

talking about how we configure the

23:15

system at the beginning of a surgery is

23:17

is a is a a problem statement that we're

23:19

trying to solve or make more efficient.

23:20

And simulation uh can really help. We

23:23

can also port top as you see here moving

23:25

from one port to another without any

23:26

re-registration. Um, and as as Axel just

23:28

alluded to the first sort of advanced

23:30

feature we've deployed, we call it scope

23:32

pallet where here as the surgeon is

23:34

moving the instrument with manually with

23:36

their right hand, the camera is

23:38

automatically following it. So, we're

23:39

giving the surgeon the ability to

23:40

control three instruments with just two

23:42

hands.

23:43

Um, Maestro adapts to any laparoscopy.

23:46

So, I showing you here just four

23:48

different surgeries. And the uh the

23:50

reason for showing this slide is that

23:52

you'll notice that the arm

23:53

configurations in every uh surgery here

23:55

is very different. And so one of the

23:57

things that we're trying to do to make

23:58

surgery more efficient is allow them to

24:00

set up the system quicker. Um we have

24:02

done over nearly at this point nearly

24:04

3,000 procedures, but you'll appreciate

24:05

different rooms, different setups,

24:07

different beds, different surgeon

24:08

heights. Uh very much uh each case is

24:11

slightly uh is slightly different. And

24:13

you either do thousands and thousands

24:14

more procedures to optimize this or you

24:16

simulate to try and make things faster.

24:18

And that's exactly what we're doing and

24:19

I'll I'll talk about in a moment. Um

24:22

before jumping to that I just want to

24:24

touch on the fact that Maestra was

24:25

designed from the beginning as an edge

24:27

compute powerhouse to be at the bedside.

24:29

So moving around this graphic from left

24:31

to right we pull in the laparoscopic

24:33

video feed so we can see inside the

24:34

patient. We have two 3D depth cameras

24:36

doing RGB and D imaging on both on the

24:39

top of the uh what we call this the

24:40

lighthouse and also down below lower in

24:42

the system so we can see outside the

24:44

patient. And I'll be referring to this

24:46

lighthouse camera a little bit later. So

24:48

just just want to point your draw your

24:49

attention to it. The robot arms

24:51

themselves obviously are are sensors in

24:53

our case high resolution force sensors.

24:55

Um the system is connected uh outbound

24:57

only at this point but it's it's

24:58

connected to Wi-Fi or 5G and all of that

25:01

is interfaced to the IGX or uh with an

25:03

A6000 GPU and an Azure capture card. So

25:06

that's the the product as it is today.

25:08

But obviously this is a session on

25:09

simulation. So I do want to touch on how

25:11

we are using simulation to close the

25:13

data gap. Um, and our first target

25:15

application is optimized arm deployment.

25:16

And and I love the suturebot from from

25:18

Axel, but in a regulated environment,

25:20

there's a crawl, walk, run element to

25:22

this. And the first place that we're

25:24

targeting is something that is a lower

25:25

patient risk. It is set up. I fully

25:28

believe we're going to keep moving

25:29

through that the risk profile to be more

25:30

and more advanced things as as Axel

25:32

touched on. Um, but where we're starting

25:34

is in optimized deployment. And this is

25:35

not currently a feature, but I wanted to

25:37

share as to how we're thinking about the

25:38

problem and what we're doing. So if you

25:40

take the you know the three computer um

25:43

concept uh on the top left we're

25:45

simulating uh in Isaac for healthcare

25:47

with the maestro embedded within uh the

25:49

simulation environment. We're then

25:51

generating synthetic and using real

25:52

world data and training um to gen

25:54

simulate lots and lots of different

25:56

configurations. Again you know we don't

25:57

want to do we don't have to get to the

25:59

point where we've done 50,000 surgeries

26:00

to simulate 50,000 surgeries and

26:02

thankfully we now have the tools to do

26:04

this. And then lastly in the bottom we

26:05

have the deploy side where the IGX at

26:07

the bedside and this video clip just

26:09

shows how an example how the arms might

26:10

deploy. So going through this the first

26:13

up you have the simulation and what's

26:15

cool about this simulation is that it's

26:17

running our full software stack. So

26:18

that's our user interface and you can

26:19

touch on the user interface uh as if you

26:21

were a user and the system let me is

26:23

that not playing?

26:26

No, it is not. Okay, no worries. Um let

26:28

me go back up. What you would see here

26:30

is that uh as you go through the user

26:32

interface, you touch on the uh the guey,

26:34

the arms will deploy, the system will

26:36

move, and you're basically getting a

26:37

full simulation within it. But you'll

26:39

appreciate we're able to move the system

26:41

with relative to the bed, the angle, the

26:43

height of the bed, the color of the

26:44

drape, all of the different

26:45

configurations that we would see in a

26:47

clinical setting. Um, and that it's been

26:49

very powerful. It is a great debug tool

26:51

if you're if you're looking for debug

26:52

tools as well because the software stack

26:54

is integrated. um and we're using it to

26:56

render uh different configurations, but

26:59

there is a need to enrich your data set.

27:00

And so what we're doing here is we're

27:02

taking um real world video from our that

27:04

top camera I referred to and we're

27:06

generating synthetic versions of that

27:08

video using Cosmos transfer. So the

27:10

video on the top is is a real video uh

27:12

as recorded, but the lower one is a

27:14

synthetic video that's been generated uh

27:16

fine-tuned Cosmos transfer. And so

27:18

things that are different, there's no

27:19

blood on the patient's skin, the skin

27:20

has changed, the profile, the drape is

27:22

different, the light cable is a

27:23

different illumination. And so this way

27:25

we're able to generate different

27:26

clinical scenarios without having to

27:28

actually uh uh do the procedures

27:30

themselves.

27:32

We're also going the other way where

27:33

we're taking the Isaac SIM uh sorry

27:37

visualization from that top camera uh

27:39

and pro applying cosmos transfer to

27:42

generate realistic or the goal is to

27:44

generate realistic uh imagery. So we

27:46

again don't have to uh create those

27:48

images even from a real setting. But

27:50

what you'll see if you look look closely

27:51

at the the hands at that top image

27:54

there's actually Cosmos transfer put

27:55

hands on the robot. Now our robot

27:57

doesn't have hands but you know I guess

27:58

it that's where it's its background is

28:00

coming from the world foundational model

28:01

and so there is a need to fine-tune that

28:04

with real world data. So this is still a

28:06

work in progress. It's it's it's

28:07

definitely not solved yet. The the best

28:08

result we've gotten yet is in the lower

28:10

video. It has improved the drape um some

28:13

of the elements of the skin but it as I

28:14

said it still has some way to go. uh but

28:17

I do hope to be back here next year and

28:18

showing results from that. Um we then

28:21

take the training data sets the

28:23

synthetic and the real and we start to

28:25

apply training and so there's there's

28:26

two or apply them to training uh the

28:28

there's two different sort of problem

28:30

statements here. One is the on the left

28:31

you got the global environment where you

28:32

got the bed the patient the the robot

28:35

that's one set we're doing and the other

28:36

side of it is actual perspective from

28:38

the camera uh because you needed to

28:40

teach the system to be able to perceive

28:41

the scene and then decide what to do

28:43

next. And our goal here is obviously to

28:45

identify an optimum pose for a given

28:46

clinical scenario.

28:48

Um from a deploy perspective, as I said,

28:50

you know, regulated environment, we do

28:52

have to take things slowly. What we'll

28:54

start with is um a workspace

28:56

optimization button in our guey. And so

28:58

if the system is able to perceive a

28:59

scene that meets certain criteria and

29:01

does have has identified an appropriate

29:04

pose, this button will appear and you'll

29:05

press it and the system would then would

29:06

then deploy to that location. Um but

29:09

where we're going next and you know as

29:11

Axel and Filipo touched on is is

29:12

training vision language action models

29:14

with this uh with this data sets uh and

29:17

deploy them directly on Maestro. And so

29:18

we we we've we have just gotten just

29:21

recently um a trained version on group

29:23

ed 1.6 and Sean told me yesterday it's

29:25

already obsolete so we have to uh talk

29:27

to the guys about that and things are

29:29

moving pretty fast. Um but that's it's

29:31

exciting to think of what we will be

29:32

able to do and again I hope I'm back

29:34

here next year to talk about the results

29:35

on that. So, so lastly, you know, where

29:37

are we going? And you know, I listed it

29:40

as VA driven capabilities and

29:41

intentionally slightly vague there

29:42

because I think there is definitely a

29:44

hierarchy of uh tasks that we'll tackle

29:47

based off the risk to the patient and

29:48

the surgeon. But we do have this data

29:50

set of cadaavver data where we can do

29:51

very controlled experiments, synthetic

29:53

data, real world data um curating that

29:55

in the LR robot format as the open H

29:57

initiative is harmonized around. We're

29:59

training both said we we did contribute

30:01

to the open H data set as well in group

30:03

and group N1.6 six and then lastly the

30:05

the evaluation and deployment and I draw

30:08

your attention to that image on the left

30:09

of the valent deploys. I took it from a

30:11

presentation I gave a couple of weeks

30:13

ago at ORC and the at the ORC surgical

30:15

AI day and wanted to call out that in a

30:17

regulated environment we will absolutely

30:19

need to consider the safety and system

30:21

architecture to deploy these things and

30:23

the AI will sit in the middle of it but

30:24

there will be nested loops uh that

30:26

you're making sure that the system is

30:28

doing exactly what it needs to do to be

30:29

safe for the surgeon and the patient. Uh

30:32

and that's it. Thank you very much.

30:34

>> [applause]

30:43

>> Hi. Hello everybody. I'm Yossi, CEO and

30:46

founder of Lamb Surgical. Lamb is a

30:49

Swiss company. Our software director was

30:51

supposed to be here today. He hurt his

30:53

back. Couldn't sustain the flight from

30:55

Switzerland. So, uh, you're stuck with

30:57

me.

31:00

We are

31:02

we are FDA cleared and uh small

31:04

disclaimer some of the things that I

31:06

will show our future development.

31:11

This is uh say our mission reason we

31:14

wake up in the morning. Um as you can

31:16

see where we are now. So red bar is

31:21

elderly population in the US above the

31:24

age of 65. The blue bar is available

31:28

surgeons per

31:30

population. So as it appears in 2025

31:34

we're okay only small number to

31:36

remember. Today 40% 40% of US orthopedic

31:41

surgeons are above the age of 60. So

31:44

take it four or five years down the road

31:46

and then you see the perfect storm

31:48

coming few years from now. there will

31:50

not be enough surgeons in our case

31:53

orthopedic surgeons to treat the elderly

31:55

population and something need to be done

31:59

this session is about physical air some

32:01

something need to be done which is not

32:03

incremental software improvement

32:05

something need to be done in the

32:06

physical world to really accommodate for

32:09

this uh coming problem I think also

32:11

Mustafa mentioned this

32:14

and many are talking about it this is

32:16

quick uh I think all of the sessions

32:18

until now were in the soft tissue area

32:20

and really the last two decade the soft

32:22

tissue robotics were the most dominant

32:25

we are in the hard tissue robotics hard

32:27

tissue spine long bones joints and there

32:31

in the hard tissue the robotic

32:33

architecture is very different from what

32:36

you're accustomed to see in soft tissue

32:38

robotics is most of the time remote

32:41

manipulation moon is a is an interesting

32:45

I don't know to call it interesting case

32:47

but in a a a large form factor of um

32:52

surgical robotics. In soft tissue, it's

32:54

two to four robotic arms. The surgeon is

32:57

not even touching the patient sitting in

32:58

a console and it's a remote

33:00

manipulation. It's a a teleoperated

33:02

procedure. In hard tissue, it's a bit

33:05

different, bigger, stronger robotic

33:08

arms.

33:10

And the the most dominant form factor

33:13

and there's only few examples here is

33:16

usually it's one robotic arms arm coming

33:19

from the bedside merged with a what is

33:21

called a navigation camera most of the

33:24

time in infrared camera and the basic

33:26

idea is there it's concept invented 30

33:28

years ago. The idea is one bone marker

33:32

attached to the bony anatomy, one bone

33:34

marker in the end defector of the robot.

33:36

And what you see here is this uh you

33:40

don't really see it on your screen.

33:41

Never mind. You see this stationary

33:43

camera. It's usually an infrared camera.

33:46

So this camera looks at at these markers

33:49

and determine where the end effect or

33:52

the tip of the robot is in relation to

33:54

the bone. And then this robot can assist

33:56

the surgeon in hard tissue. usually is

33:59

to drill to apply saw to cut planes and

34:02

this is how it is done today. The

34:04

equivalent for this is is if I would

34:07

tell you imagine yourself trying to

34:09

hammer a nail to the wall with one hand

34:12

because this is what you have today. One

34:13

robotic arm and you're trying to hammer

34:15

a nail to the wall or to tie your

34:18

shoelaces. Never mind the analogy. One

34:20

arm one robotic arm is very limiting.

34:24

It's a very limiting architecture. That

34:26

that is why

34:28

While you can see today in intuitive and

34:31

soft tissue robotics really

34:34

um proliferated significantly in hard

34:37

tissue robotics it's a different

34:39

picture because these robots are

34:41

relatively limited can do only simple

34:44

tasks.

34:46

So what do we do in LM? We bring a first

34:50

of its kind upper torso humanoid

34:53

architecture meaning two robotic arms

34:56

synchronized with a third arm which is

34:58

the vision and then you can have the

35:00

human architecture that you are so

35:03

familiar with and use use it every day

35:05

without noticing it your uh something so

35:08

simple like to tie your shoelaces again

35:10

try to do it with one hand and this

35:13

architecture is missing in heart tissue

35:14

so two operating arms

35:17

dynamic vision that is synchronized with

35:20

the arms. Don't have enough time to

35:22

explain why is it important. Uh one of

35:24

the key elements is instrument and

35:27

implant agnostic. What does it mean? The

35:30

ability to operate any tool and not only

35:32

proprietary tool that we know um the

35:37

equivalent is you if I give you 30

35:40

different markers or pens, you can

35:43

operate all of them. No problem at all.

35:45

You take it and within a second you

35:47

calibrate it and you operate it. In

35:49

robotics,

35:51

robotics, heart tissue robotics, this

35:52

capability doesn't exist. You can

35:54

operate with the million-doll robot that

35:56

you bought only my proprietary implants

35:59

or instruments and this is a very

36:02

limiting thing and we are also doing

36:05

this where the idea is last bullet. We

36:07

want to have the pathway towards

36:09

supervised autonomous surgery. We want

36:10

to raise the bar so the robot can do

36:12

more sophisticated tasks and to automate

36:15

parts of the procedure.

36:18

Um how does it look? So let me operate

36:22

like this. [snorts]

36:23

So now you have this upper torso

36:26

humanoid architect the two operating

36:27

arms. So you have the two arms one in

36:31

this case that you see one can stabilize

36:33

a specific bone monitor it have

36:36

plurality of cameras and sensors. So

36:38

while the other arm doing what you see

36:40

here as as a bone milling the specific

36:44

bony element is controlled stabilized

36:46

and then you can have predictable

36:48

results. This is the reason why current

36:50

robotics cannot do it. Uh you need to

36:53

have predictable results and that's what

36:57

we're doing. Let me jump this forward.

37:00

Now where's the where's the challenge

37:02

and uh where we are cooperating with

37:06

Mustafa and the guys in Nvidia. One of

37:09

the things that at least occurred to me

37:11

at some point that a simple robot can do

37:14

simple tasks but also can do simple

37:15

mistakes. That's relatively easy for the

37:18

surgeon you have one arm come stand here

37:21

do this very simple. Now you have this

37:23

sophisticated robot that can do multiple

37:25

things and and the equivalent as as a

37:29

father to small girls, the equivalent

37:31

you you give your your your kids a fork

37:34

and knife and you you put it in the hand

37:36

in your in their hands. You say eat

37:38

nicely. What does it mean? So you say,

37:40

"No, no, you see, so you tell them this

37:42

is how you hold the fork. This is how

37:44

you hold the knife and eat your pasta."

37:46

Okay? So they learn to do it and but now

37:48

it's not pasta, now it's rice. So it's a

37:51

complete different action and now it's

37:53

different food. So and I'm not even

37:55

talking about more complicated thing

37:57

than a fork a fork and knife. So this

37:59

training is is a challenge and it's an

38:02

iterative

38:04

process that needs to be done and uh and

38:07

again I go back to our situation

38:10

when you have this three arms system and

38:14

you want to deploy it in the field

38:17

this system needs to be had to needs to

38:19

have a basic training of how to behave

38:22

how to move and one thing that

38:25

differentiate

38:27

hard tissue significantly from soft

38:28

issue is the the fact that it is a

38:32

collaborative robot. There is people

38:34

always there between five to five to 10

38:37

people working together with the robot.

38:39

So I will do this uh part uh fast

38:42

because I see I'm running out of time.

38:44

So uh I I will jump this forward. But

38:48

the idea is you have this complicated or

38:51

or very capable robot cooperating with

38:53

several people around the table and you

38:56

need to train it. One second. Sorry. I

38:58

will skip this quickly.

39:01

There you go. So, uh, one of the things

39:04

that at least in our list we first to

39:07

solve before we go to the clinical

39:09

application and to the bone and to the

39:11

tissue, how this robot can be can

39:14

collaborate with the people around.

39:16

There are people like the one that you

39:18

see on the right as sterile person.

39:20

Okay. So, the robot can come close even

39:22

touch it. But the the other guy is

39:24

non-sterile. You cannot come close to it

39:28

when you're sterile. There are

39:30

instruments, sharp instruments. Again,

39:33

one thing very common to a heart tissue

39:36

surgery, orthopedic surgery, hundreds of

39:38

different instrument, not your

39:40

instruments, other vendors instruments.

39:42

So, this robot needs to know how to

39:44

collaborate with all these people going

39:46

around you, the instruments that are

39:49

involved always in the surgery and then

39:51

the clinical application. So with

39:54

current

39:56

decision matrix we cannot this the

39:58

number of variable do the math is is

40:03

not scalable this solution and here AI

40:06

is something that we see with a high

40:08

value how to create what the guys can

40:11

explain or explain better than me to

40:14

create this uh and I will skip the

40:16

flowchart but to create this simulator

40:18

create this synthetic data and now train

40:21

the robot in a scale scalable way. So in

40:24

the first time that we deploy it in and

40:26

I'm finishing here. The first time that

40:28

we deploy it in the O, it has the basic

40:31

ability, the basic knowledge of how to

40:33

behave so the surgeon and the surgical

40:37

team can uh collaborate with it

40:39

successfully. Thank you very much.

40:43

[applause]

40:48

Just uh a couple of quick questions if

40:51

there is any. The amazing speakers are

40:55

here. Uh Jade is going to pass the

40:57

microphone and you guys can ask

40:59

questions. Just two short questions

41:01

please if there is any.

41:06

>> This is all

41:08

that should be on. Okay.

41:10

>> I'm Jonathan Strong. I'm actually a home

41:12

and community based provider up in

41:14

Alaska. Are there any robotic stuff

41:16

going on in home health care? Cuz that's

41:18

where I'm at. But anything that you guys

41:20

have heard of in home health care to

41:22

help like elderly and disabled people in

41:24

in their own homes? Anything like that

41:26

you've heard of?

41:30

>> Yes, there there are companies that are

41:32

working on those. Yes. Uh across the

41:34

board you can think of like from even

41:36

infection control to delivery within the

41:39

hospitals and beyond, right? And then

41:41

you can even think about exoskeleton

41:43

type of robots that are going to help

41:44

you and that is not within the

41:46

confinements of the hospital. But yes,

41:48

there are the robotic uh companies

41:50

working on that and the tools and stack

41:53

is the same, right? You still need data,

41:55

you still need to train.

41:59

>> One more question.

42:01

>> Hey there. Um I work with a lot of

42:04

startup companies and uh leveraging the

42:06

Jetson Thor platform. I'm curious on two

42:09

ways. is when I help out with smaller

42:11

companies who don't have simulation data

42:12

and are trying to leverage that to train

42:14

models uh are they able to access like

42:16

the open h or different things to have a

42:18

data set to work off of that's public

42:21

and then secondly you know when you guys

42:23

are using the Jetson for the multi-ensor

42:26

input processing is that running the

42:28

full um like learning process on top of

42:31

those as well or do you have to have two

42:33

machines communicating

42:34

>> two two machines communicating but

42:35

that's a great question because our next

42:37

session and you see Jay here uh please

42:41

be seated uh or just take a quick break

42:44

and come back. These are the same

42:45

sessions back to back. We're going to go

42:47

in depth about the runtime computer and

42:50

kind of differentiating between what you

42:52

need for the runtime versus what we

42:53

would need in the training.

42:56

That's all the questions we have uh

42:58

right now. What we're going to do is

43:00

we're going to get these guys off, get

43:02

uh next speakers miked up. everyone can

43:04

kind of stand up, stretch their legs,

43:06

and uh you know, feel free to mingle,

43:08

mix and mingle between sessions. And um

43:11

we'll see you back here in at the top of

43:13

the hour. Thank you so much.

Interactive Summary

The session introduces healthcare robotics as a critical application for physical AI, addressing a massive demand and supply crisis. The primary challenge identified is data (curation, annotation, generation), which can be overcome through foundational models, generative physics, and accelerated simulation. Speakers from Northwell Health, Johns Hopkins/Asensus Surgical, Moon Surgical, and Lamb Surgical then detail their innovative approaches. Filipo Fukori highlights the use of robotic data for surgeon training and the challenges of data access and curation, leveraging video language models and Cosmos for better simulations and surgical scene understanding. Axel Krieger from Asensus Surgical discusses moving from model-based to learning-based autonomous surgery, showcasing high-fidelity performance on fundamental tasks and procedure-level autonomy using hierarchical policies and large datasets like OpenAGE. Jay from Moon Surgical presents the Maestro system, a collaborative robot leveraging edge compute and simulation for optimized arm deployment, emphasizing safety in a regulated environment. Yossi from Lamb Surgical introduces a novel upper torso humanoid architecture for hard tissue robotics, aiming to overcome the limitations of single-arm systems and enable supervised autonomous surgery, focusing on scalable training for complex human-robot collaboration. The session concludes with a brief Q&A, touching on home healthcare robotics and data access for startups.

Suggested questions

7 ready-made prompts