HomeVideos

GTC SJ 2026: Next Generation Intelligent Surgical Robots

Now Playing

GTC SJ 2026: Next Generation Intelligent Surgical Robots

Transcript

1109 segments

0:05

Hello everyone. My name is Mati Azizzian

0:07

and I'm pleased today to talk to you

0:09

about the next generation intelligent

0:11

surgical robots.

0:12

>> [clears throat]

0:13

>> Uh here's a brief overview of the agenda

0:15

we're going to go through. Uh first

0:16

we'll take a look at the state of the

0:18

art surgical robotic systems. How do

0:20

they look like? And then we talk about

0:22

uh

0:23

the next generation of surgical robots

0:24

which we think will be data-centric uh

0:27

architectures. And then we briefly touch

0:29

on the pathway from development to

0:31

deployment and what are the computer

0:33

requirements for that

0:35

and how to get from clinical decision

0:37

making uh to AI that is used in action.

0:40

Then I'll pass it over to my colleague

0:41

Shawn Hoover who will talk about data

0:43

and next generation models and some of

0:45

the announcements that you have seen

0:47

already uh at NVIDIA. He will dive deep

0:49

into those. And then uh we'll basically

0:52

uh go over

0:53

agentic systems in uh smart hospitals

0:56

and then we'll summarize and uh wrap it

0:58

up for some uh questions and answers.

1:02

So if you look at the surgical robots of

1:04

today, uh you can categorize them from

1:07

uh a few different angles. One is by the

1:09

access method to the human body.

1:12

Um obviously open access uh robotics has

1:14

been around for quite a long time in

1:16

orthopedics, cranial, spine um

1:18

applications.

1:20

Multiport robotics in laparoscopy,

1:21

thoracoscopy, arthroscopy have been uh

1:25

around for the past three decades or so

1:27

and uh and in use. And single port

1:29

robots are a little bit uh more uh

1:33

basically

1:34

uh recent uh with the likes of da Vinci

1:37

SP and equivalents of those that are

1:39

being used. And endoluminal robots,

1:40

there's a wide range of things from

1:42

notes to endobronchial, endovascular,

1:45

and uh different types. And incisionless

1:48

robotics, uh the types of focal therapy,

1:50

um HIFU, histotripsy, and those. And

1:53

each of these basically have different

1:54

types of requirements.

1:56

Uh so mostly we're going to focus on the

1:58

left side of the

1:59

uh this screen with multiple single port

2:02

like systems.

2:03

But the phenomena that we talk about

2:05

applies to all of these basically.

2:07

You can also categorize by embodiments.

2:10

Uh so you can envision

2:11

a surgical robot could be

2:13

uh basically multiple arms on on a

2:15

single cart on wheels. This is probably

2:18

one of the most popular

2:20

types of systems that is out there. Many

2:22

of these are manufactured by different

2:24

uh

2:25

medical device companies.

2:27

Or it could be actually on multiple

2:28

carts on wheels that actually get around

2:30

the operating room table.

2:32

Uh it could be also mounted on the

2:34

operating room table

2:36

like the Ottawa system from J&J. Or

2:39

there are different types that are

2:40

basically bolted to the floor or to the

2:42

ceiling. Especially focal

2:44

therapy

2:45

robots have those type of embodiments a

2:48

lot.

2:49

Uh the other categorization you can look

2:51

into is actually from

2:54

the levels of autonomy.

2:55

This is actually from a paper in Nature

2:58

by Band Rapport and team from a couple

3:01

of years ago.

3:02

Uh where they have built on top of the

3:04

seminal paper from Guang-Zhong Yang and

3:07

team in Science Robotics in 2017 where

3:09

they tried to define and get into a

3:11

consensus on different levels of

3:13

autonomy for surgical robotics.

3:15

Uh the green curve that you see over

3:17

there shows basically how much of the

3:20

existing robotic systems are in each

3:22

category and majority as you can see are

3:24

in level one which is uh robots that

3:27

assist you mostly teleoperated robots.

3:30

And obviously

3:31

there are few that are in level two and

3:33

level three with task autonomy and

3:35

conditional autonomy. Uh but there are

3:37

very few and mostly they're in

3:40

rigid body type applications. There's

3:42

pretty much nothing commercially

3:43

available with level four or level five

3:45

autonomy. But those there's a lot of

3:48

research groups and startups working on

3:50

that and those will be coming in future

3:53

as well. And some of the architectures

3:54

we are talking about is actually

3:57

enabling not only the level one systems,

4:00

but also enabling the higher level

4:01

autonomy systems.

4:04

So, a typical surgical robotic system,

4:07

a tele-operated one,

4:09

has one or two surgeon consoles

4:11

that the primary and secondary surgeon

4:13

are sitting at there. That could that

4:14

could be a trainee and a mentor or

4:17

basically two surgeons were sharing

4:19

responsibilities or just a fellow who's

4:22

actually watching the surgery.

4:24

There's usually a vision tower that's

4:26

endoscope and all accessories are

4:28

connected to and a surgical robot. So,

4:30

this is a very simple architecture of

4:31

what exists today. Often the modern

4:33

surgical robotic systems have a

4:34

connection to the cloud, a virtual

4:36

private cloud that is used for

4:39

over-the-air updates,

4:41

maintenance, collecting data, and

4:43

sometimes basically

4:45

also

4:46

kind of remotely connecting to the

4:47

system and doing things like for example

4:50

tele-mentoring.

4:51

And we see more and more of edge AI

4:54

sidecars being added to these systems as

4:56

an afterthought to bring

4:59

AI capabilities to the edge of surgical

5:01

robotics.

5:02

Uh

5:04

If you look at

5:05

a little bit more details of how

5:07

different components of such a system

5:09

are connected to, this is just a

5:11

representation and an oversimplified

5:14

way of talking about how different

5:16

components of a surgical robot are

5:17

connected to.

5:19

There are many different components.

5:20

This is by definition a distributed

5:22

system.

5:23

Uh

5:24

And each of those are connected with

5:27

often proprietary

5:29

type of connections to each other. So,

5:30

the surgeon console

5:32

to the vision card, to the

5:35

surgical robot itself, and the

5:37

connection between different parts of

5:39

that. Sometimes there are star

5:41

configurations or daisy chains with

5:43

different types of uh,

5:45

basically cables, signals, protocols.

5:48

And this makes it extremely difficult to

5:50

actually move from one generation of a,

5:52

robotic system to the next generation

5:55

because you basically have to maintain

5:57

something that A to Z is proprietary, is

5:59

not interoperable with other things, and

6:02

uh,

6:03

often you deal with obsolescence of

6:06

basically parts that are supporting

6:07

that.

6:09

So,

6:10

uh, that's one of the reasons we're

6:11

actually thinking that

6:13

>> [clears throat]

6:13

>> the next generation of surgical robots,

6:15

like many other robots, are going to be

6:17

data centers on wheels or data centers

6:20

bolted to the operating room depending

6:22

on what embodiments that you're talking

6:24

about. Data centers are very extensible.

6:27

Uh, let's take a look at few parallels

6:29

from different industries. If you look

6:30

at the autonomous vehicle industry,

6:32

especially in the Bay Area, you see a

6:33

lot of startups who are working on

6:36

uh, AV. Some of them now are big

6:38

companies that have products in the

6:40

market. You probably have

6:41

uh,

6:42

took a ride with Waymo here.

6:44

Uh, that you have seen how it works, but

6:47

often,

6:48

uh, if you have the inner guts of a

6:49

prototype, it's actually that ugly, you

6:52

know, in the trunk. This is a trunk of

6:53

a, uh, self-driving car under

6:55

development. You see very different

6:57

types of cable, cables, different

6:59

protocols, and uh, different

7:01

connections. Extremely difficult to

7:04

maintain and deal with multiple

7:05

different vendors, uh, and uh,

7:08

complexities that comes with that. And

7:10

here is actually a view of a

7:12

data-centric, uh, autonomous vehicle.

7:14

So, you can see actually that the ECUs,

7:18

or the brain of the the car itself, is

7:21

basically connected over Ethernet to

7:23

pretty much all the sensors and

7:24

actuators in the car. And that enables

7:27

you to easily expand if you want to go

7:29

from one generation to the next

7:30

generation, add more sensors, upgrade

7:32

your sensors, change your actuators,

7:35

increase the capacity of compute. This

7:37

is simply a data center. You just change

7:39

the compute, you You the sensors and

7:40

actuators, and everything stays the

7:42

same. It's all ethernet based.

7:44

It's interoperable.

7:46

Makes you capable of actually working

7:48

with different vendors and even have a

7:51

basically a system that is built by

7:54

multiple different companies. Same goes

7:57

with humanoids. If you look at legacy

7:59

humanoids, it's actually

8:01

kind of funny to call humanoids legacy.

8:03

They're They're newer by definition, but

8:06

what is in the market has been following

8:08

what has existed in

8:10

kind of robotic manipulators and also

8:12

AV.

8:13

So often there's fixed input and output

8:15

that limits basically the sensor count

8:16

that you have. You can't easily add more

8:19

sensors.

8:20

You have separate buses for sensing and

8:22

actuation and then you basically end up

8:25

with something complex and

8:27

expensive to deal with. You have limited

8:29

safety and fault tolerance and there's

8:32

heavy usage of CPU for handling the

8:34

traffic versus a data-centric humanoid

8:37

design that we're

8:40

basically working on with some of our

8:41

partners

8:42

is all ethernet based. Again, you have a

8:45

brain of the

8:47

of the humanoid, that green box in the

8:48

chest of the humanoid, that has flexible

8:51

IO over ethernet. So you're only limited

8:53

by the bandwidth of the ethernet links

8:55

that you put in there. You have a

8:57

unified network for sensing and

8:58

actuation. Everything is synchronized on

9:00

the same network.

9:02

This

9:03

drastically reduces the complexity and

9:05

makes it more fault tolerant and also

9:08

because of the network acceleration that

9:11

basically

9:12

brains of systems bring in that we'll

9:14

talk about, it offloads the CPU usage.

9:17

So you basically have direct access in

9:20

and out of GPU memory.

9:22

So

9:23

here's actually

9:24

a view of that, a a very simplified view

9:27

of what a data-centric surgical robot

9:28

would look like. You can actually see in

9:31

the bottom

9:32

that you can have different edge

9:33

computes for different zones of a

9:35

surgical robot for lack of a better

9:37

word. You have a simulation zone,

9:39

uh surgical console zone, a robotic uh

9:42

zone, endoscopy zone, and accessory

9:44

zone.

9:45

You have a backbone of uh basically an

9:47

Ethernet

9:49

uh

9:49

uh network, and you can have a redundant

9:51

backbone as well for safety purposes.

9:54

And pretty much all the input-output in

9:56

the system is dealt with uh Ethernet. Uh

9:59

Host and sensor bridge that we briefly

10:01

touch upon

10:02

after this enables you to pretty much

10:04

convert any type of signal in uh and out

10:08

of uh Ethernet and basically make that

10:11

RDMA-able. RDMA stands for remote direct

10:13

memory access to uh GPU memories. So,

10:16

we'll uh double down on that a little

10:19

bit.

10:19

>> [clears throat]

10:19

>> So, let's get a bit geeky here for those

10:22

of you from Medtronic. If you want to

10:24

look at the highest level of

10:25

requirements for a data-centric robotic

10:27

system,

10:28

uh you want the robotic network uh to be

10:30

end-to-end Ethernet-based.

10:32

And you want to uh have the data

10:35

accessible in real time anywhere within

10:37

the system, so you don't have to

10:38

redesign things to get, for example, the

10:41

endoscopic image on the other node or

10:43

this joint encoders basically in another

10:45

node. Everything should be accessible

10:46

everywhere with software configuration.

10:48

Ingress data from sensor should be

10:50

accessible in any processing unit with

10:52

minimal latency and jitter. That's uh

10:54

basically a requirement,

10:56

>> [clears throat]

10:56

>> and vice versa. Egress data from a

10:58

processing unit should be accessible to

11:00

controller units, whether it's a motor

11:01

controller or display controller, again

11:04

with minimal latency and jitter.

11:05

There shall be redundant access passes,

11:08

obviously, for safety-critical data, and

11:10

all of ingress-egress data should be

11:12

synchronized, obviously.

11:14

You want to have data authenticity and

11:17

security. Uh that's an obvious

11:19

requirement because you don't want

11:20

someone to connect with their laptop to

11:22

the system and and try to uh hack into

11:24

that.

11:25

And switching uh this one is probably uh

11:28

one of the most important features that

11:30

this enables, that switching between a

11:32

digital and physical twin of a robot

11:34

becomes seamless with this type of

11:36

architecture. With the change of IP

11:37

address and port, you'll be able to

11:39

actually switch from a digital version

11:41

of a sensor to its physical version. So,

11:44

development cycles become much faster

11:46

because of that. Uh you don't need to

11:48

create

11:49

uh hardware abstraction layers and uh

11:51

basically make ad hoc simulation

11:53

bridges.

11:54

The processing units obviously shall be

11:55

able to record data and playback that

11:58

data seamlessly with synchronization and

12:00

consistency. So,

12:02

let's take a closer look into

12:04

performance. Uh I'll just give you a

12:05

very simplified example of that. Imagine

12:08

that

12:09

you want to talk about uh

12:10

event-to-action performance on a

12:12

surgical robot, which is the core of

12:14

what it does. You know, a teleoperated

12:15

robot is detecting something, you take

12:17

an action based on that and and go

12:19

through. [snorts] So, if you have a

12:20

chip-on-tip sensor, imagine an

12:21

endoscope, in a regular uh

12:25

surgical [clears throat] robotic system,

12:26

you have it go through an ISP, an image

12:28

signal processing pipeline, going

12:29

through a video router,

12:31

uh to a frame grabber often on a

12:33

different physical node, to a processing

12:35

unit doing some processing on that, and

12:37

then to a controller unit, and then back

12:40

to a motor controller, to a robotic arm

12:42

to do something. And I I'm not showing

12:44

surgeon in loop here just to be able to

12:45

talk solid about the

12:48

uh performance that you get. The best

12:49

things that you find in the market for

12:51

this is like 50 ms uh latency

12:55

from basically

12:56

>> [clears throat]

12:56

>> the start to the end. Um

12:59

uh in terms of what's uh performance

13:01

that you get. The data-centric version

13:03

of this will be something like this,

13:05

that you have uh the chip-on-tip sensor,

13:07

the raw data over, say, CSI MIP, goes

13:10

through whole scan sensor bridge

13:11

directly written to GPU memory. So, as

13:14

the lines are being uh scanned in your

13:16

camera, they're basically written in GPU

13:19

memory. In another word, the GPU memory

13:20

becomes extension of your sensor memory.

13:23

And then it goes to a whole scan SDK

13:24

pipeline on the GPU, which is basically

13:26

a graph for processing, and then you can

13:29

basically

13:30

>> [clears throat]

13:30

>> go back out over Ethernet to a

13:32

microcontroller that is controlling your

13:34

robotic arm.

13:36

And benchmarks on this that we have done

13:37

shows a speed of light implementation of

13:40

1.2 milliseconds from the input to the

13:42

output. So, you see the difference that

13:43

you get uh going from 50 to uh 1.2 uh

13:48

milliseconds uh on this.

13:50

>> [clears throat]

13:50

>> How is this enabled?

13:52

Uh we have a smaller version of of uh

13:55

our data centers on the edge for these

13:57

type of robotics deployments. This is

14:00

>> [clears throat]

14:00

>> the high-end version of that called IGX

14:02

Tour T7000. There are different versions

14:05

of it with uh basically different levels

14:08

of compute and and price points,

14:09

obviously.

14:10

Uh this one brings over 5000 uh

14:14

uh teraflops of FE4 uh for inference,

14:16

and you have the option of adding an RTX

14:18

Pro 6000 if you need more uh compute.

14:22

Uh you get 128 uh GB of memory, and more

14:25

importantly, you basically get up to 400

14:28

Gbit/s of Ethernet bandwidth to this,

14:30

which is twice of the previous

14:32

generation, which was IGX Orin

14:35

uh 700. And then next generation, it'll

14:37

be doubled again, it'll go to 800

14:39

Gbit/s. So, without doing anything, this

14:41

data-centric uh approach will give you

14:43

future proofness.

14:45

Uh

14:46

the IO parts is enabled by Holoscan

14:49

Sensor Bridge. This is a very simplified

14:51

uh reference design of that.

14:53

Uh it is an FPGA IP that we have made

14:56

public and freely available for everyone

14:58

who wants to put it in their FPGA. It's

15:00

pretty much supported uh by any FPGA

15:03

vendor. Uh there are many of our

15:05

partners who are building ASICs out of

15:06

that, and

15:08

also it comes with a software emulator

15:10

that you can run it on a microcontroller

15:12

or on any CPU, basically.

15:14

Uh so, you can convert pretty much any

15:16

type of uh signal that you get from I2C,

15:19

SPI, and GPIO to just the maybe CSI,

15:23

LVDS and so. And on the other side you

15:25

will have ethernet in and out uh time

15:28

synchronized if you have multiple of

15:29

these and directly accessing uh GPU

15:32

memory. This is enabled by NVIDIA

15:34

networking acceleration

15:36

uh provided with with ConnectX.

15:39

So, this is the

15:40

1.2 [clears throat] millisecond pipeline

15:41

that I was talking about.

15:43

If you look at uh the top left you see

15:46

the stereo camera connected to a whole

15:48

scan sensor bridge implemented on a this

15:50

is a Lattice ECP3.

15:52

It goes over an ethernet link to an IGX

15:55

and gets processed in a whole scan

15:57

pipeline. Uh what you see on uh the

16:00

right side of the screen, the green is a

16:02

laser pointer controlled by a uh by a

16:05

human and then the red is basically what

16:07

that galvo um

16:09

uh laser pointer with at a two degree of

16:11

freedom robot is uh is uh basically

16:15

shining on the wall to track the

16:18

Let me try to play it again if

16:20

to track basically the green. And you

16:22

can see that's almost perfectly tracking

16:24

that. The little offset is uh because of

16:27

calibration.

16:29

Uh so, that 1.2 millisecond is getting

16:32

into uh the basically from shutter close

16:34

to

16:36

uh the GPIO command sending to uh the

16:38

MCU. And the photon to photon here is 10

16:41

milliseconds which is mainly driven by

16:44

the time that the camera shutter

16:46

basically needs to be uh to be open.

16:49

Uh

16:50

if you look at this obviously in a

16:51

regular teleoperated robot this can be

16:53

used but also

16:55

you hear a lot about a lot about tele

16:56

surgery because um robots and surgeons

16:59

are both scarce and you want and

17:01

surgeons are more scarce and you want to

17:03

have the surgeons be able to operate in

17:05

remote locations. So, this is a a very

17:08

simplified uh kind of diagram of showing

17:11

how state of the art tele surgery works.

17:13

Uh basically on the patient side you

17:14

have the endoscope being processed and

17:18

compressed, sent over the network,

17:19

displayed to the surgeon, surgeon taking

17:21

an action, that action goes again over

17:23

the network back to the other side, and

17:25

the robot follows that.

17:26

How this can be accelerated using this

17:29

type of data-centric technology is that

17:31

you can actually apply Holoscan sensor

17:33

bridge and the Nvidia acceleration to

17:35

basically shove off latency on both the

17:37

patient side and the surgeon side,

17:39

because there's not much that we can do

17:41

about network. You can

17:43

obviously work with the likes of

17:45

Intuitive and other telesurgery

17:46

companies and telecom companies to get

17:48

the best network possible,

17:50

but speed of light is limited based on

17:52

the distance that you have. What we can

17:54

do is shove off latency on basically the

17:56

proximal and distal side of the

17:57

pipeline, and give a better performance

18:00

and better safety to the surgeon who's

18:02

operating.

18:04

Let's quickly touch on the development

18:06

workflow and how

18:08

we're actually thinking about that. We

18:10

see this type of problem, especially as

18:12

you go to the higher levels of autonomy,

18:14

as the three computer solution problem.

18:18

What it means is basically

18:20

if you're bringing AI to the loop, you

18:22

need to train that AI algorithm. Often,

18:25

you need to actually involve some sort

18:27

of a simulation,

18:28

and then at the end of the day, you need

18:30

to deploy it. But this is actually a

18:31

cycle that keeps going on during your

18:34

product development and after the

18:35

product launch as well. So you may want

18:37

to make sure that this cycle is very

18:40

fluid, and you have a flywheel running

18:42

over there. This data-centric approach

18:44

enables you to seamlessly, as I

18:46

mentioned, switch between simulation and

18:48

real world. So

18:51

regardless of if you do it on prem or on

18:53

cloud, you will be able to use the same

18:55

brain of the robot, say based on that

18:57

IGX,

18:58

working with a simulation environment

18:59

without knowing

19:01

that is simulated and not real. Shawn is

19:03

going to talk about the basically a

19:06

modern approach to surgical simulation

19:09

next after this.

19:10

And

19:11

you can actually envision that once you

19:13

have this uh ready for deployment, you

19:16

can implement it with a distributed

19:18

compute for layered intelligence having

19:20

the embedded intelligence in the robots

19:23

whether they're surgical robots or

19:24

cobots

19:26

in the operating rooms also having the

19:28

signals from the room like room cameras

19:31

on prime compute in the hospitals and

19:34

cloud compute and this will enable you

19:36

to actually scale the intelligence as

19:38

needed for your systems.

19:40

Clinical decision support has been the

19:41

main focus for many years with AI

19:43

development for surgery.

19:45

The likes of Cosmo AMD genius platform

19:48

or Medtronic digital surgery one have

19:50

been adding edge compute as a side car

19:52

to medical devices and then displaying

19:55

the results of for example detecting a

19:57

polyp in colonoscopy to the surgeon to

19:59

help them with the clinical decision.

20:03

But you can actually envision that AI

20:04

can be used in action closing the loop.

20:07

This is Moon Surgical scope pilot where

20:10

tracking of these instruments results in

20:12

detecting of where the surgeon wants to

20:14

look at and then autonomously moving the

20:16

endoscope to basically help them locate

20:19

that. With that I'm going to pass it to

20:20

my colleague Shawn to talk about the

20:22

data and the next generation models.

20:27

>> [applause]

20:31

>> Thank you Marty.

20:32

Yeah, so if you attended Kimberly's

20:34

keynote on Monday you've probably seen

20:37

kind of the high level of this but we'll

20:38

be talking about the data that we

20:40

released this week and how how that's

20:42

helping to fuel the next generation of

20:45

AI [snorts] models.

20:46

So if if you look at data available to

20:49

train different types of generative

20:50

models if you scrape the entire internet

20:53

there's roughly a billion hours of kind

20:55

of human existence for for training data

20:59

across math, across science, philosophy,

21:02

medicine etc.

21:03

Now if you look at what's available for

21:05

general robotics then that drops down to

21:07

on the order of tens of thousands of

21:09

hours. So, this is why Frontier Labs in

21:12

the general robotics space

21:14

pay people to come in off the street and

21:16

teleoperate robots for training data

21:18

such as folding t-shirts, doing the

21:20

dishes, etc.

21:23

Now, of course, when you look at what's

21:24

openly available for healthcare

21:25

robotics, you know, this is especially

21:27

dire. You know, if you're researcher, if

21:29

you're a PhD student, and you happen to

21:33

be at a university that does not have a

21:35

DVRK, you're kind of out of luck.

21:37

There's only 3 and 1/2 hours of open

21:39

source data. This is a 10-year-old data

21:41

set. That's it's still an amazing data

21:43

set put out by Johns Hopkins called

21:45

called the JIGSAWS data set.

21:47

So, when we were looking at, you know,

21:49

what we could do roughly a year ago in

21:52

the space, then, you know, we we decided

21:54

to reach out to some of the luminary

21:56

figures

21:57

in healthcare robotics,

21:59

folks such as Axel Krieger at JHU,

22:02

Nasir Navab at TUM, and we decided to

22:05

build a coalition that we call

22:07

Open Embodyment, this community-driven

22:10

effort. And I think kind of beyond our

22:12

wildest dreams, we grew from this

22:14

initial group of of three organizations

22:16

to 35 partners spanning industry,

22:19

academia, as well as healthcare

22:21

partners.

22:22

Filippo Filicori, for instance, here in

22:25

in the audience and his team of surgical

22:27

residents helping to do annotations. So,

22:29

really this great group came together.

22:33

Actually became, I think, an easy sell

22:35

to get people involved.

22:37

And so, the data set itself has eight

22:39

different robotic embodiments

22:41

represented, all sorts of different data

22:43

spanning from simulation to tabletop

22:46

to ex-vivo animal to even clinical. So,

22:51

CMR Surgical made an incredibly generous

22:54

donation of almost 500 hours of clinical

22:57

data. So,

22:59

you know, robotics kinematics paired

23:00

with with video.

23:02

So, this data set went live on hugging

23:04

face. I just checked right right before

23:07

this this morning. Already over a

23:09

thousand downloads of 4 terabyte data

23:11

set in just a few days.

23:14

So now that we have this great data set

23:16

to work with, you know, what are some of

23:17

the things that we can do? One of the

23:19

first models that we experimented with

23:22

building is a vision language action

23:24

model of VLA. So Nvidia has

23:28

series of models called Groot for

23:29

general humanoid robotics. So we took

23:33

building on that foundation. We then

23:36

trained on all of the open age data. To

23:39

see

23:41

what it could do in the lab

23:43

with our with our partners such as Johns

23:45

Hopkins. So this is, you know, being

23:48

2026 and AI moving as quickly as it is.

23:52

I'm going to refer to this type of VLA

23:54

architecture as a classical VLA meaning

23:57

that it's built upon a vision language

23:59

model at its base of VLM if you look in

24:01

the bottom left hand corner.

24:03

And

24:05

this is this is what we saw when we

24:08

pre-trained on all of the open age data

24:10

and then did a little bit of post

24:11

training on suture data from an effort

24:14

from JHU called suture bot.

24:16

So hopefully this will will play.

24:18

There we go.

24:20

So this is

24:22

running on a DVRK

24:24

example and this is

24:26

an end to end suturing.

24:27

So the model itself probably won't do

24:30

great at zero shot straight out of the

24:32

box, but it's been pre-trained on all of

24:34

these different embodiments on all of

24:35

these different tasks.

24:37

That it requires less data than you

24:40

would need otherwise.

24:42

So you know, this is a good example. The

24:45

pre-trained model is then post-trained

24:47

on just a little bit of the suture

24:48

suturing data. And what you see here is

24:50

a successful example. And of course full

24:53

disclosure, there were many unsuccessful

24:55

examples. We're quantifying the the

24:58

entire performance right now as part of

25:00

an upcoming paper.

25:02

So before I I I referred to this VLA as

25:04

a classical VLA because this is it's

25:07

it's almost out of date. We we we just

25:09

published this model and here's an

25:11

example of what we're looking at in the

25:12

future. Now VLAs rather than having the

25:15

vision language model as their backbone

25:18

have world foundation models as their

25:20

backbone. So if you're not familiar with

25:21

world foundation model is a generative

25:23

AI model that can generate video going

25:25

into the future. So not only is this

25:28

model from Semaphor Surgical which is a

25:31

company founded by Axel Krieger and some

25:34

other JHU

25:36

collaborators.

25:38

So not only is it generating the

25:40

kinematics of what the robot should do

25:42

in the future. It's also generating the

25:44

video. And some early papers are showing

25:47

that these types of models, this type of

25:48

architecture

25:50

is actually doing better than what we're

25:52

seeing with the classical VLA.

25:55

So now I'm going to switch gears a

25:56

little bit and

25:58

talk about rather than a VLA what world

26:00

foundation models can do as as far as

26:03

simulating surgical robots.

26:05

So this is our second model that we also

26:08

open source this week and it's called

26:09

Cosmos H Surgical Simulator.

26:12

So when world foundation models first

26:14

appeared they were typically text

26:16

driven. You would write a description of

26:18

what you want. You may have seen this

26:19

from Open AI's Sora and it would

26:21

generate a video describe

26:23

a video based on your description.

26:25

So this is different in that it's driven

26:26

by the kinematics of the robot. So you

26:29

give it a ground truth single frame

26:31

image and then the kinematics of of your

26:34

robot and it actually generates each

26:35

successive frame.

26:37

So this model has been trained on all of

26:40

the data from from Open H. So it's a

26:43

single model that can do eight different

26:45

embodiment simulations across the

26:47

different tasks that it was trained on.

26:50

What you're looking at is suturing and

26:52

one thing I want you to take notice of

26:53

here is again, we're just giving it the

26:56

first frame and it can actually keep

26:59

temporal coherence with how the ground

27:02

truth evolved over a few seconds and

27:05

that'll be important for what I talk

27:06

about a little bit later.

27:08

So again, here's kind of an example of

27:10

all different types of scenarios. These

27:12

are all holdout test sets of of what

27:15

this model can can simulate.

27:19

Um where we're taking this type of model

27:21

next is these models are getting closer

27:23

and closer to being real-time.

27:26

Currently,

27:27

you might need to

27:29

spend 10 minutes to generate 30-second

27:31

clip. Uh but we can distill these

27:34

models, make them smaller, and make them

27:35

much faster. So our fastest version that

27:38

runs on a single GPU today runs at about

27:40

12 frames per second on an RTX Pro 6000.

27:44

Uh but if you've been out on the the

27:46

show floor, there's a really cool demo

27:48

for a driving simulator

27:50

in in set of Cosmos. And so if you've

27:53

seen that,

27:54

that's actually running at almost 30

27:56

frames per second. Although full

27:58

disclosure, that's running on 12 Vera

28:00

Rubin GPUs. So it's a lot of expensive

28:03

firepower to run that. But of course,

28:05

this is this entire field is

28:07

accelerating so quickly that we're

28:09

reasonably confident that we'll be able

28:11

to have a real-time version within the

28:13

next 6 months. And once we have a

28:14

real-time version, you can actually

28:16

teleoperate these. So just like that

28:18

driving simulator has a gas pedal and a

28:20

steering wheel where you can interact

28:22

with the world foundation model, we'll

28:24

be able to connect happily devices for

28:26

instance and treat this as a surgical

28:28

simulator.

28:30

So

28:31

um

28:32

and uh

28:33

switching gears again, our third type of

28:35

model and this was just actually

28:38

released this morning uh is the

28:41

tokenizer for real-time surgery. Uh so

28:43

as tele surgery so as many of you are

28:46

familiar, in telesurgery,

28:48

uh the the the the the surgeon

28:51

uh controlling the robot

28:53

uh

28:54

you're use a classical encoder to encode

28:57

the video to try and make the the video

28:59

packets as small as possible,

29:01

uh reduce latency.

29:03

So, um just like uh tokenizers are used

29:06

for large language models, when you type

29:07

a request, uh a text tokenizer turns

29:10

your request into tokens where it lives

29:13

in a latent space that AI model

29:14

understands.

29:16

Uh we now have the same thing for for

29:18

video uh for for for telesurgery. So,

29:20

this has been trained on surgical video

29:22

data,

29:23

and these models are are are getting

29:25

really great at high compression and

29:28

high quality. Uh this is changing very

29:30

very quickly.

29:32

Uh so, this is an example of the model

29:34

that we uh open sourced and put in on

29:36

Hugging Face. Uh it was made public this

29:38

morning. We're really uh interested to

29:40

see what types of experiments people do

29:42

with this. And this is actually rivaling

29:44

JPEG compression in terms of quality

29:46

here.

29:48

So, if we put some of these technologies

29:50

that I just talked about together, I'm

29:52

going to describe real quickly kind of

29:54

uh this uh idealistic moonshot

29:56

experiment that we're hoping to do.

29:58

Uh so, having these neural tokenizers

30:00

means that we can put not just video

30:02

data, but in the future

30:04

uh all all types of multimodal data,

30:06

whe- whether it's audio, uh whether it's

30:08

the kinematics that all live in the same

30:10

latent space that generative AI models

30:12

can understand.

30:14

And so, if we feed those tokens into the

30:17

world foundation model that I described

30:19

earlier, as that gets to be real time,

30:21

and that temporal coherence is good

30:23

enough uh that we saw earlier, we might

30:26

be able to, you know,

30:28

predict what happens a few frames into

30:30

the second. And so, this this moonshot

30:33

experiment that we're hoping to do later

30:34

this year is to see, can we shave

30:36

latency off of telesurgery using these

30:39

technologies together?

30:41

So, again, it's it's it's very

30:42

aspirational, uh but this is kind of

30:44

where we're headed in in some of our

30:46

ideas.

30:48

All right, and I'll switch it give give

30:50

it back over to Madi for the next

30:52

section.

30:53

>> Thank you, Sean.

30:54

>> [applause]

30:59

>> Thank you, Sean. So, I'm going to spend

31:01

few minutes to talk about the agentic

31:03

systems in smart hospitals, how to

31:05

connect all of this intelligence that

31:06

the your system would architecturally

31:08

support and the models that Sean talked

31:11

about would actually bring it on the

31:13

edge. In 2024 at GTC, just 2 years ago,

31:17

we actually showed surgical agentic

31:20

co-pilot here.

31:22

It was mainly think of it as a chatbot

31:25

for the surgeon and the operating room

31:26

staff,

31:27

bringing all of the EHR and PACS data,

31:30

the surgical endoscopy

31:33

feed live basically into a VLM and also

31:36

loading a rag of the medical device

31:38

documentation, and allowing basically

31:41

the model to understand what is

31:42

happening in the surgery and providing

31:45

answers to any questions that people

31:47

have. For example, what phase of surgery

31:48

you're in, what instruments you're going

31:50

to need in the next phase,

31:53

what complications the patient might

31:55

face.

31:56

This year actually we have expanded and

31:59

built on top of that work and brought

32:01

physical agents to basically this

32:04

framework. If you look at the problems

32:06

in the that we have in the hospitals, a

32:08

lot of that is around workflow

32:09

efficiency.

32:11

Just a very brief overview,

32:13

every minute in operating room in US

32:15

costs around $46, and there are over

32:18

50,000 operating rooms. Just do the

32:20

math. Surgical suites represent 60% or

32:25

more of the hospital revenue, and

32:27

there's shortage of 15 million health

32:28

care workers globally by 2030, which is

32:31

a major shortage. And the global

32:34

surgical volume is is

32:35

to rise by 50% in the next 9 years.

32:39

So, if you put all of this together,

32:42

any technology that you can bring in to

32:44

actually make operating room more

32:45

efficient is going to be extremely

32:47

valuable both economically and also from

32:50

a patient outcome and and society point

32:52

of view. So, the real workflow is a

32:54

blueprint for hospital automation that

32:56

we have published as part of Isaac for

32:58

healthcare 0.5 and a physical demo of

33:01

that is at the Peritus booth on the

33:03

exhibit floor if you have not seen it. I

33:05

would suggest to

33:06

go and take a look at that. We start

33:08

with a digital twin of a hospital.

33:10

We train these robots to be able to do

33:13

chores and be connected to these agentic

33:15

frameworks. You can envision,

33:16

for example, a

33:18

robot that can go between the operating

33:20

room and the sterilization processing

33:22

department SPD to bring a tray back with

33:24

instruments that are needed for the next

33:26

phase of the surgery.

33:27

In order to do that, obviously, you need

33:30

a simulation environment. That's what

33:32

Isaac for healthcare provides. And this

33:35

enables you with combination of

33:37

generative AI with Cosmos transfer to

33:39

have synthetic data generation with

33:41

thousands, if not millions, of different

33:44

types of scenarios that you can

33:47

use for training of the robot. And let's

33:49

take a quick look at

33:51

this basically blueprint in action.

33:53

This is

33:56

a Peritus robot

33:58

that is trained to be connected to this

34:00

agentic framework, which, for example,

34:02

is monitoring what is happening inside

34:04

the operating room and the surgery,

34:06

knowing what phase you're in and what is

34:08

the next phase, what instruments and

34:10

accessories are required, and also

34:12

checking the room camera to see if those

34:14

instrument and accessories are on the

34:15

table and available to the surgeon and

34:18

staff. If it's not there, it will

34:20

autonomously send over a message to the

34:23

fleet of robots to basically go get it

34:25

from the SPD department. It basically

34:28

knows where to find it, uh, goes and

34:30

gets that, and brings it over to the

34:32

operating room. And it can get into more

34:35

detail type of actions of, for example,

34:37

opening the tray and recognizing what

34:40

instrument it is in there and putting it

34:42

on the sterile table for the surgeon.

34:45

Obviously, this is representative and

34:48

uh, requires much more productization

34:50

work to be available at scale, but you

34:52

can envision that what once you have

34:53

this fleet of physical agents in the in

34:55

the hospital, they can do a lot of

34:57

chores for you that are technically not

35:00

medical device functionality, but it

35:02

would help with the workflow in general.

35:05

Uh, to summarize, we uh, discussed the

35:08

anatomy of existing surgical robots, the

35:10

new way of data-centric approach. We

35:12

talked about data and next generation

35:14

models and agentic systems in hospitals.

35:16

A few points as food for thought,

35:19

we think that future of surgical

35:20

robotics will have data-centric

35:22

architecture.

35:23

Uh,

35:24

the we think that they'll use raw data

35:26

and tokens for data exchange rather than

35:28

all different weird formats of things

35:30

that uh, that exist today. Uh, and then

35:33

we think they'll be part of an agentic

35:35

framework uh, within the clinics.

35:38

In terms of regulatory considerations,

35:39

we're seeing signs that there will be

35:41

pre-cleared foundational models that

35:43

will help you with easily or uh,

35:46

basically with an easier flow to get

35:49

clearance on a fine-tuning of that

35:51

already cleared foundational model. And

35:53

then uh, that enables potential for

35:55

subsystem clearances, which is super

35:57

important. It enables new business

35:59

opportunities by bringing

36:00

interoperability, multi-vendor systems,

36:03

and continuous feature rollout. And uh,

36:06

also as starts changing the mentality of

36:09

medical device companies to become

36:10

software defined, enables SaMD or

36:12

software as medical device marketplaces,

36:15

and potentially opens the road for

36:16

general purpose compute platform where

36:19

uh, the

36:20

uh, regulation regulation uh, burden is

36:22

going to be less.

36:24

A few links that I I suggest you watch

36:26

for if you're interested in this. This

36:28

is an video Holoscan platform which all

36:30

of what we talked about is built on top

36:33

of

36:34

the NVIDIA Isaac for healthcare which a

36:35

lot of these workflows

36:37

going from simulation training to

36:39

deployment are basically being shared

36:41

and open sourced over there. And the

36:43

NVIDIA MedTech where you will find all

36:46

these foundational models and links to

36:47

the data sets for different efforts that

36:51

we're doing in collaboration with our

36:53

partners. With that, if you have more

36:54

questions, we have few minutes left

36:57

if not we can find time after or email

37:00

us. Thank you.

37:01

>> [applause]

37:06

>> Thank you Marian and Shawn. And yeah, if

37:08

you have any questions please come up to

37:09

one of the microphones that are in the

37:10

aisles. We have about 2 minutes and if

37:13

not we can also you can also speak with

37:14

them afterward out in the hall as well.

37:16

But

37:17

any questions from come to the the

37:19

microphone here.

37:24

>> Thank for the presentation. My name is

37:25

Onur from Lam Surgical. You are using

37:29

mainly

37:30

Holoscan sensor bridge for connecting

37:32

the physical world to the GPU.

37:34

And for the medical devices you're using

37:37

IGX platform for the functional safety.

37:39

Do you have plans to also embed the

37:41

functional safety in the sensor bridge

37:43

also to allow IGX platform also to be

37:46

used in the surgical platforms?

37:48

>> That's a great question. So Holoscan

37:51

sensor bridge is an IP. So we have

37:53

customers who actually have put that

37:55

into their FPGAs and have added their

37:57

own logic. Like for example, they have

37:59

added a a soft core that they basically

38:01

implement their logic inside the FPGA

38:03

alongside of that. So it's completely

38:05

configurable. For us basically we're

38:08

adding some safety features. For

38:10

example, watermarking, CRC checking and

38:12

certain things that we need to get a

38:14

seal to IEC 61508 level certification

38:17

for HSB. That's something we're looking

38:19

into. But uh broader safety solution

38:22

will come from the medical device

38:24

manufacturer because it's specific to

38:26

your own solution. But, it is

38:27

configurable and everything, including

38:29

the RTL for the IP itself, is all open

38:31

source. So,

38:33

uh you'll be able to use it towards your

38:35

uh design.

38:36

>> Thank you.

38:37

>> Thank you.

38:40

>> I think that might have been all the

38:41

time we have for questions. We're about

38:42

to run out of time now, but if you have

38:43

any other questions, you can follow up

38:44

with them outside in the hall. But,

38:47

thank you, Mati and Sean, for the talk.

38:48

Please give them another hand.

38:50

>> Thank you.

38:51

>> [applause]

38:51

>> Thank you.

Interactive Summary

This presentation explores the transition toward data-centric architectures for next-generation surgical robots, comparing them to data centers on wheels. The speakers discuss hardware advancements like NVIDIA's IGX and the Holoscan Sensor Bridge, which enable sub-millisecond latency. They also detail the use of open-source datasets and generative AI, including Vision Language Action (VLA) models and 'World Foundation Models' for surgical simulation. Finally, the session highlights the integration of these technologies into agentic systems to improve operating room efficiency and hospital automation.

Suggested questions

3 ready-made prompts