GTC SJ 2026: Next Generation Intelligent Surgical Robots
1109 segments
Hello everyone. My name is Mati Azizzian
and I'm pleased today to talk to you
about the next generation intelligent
surgical robots.
>> [clears throat]
>> Uh here's a brief overview of the agenda
we're going to go through. Uh first
we'll take a look at the state of the
art surgical robotic systems. How do
they look like? And then we talk about
uh
the next generation of surgical robots
which we think will be data-centric uh
architectures. And then we briefly touch
on the pathway from development to
deployment and what are the computer
requirements for that
and how to get from clinical decision
making uh to AI that is used in action.
Then I'll pass it over to my colleague
Shawn Hoover who will talk about data
and next generation models and some of
the announcements that you have seen
already uh at NVIDIA. He will dive deep
into those. And then uh we'll basically
uh go over
agentic systems in uh smart hospitals
and then we'll summarize and uh wrap it
up for some uh questions and answers.
So if you look at the surgical robots of
today, uh you can categorize them from
uh a few different angles. One is by the
access method to the human body.
Um obviously open access uh robotics has
been around for quite a long time in
orthopedics, cranial, spine um
applications.
Multiport robotics in laparoscopy,
thoracoscopy, arthroscopy have been uh
around for the past three decades or so
and uh and in use. And single port
robots are a little bit uh more uh
basically
uh recent uh with the likes of da Vinci
SP and equivalents of those that are
being used. And endoluminal robots,
there's a wide range of things from
notes to endobronchial, endovascular,
and uh different types. And incisionless
robotics, uh the types of focal therapy,
um HIFU, histotripsy, and those. And
each of these basically have different
types of requirements.
Uh so mostly we're going to focus on the
left side of the
uh this screen with multiple single port
like systems.
But the phenomena that we talk about
applies to all of these basically.
You can also categorize by embodiments.
Uh so you can envision
a surgical robot could be
uh basically multiple arms on on a
single cart on wheels. This is probably
one of the most popular
types of systems that is out there. Many
of these are manufactured by different
uh
medical device companies.
Or it could be actually on multiple
carts on wheels that actually get around
the operating room table.
Uh it could be also mounted on the
operating room table
like the Ottawa system from J&J. Or
there are different types that are
basically bolted to the floor or to the
ceiling. Especially focal
therapy
robots have those type of embodiments a
lot.
Uh the other categorization you can look
into is actually from
the levels of autonomy.
This is actually from a paper in Nature
by Band Rapport and team from a couple
of years ago.
Uh where they have built on top of the
seminal paper from Guang-Zhong Yang and
team in Science Robotics in 2017 where
they tried to define and get into a
consensus on different levels of
autonomy for surgical robotics.
Uh the green curve that you see over
there shows basically how much of the
existing robotic systems are in each
category and majority as you can see are
in level one which is uh robots that
assist you mostly teleoperated robots.
And obviously
there are few that are in level two and
level three with task autonomy and
conditional autonomy. Uh but there are
very few and mostly they're in
rigid body type applications. There's
pretty much nothing commercially
available with level four or level five
autonomy. But those there's a lot of
research groups and startups working on
that and those will be coming in future
as well. And some of the architectures
we are talking about is actually
enabling not only the level one systems,
but also enabling the higher level
autonomy systems.
So, a typical surgical robotic system,
a tele-operated one,
has one or two surgeon consoles
that the primary and secondary surgeon
are sitting at there. That could that
could be a trainee and a mentor or
basically two surgeons were sharing
responsibilities or just a fellow who's
actually watching the surgery.
There's usually a vision tower that's
endoscope and all accessories are
connected to and a surgical robot. So,
this is a very simple architecture of
what exists today. Often the modern
surgical robotic systems have a
connection to the cloud, a virtual
private cloud that is used for
over-the-air updates,
maintenance, collecting data, and
sometimes basically
also
kind of remotely connecting to the
system and doing things like for example
tele-mentoring.
And we see more and more of edge AI
sidecars being added to these systems as
an afterthought to bring
AI capabilities to the edge of surgical
robotics.
Uh
If you look at
a little bit more details of how
different components of such a system
are connected to, this is just a
representation and an oversimplified
way of talking about how different
components of a surgical robot are
connected to.
There are many different components.
This is by definition a distributed
system.
Uh
And each of those are connected with
often proprietary
type of connections to each other. So,
the surgeon console
to the vision card, to the
surgical robot itself, and the
connection between different parts of
that. Sometimes there are star
configurations or daisy chains with
different types of uh,
basically cables, signals, protocols.
And this makes it extremely difficult to
actually move from one generation of a,
robotic system to the next generation
because you basically have to maintain
something that A to Z is proprietary, is
not interoperable with other things, and
uh,
often you deal with obsolescence of
basically parts that are supporting
that.
So,
uh, that's one of the reasons we're
actually thinking that
>> [clears throat]
>> the next generation of surgical robots,
like many other robots, are going to be
data centers on wheels or data centers
bolted to the operating room depending
on what embodiments that you're talking
about. Data centers are very extensible.
Uh, let's take a look at few parallels
from different industries. If you look
at the autonomous vehicle industry,
especially in the Bay Area, you see a
lot of startups who are working on
uh, AV. Some of them now are big
companies that have products in the
market. You probably have
uh,
took a ride with Waymo here.
Uh, that you have seen how it works, but
often,
uh, if you have the inner guts of a
prototype, it's actually that ugly, you
know, in the trunk. This is a trunk of
a, uh, self-driving car under
development. You see very different
types of cable, cables, different
protocols, and uh, different
connections. Extremely difficult to
maintain and deal with multiple
different vendors, uh, and uh,
complexities that comes with that. And
here is actually a view of a
data-centric, uh, autonomous vehicle.
So, you can see actually that the ECUs,
or the brain of the the car itself, is
basically connected over Ethernet to
pretty much all the sensors and
actuators in the car. And that enables
you to easily expand if you want to go
from one generation to the next
generation, add more sensors, upgrade
your sensors, change your actuators,
increase the capacity of compute. This
is simply a data center. You just change
the compute, you You the sensors and
actuators, and everything stays the
same. It's all ethernet based.
It's interoperable.
Makes you capable of actually working
with different vendors and even have a
basically a system that is built by
multiple different companies. Same goes
with humanoids. If you look at legacy
humanoids, it's actually
kind of funny to call humanoids legacy.
They're They're newer by definition, but
what is in the market has been following
what has existed in
kind of robotic manipulators and also
AV.
So often there's fixed input and output
that limits basically the sensor count
that you have. You can't easily add more
sensors.
You have separate buses for sensing and
actuation and then you basically end up
with something complex and
expensive to deal with. You have limited
safety and fault tolerance and there's
heavy usage of CPU for handling the
traffic versus a data-centric humanoid
design that we're
basically working on with some of our
partners
is all ethernet based. Again, you have a
brain of the
of the humanoid, that green box in the
chest of the humanoid, that has flexible
IO over ethernet. So you're only limited
by the bandwidth of the ethernet links
that you put in there. You have a
unified network for sensing and
actuation. Everything is synchronized on
the same network.
This
drastically reduces the complexity and
makes it more fault tolerant and also
because of the network acceleration that
basically
brains of systems bring in that we'll
talk about, it offloads the CPU usage.
So you basically have direct access in
and out of GPU memory.
So
here's actually
a view of that, a a very simplified view
of what a data-centric surgical robot
would look like. You can actually see in
the bottom
that you can have different edge
computes for different zones of a
surgical robot for lack of a better
word. You have a simulation zone,
uh surgical console zone, a robotic uh
zone, endoscopy zone, and accessory
zone.
You have a backbone of uh basically an
Ethernet
uh
uh network, and you can have a redundant
backbone as well for safety purposes.
And pretty much all the input-output in
the system is dealt with uh Ethernet. Uh
Host and sensor bridge that we briefly
touch upon
after this enables you to pretty much
convert any type of signal in uh and out
of uh Ethernet and basically make that
RDMA-able. RDMA stands for remote direct
memory access to uh GPU memories. So,
we'll uh double down on that a little
bit.
>> [clears throat]
>> So, let's get a bit geeky here for those
of you from Medtronic. If you want to
look at the highest level of
requirements for a data-centric robotic
system,
uh you want the robotic network uh to be
end-to-end Ethernet-based.
And you want to uh have the data
accessible in real time anywhere within
the system, so you don't have to
redesign things to get, for example, the
endoscopic image on the other node or
this joint encoders basically in another
node. Everything should be accessible
everywhere with software configuration.
Ingress data from sensor should be
accessible in any processing unit with
minimal latency and jitter. That's uh
basically a requirement,
>> [clears throat]
>> and vice versa. Egress data from a
processing unit should be accessible to
controller units, whether it's a motor
controller or display controller, again
with minimal latency and jitter.
There shall be redundant access passes,
obviously, for safety-critical data, and
all of ingress-egress data should be
synchronized, obviously.
You want to have data authenticity and
security. Uh that's an obvious
requirement because you don't want
someone to connect with their laptop to
the system and and try to uh hack into
that.
And switching uh this one is probably uh
one of the most important features that
this enables, that switching between a
digital and physical twin of a robot
becomes seamless with this type of
architecture. With the change of IP
address and port, you'll be able to
actually switch from a digital version
of a sensor to its physical version. So,
development cycles become much faster
because of that. Uh you don't need to
create
uh hardware abstraction layers and uh
basically make ad hoc simulation
bridges.
The processing units obviously shall be
able to record data and playback that
data seamlessly with synchronization and
consistency. So,
let's take a closer look into
performance. Uh I'll just give you a
very simplified example of that. Imagine
that
you want to talk about uh
event-to-action performance on a
surgical robot, which is the core of
what it does. You know, a teleoperated
robot is detecting something, you take
an action based on that and and go
through. [snorts] So, if you have a
chip-on-tip sensor, imagine an
endoscope, in a regular uh
surgical [clears throat] robotic system,
you have it go through an ISP, an image
signal processing pipeline, going
through a video router,
uh to a frame grabber often on a
different physical node, to a processing
unit doing some processing on that, and
then to a controller unit, and then back
to a motor controller, to a robotic arm
to do something. And I I'm not showing
surgeon in loop here just to be able to
talk solid about the
uh performance that you get. The best
things that you find in the market for
this is like 50 ms uh latency
from basically
>> [clears throat]
>> the start to the end. Um
uh in terms of what's uh performance
that you get. The data-centric version
of this will be something like this,
that you have uh the chip-on-tip sensor,
the raw data over, say, CSI MIP, goes
through whole scan sensor bridge
directly written to GPU memory. So, as
the lines are being uh scanned in your
camera, they're basically written in GPU
memory. In another word, the GPU memory
becomes extension of your sensor memory.
And then it goes to a whole scan SDK
pipeline on the GPU, which is basically
a graph for processing, and then you can
basically
>> [clears throat]
>> go back out over Ethernet to a
microcontroller that is controlling your
robotic arm.
And benchmarks on this that we have done
shows a speed of light implementation of
1.2 milliseconds from the input to the
output. So, you see the difference that
you get uh going from 50 to uh 1.2 uh
milliseconds uh on this.
>> [clears throat]
>> How is this enabled?
Uh we have a smaller version of of uh
our data centers on the edge for these
type of robotics deployments. This is
>> [clears throat]
>> the high-end version of that called IGX
Tour T7000. There are different versions
of it with uh basically different levels
of compute and and price points,
obviously.
Uh this one brings over 5000 uh
uh teraflops of FE4 uh for inference,
and you have the option of adding an RTX
Pro 6000 if you need more uh compute.
Uh you get 128 uh GB of memory, and more
importantly, you basically get up to 400
Gbit/s of Ethernet bandwidth to this,
which is twice of the previous
generation, which was IGX Orin
uh 700. And then next generation, it'll
be doubled again, it'll go to 800
Gbit/s. So, without doing anything, this
data-centric uh approach will give you
future proofness.
Uh
the IO parts is enabled by Holoscan
Sensor Bridge. This is a very simplified
uh reference design of that.
Uh it is an FPGA IP that we have made
public and freely available for everyone
who wants to put it in their FPGA. It's
pretty much supported uh by any FPGA
vendor. Uh there are many of our
partners who are building ASICs out of
that, and
also it comes with a software emulator
that you can run it on a microcontroller
or on any CPU, basically.
Uh so, you can convert pretty much any
type of uh signal that you get from I2C,
SPI, and GPIO to just the maybe CSI,
LVDS and so. And on the other side you
will have ethernet in and out uh time
synchronized if you have multiple of
these and directly accessing uh GPU
memory. This is enabled by NVIDIA
networking acceleration
uh provided with with ConnectX.
So, this is the
1.2 [clears throat] millisecond pipeline
that I was talking about.
If you look at uh the top left you see
the stereo camera connected to a whole
scan sensor bridge implemented on a this
is a Lattice ECP3.
It goes over an ethernet link to an IGX
and gets processed in a whole scan
pipeline. Uh what you see on uh the
right side of the screen, the green is a
laser pointer controlled by a uh by a
human and then the red is basically what
that galvo um
uh laser pointer with at a two degree of
freedom robot is uh is uh basically
shining on the wall to track the
Let me try to play it again if
to track basically the green. And you
can see that's almost perfectly tracking
that. The little offset is uh because of
calibration.
Uh so, that 1.2 millisecond is getting
into uh the basically from shutter close
to
uh the GPIO command sending to uh the
MCU. And the photon to photon here is 10
milliseconds which is mainly driven by
the time that the camera shutter
basically needs to be uh to be open.
Uh
if you look at this obviously in a
regular teleoperated robot this can be
used but also
you hear a lot about a lot about tele
surgery because um robots and surgeons
are both scarce and you want and
surgeons are more scarce and you want to
have the surgeons be able to operate in
remote locations. So, this is a a very
simplified uh kind of diagram of showing
how state of the art tele surgery works.
Uh basically on the patient side you
have the endoscope being processed and
compressed, sent over the network,
displayed to the surgeon, surgeon taking
an action, that action goes again over
the network back to the other side, and
the robot follows that.
How this can be accelerated using this
type of data-centric technology is that
you can actually apply Holoscan sensor
bridge and the Nvidia acceleration to
basically shove off latency on both the
patient side and the surgeon side,
because there's not much that we can do
about network. You can
obviously work with the likes of
Intuitive and other telesurgery
companies and telecom companies to get
the best network possible,
but speed of light is limited based on
the distance that you have. What we can
do is shove off latency on basically the
proximal and distal side of the
pipeline, and give a better performance
and better safety to the surgeon who's
operating.
Let's quickly touch on the development
workflow and how
we're actually thinking about that. We
see this type of problem, especially as
you go to the higher levels of autonomy,
as the three computer solution problem.
What it means is basically
if you're bringing AI to the loop, you
need to train that AI algorithm. Often,
you need to actually involve some sort
of a simulation,
and then at the end of the day, you need
to deploy it. But this is actually a
cycle that keeps going on during your
product development and after the
product launch as well. So you may want
to make sure that this cycle is very
fluid, and you have a flywheel running
over there. This data-centric approach
enables you to seamlessly, as I
mentioned, switch between simulation and
real world. So
regardless of if you do it on prem or on
cloud, you will be able to use the same
brain of the robot, say based on that
IGX,
working with a simulation environment
without knowing
that is simulated and not real. Shawn is
going to talk about the basically a
modern approach to surgical simulation
next after this.
And
you can actually envision that once you
have this uh ready for deployment, you
can implement it with a distributed
compute for layered intelligence having
the embedded intelligence in the robots
whether they're surgical robots or
cobots
in the operating rooms also having the
signals from the room like room cameras
on prime compute in the hospitals and
cloud compute and this will enable you
to actually scale the intelligence as
needed for your systems.
Clinical decision support has been the
main focus for many years with AI
development for surgery.
The likes of Cosmo AMD genius platform
or Medtronic digital surgery one have
been adding edge compute as a side car
to medical devices and then displaying
the results of for example detecting a
polyp in colonoscopy to the surgeon to
help them with the clinical decision.
But you can actually envision that AI
can be used in action closing the loop.
This is Moon Surgical scope pilot where
tracking of these instruments results in
detecting of where the surgeon wants to
look at and then autonomously moving the
endoscope to basically help them locate
that. With that I'm going to pass it to
my colleague Shawn to talk about the
data and the next generation models.
>> [applause]
>> Thank you Marty.
Yeah, so if you attended Kimberly's
keynote on Monday you've probably seen
kind of the high level of this but we'll
be talking about the data that we
released this week and how how that's
helping to fuel the next generation of
AI [snorts] models.
So if if you look at data available to
train different types of generative
models if you scrape the entire internet
there's roughly a billion hours of kind
of human existence for for training data
across math, across science, philosophy,
medicine etc.
Now if you look at what's available for
general robotics then that drops down to
on the order of tens of thousands of
hours. So, this is why Frontier Labs in
the general robotics space
pay people to come in off the street and
teleoperate robots for training data
such as folding t-shirts, doing the
dishes, etc.
Now, of course, when you look at what's
openly available for healthcare
robotics, you know, this is especially
dire. You know, if you're researcher, if
you're a PhD student, and you happen to
be at a university that does not have a
DVRK, you're kind of out of luck.
There's only 3 and 1/2 hours of open
source data. This is a 10-year-old data
set. That's it's still an amazing data
set put out by Johns Hopkins called
called the JIGSAWS data set.
So, when we were looking at, you know,
what we could do roughly a year ago in
the space, then, you know, we we decided
to reach out to some of the luminary
figures
in healthcare robotics,
folks such as Axel Krieger at JHU,
Nasir Navab at TUM, and we decided to
build a coalition that we call
Open Embodyment, this community-driven
effort. And I think kind of beyond our
wildest dreams, we grew from this
initial group of of three organizations
to 35 partners spanning industry,
academia, as well as healthcare
partners.
Filippo Filicori, for instance, here in
in the audience and his team of surgical
residents helping to do annotations. So,
really this great group came together.
Actually became, I think, an easy sell
to get people involved.
And so, the data set itself has eight
different robotic embodiments
represented, all sorts of different data
spanning from simulation to tabletop
to ex-vivo animal to even clinical. So,
CMR Surgical made an incredibly generous
donation of almost 500 hours of clinical
data. So,
you know, robotics kinematics paired
with with video.
So, this data set went live on hugging
face. I just checked right right before
this this morning. Already over a
thousand downloads of 4 terabyte data
set in just a few days.
So now that we have this great data set
to work with, you know, what are some of
the things that we can do? One of the
first models that we experimented with
building is a vision language action
model of VLA. So Nvidia has
series of models called Groot for
general humanoid robotics. So we took
building on that foundation. We then
trained on all of the open age data. To
see
what it could do in the lab
with our with our partners such as Johns
Hopkins. So this is, you know, being
2026 and AI moving as quickly as it is.
I'm going to refer to this type of VLA
architecture as a classical VLA meaning
that it's built upon a vision language
model at its base of VLM if you look in
the bottom left hand corner.
And
this is this is what we saw when we
pre-trained on all of the open age data
and then did a little bit of post
training on suture data from an effort
from JHU called suture bot.
So hopefully this will will play.
There we go.
So this is
running on a DVRK
example and this is
an end to end suturing.
So the model itself probably won't do
great at zero shot straight out of the
box, but it's been pre-trained on all of
these different embodiments on all of
these different tasks.
That it requires less data than you
would need otherwise.
So you know, this is a good example. The
pre-trained model is then post-trained
on just a little bit of the suture
suturing data. And what you see here is
a successful example. And of course full
disclosure, there were many unsuccessful
examples. We're quantifying the the
entire performance right now as part of
an upcoming paper.
So before I I I referred to this VLA as
a classical VLA because this is it's
it's almost out of date. We we we just
published this model and here's an
example of what we're looking at in the
future. Now VLAs rather than having the
vision language model as their backbone
have world foundation models as their
backbone. So if you're not familiar with
world foundation model is a generative
AI model that can generate video going
into the future. So not only is this
model from Semaphor Surgical which is a
company founded by Axel Krieger and some
other JHU
collaborators.
So not only is it generating the
kinematics of what the robot should do
in the future. It's also generating the
video. And some early papers are showing
that these types of models, this type of
architecture
is actually doing better than what we're
seeing with the classical VLA.
So now I'm going to switch gears a
little bit and
talk about rather than a VLA what world
foundation models can do as as far as
simulating surgical robots.
So this is our second model that we also
open source this week and it's called
Cosmos H Surgical Simulator.
So when world foundation models first
appeared they were typically text
driven. You would write a description of
what you want. You may have seen this
from Open AI's Sora and it would
generate a video describe
a video based on your description.
So this is different in that it's driven
by the kinematics of the robot. So you
give it a ground truth single frame
image and then the kinematics of of your
robot and it actually generates each
successive frame.
So this model has been trained on all of
the data from from Open H. So it's a
single model that can do eight different
embodiment simulations across the
different tasks that it was trained on.
What you're looking at is suturing and
one thing I want you to take notice of
here is again, we're just giving it the
first frame and it can actually keep
temporal coherence with how the ground
truth evolved over a few seconds and
that'll be important for what I talk
about a little bit later.
So again, here's kind of an example of
all different types of scenarios. These
are all holdout test sets of of what
this model can can simulate.
Um where we're taking this type of model
next is these models are getting closer
and closer to being real-time.
Currently,
you might need to
spend 10 minutes to generate 30-second
clip. Uh but we can distill these
models, make them smaller, and make them
much faster. So our fastest version that
runs on a single GPU today runs at about
12 frames per second on an RTX Pro 6000.
Uh but if you've been out on the the
show floor, there's a really cool demo
for a driving simulator
in in set of Cosmos. And so if you've
seen that,
that's actually running at almost 30
frames per second. Although full
disclosure, that's running on 12 Vera
Rubin GPUs. So it's a lot of expensive
firepower to run that. But of course,
this is this entire field is
accelerating so quickly that we're
reasonably confident that we'll be able
to have a real-time version within the
next 6 months. And once we have a
real-time version, you can actually
teleoperate these. So just like that
driving simulator has a gas pedal and a
steering wheel where you can interact
with the world foundation model, we'll
be able to connect happily devices for
instance and treat this as a surgical
simulator.
So
um
and uh
switching gears again, our third type of
model and this was just actually
released this morning uh is the
tokenizer for real-time surgery. Uh so
as tele surgery so as many of you are
familiar, in telesurgery,
uh the the the the the surgeon
uh controlling the robot
uh
you're use a classical encoder to encode
the video to try and make the the video
packets as small as possible,
uh reduce latency.
So, um just like uh tokenizers are used
for large language models, when you type
a request, uh a text tokenizer turns
your request into tokens where it lives
in a latent space that AI model
understands.
Uh we now have the same thing for for
video uh for for for telesurgery. So,
this has been trained on surgical video
data,
and these models are are are getting
really great at high compression and
high quality. Uh this is changing very
very quickly.
Uh so, this is an example of the model
that we uh open sourced and put in on
Hugging Face. Uh it was made public this
morning. We're really uh interested to
see what types of experiments people do
with this. And this is actually rivaling
JPEG compression in terms of quality
here.
So, if we put some of these technologies
that I just talked about together, I'm
going to describe real quickly kind of
uh this uh idealistic moonshot
experiment that we're hoping to do.
Uh so, having these neural tokenizers
means that we can put not just video
data, but in the future
uh all all types of multimodal data,
whe- whether it's audio, uh whether it's
the kinematics that all live in the same
latent space that generative AI models
can understand.
And so, if we feed those tokens into the
world foundation model that I described
earlier, as that gets to be real time,
and that temporal coherence is good
enough uh that we saw earlier, we might
be able to, you know,
predict what happens a few frames into
the second. And so, this this moonshot
experiment that we're hoping to do later
this year is to see, can we shave
latency off of telesurgery using these
technologies together?
So, again, it's it's it's very
aspirational, uh but this is kind of
where we're headed in in some of our
ideas.
All right, and I'll switch it give give
it back over to Madi for the next
section.
>> Thank you, Sean.
>> [applause]
>> Thank you, Sean. So, I'm going to spend
few minutes to talk about the agentic
systems in smart hospitals, how to
connect all of this intelligence that
the your system would architecturally
support and the models that Sean talked
about would actually bring it on the
edge. In 2024 at GTC, just 2 years ago,
we actually showed surgical agentic
co-pilot here.
It was mainly think of it as a chatbot
for the surgeon and the operating room
staff,
bringing all of the EHR and PACS data,
the surgical endoscopy
feed live basically into a VLM and also
loading a rag of the medical device
documentation, and allowing basically
the model to understand what is
happening in the surgery and providing
answers to any questions that people
have. For example, what phase of surgery
you're in, what instruments you're going
to need in the next phase,
what complications the patient might
face.
This year actually we have expanded and
built on top of that work and brought
physical agents to basically this
framework. If you look at the problems
in the that we have in the hospitals, a
lot of that is around workflow
efficiency.
Just a very brief overview,
every minute in operating room in US
costs around $46, and there are over
50,000 operating rooms. Just do the
math. Surgical suites represent 60% or
more of the hospital revenue, and
there's shortage of 15 million health
care workers globally by 2030, which is
a major shortage. And the global
surgical volume is is
to rise by 50% in the next 9 years.
So, if you put all of this together,
any technology that you can bring in to
actually make operating room more
efficient is going to be extremely
valuable both economically and also from
a patient outcome and and society point
of view. So, the real workflow is a
blueprint for hospital automation that
we have published as part of Isaac for
healthcare 0.5 and a physical demo of
that is at the Peritus booth on the
exhibit floor if you have not seen it. I
would suggest to
go and take a look at that. We start
with a digital twin of a hospital.
We train these robots to be able to do
chores and be connected to these agentic
frameworks. You can envision,
for example, a
robot that can go between the operating
room and the sterilization processing
department SPD to bring a tray back with
instruments that are needed for the next
phase of the surgery.
In order to do that, obviously, you need
a simulation environment. That's what
Isaac for healthcare provides. And this
enables you with combination of
generative AI with Cosmos transfer to
have synthetic data generation with
thousands, if not millions, of different
types of scenarios that you can
use for training of the robot. And let's
take a quick look at
this basically blueprint in action.
This is
a Peritus robot
that is trained to be connected to this
agentic framework, which, for example,
is monitoring what is happening inside
the operating room and the surgery,
knowing what phase you're in and what is
the next phase, what instruments and
accessories are required, and also
checking the room camera to see if those
instrument and accessories are on the
table and available to the surgeon and
staff. If it's not there, it will
autonomously send over a message to the
fleet of robots to basically go get it
from the SPD department. It basically
knows where to find it, uh, goes and
gets that, and brings it over to the
operating room. And it can get into more
detail type of actions of, for example,
opening the tray and recognizing what
instrument it is in there and putting it
on the sterile table for the surgeon.
Obviously, this is representative and
uh, requires much more productization
work to be available at scale, but you
can envision that what once you have
this fleet of physical agents in the in
the hospital, they can do a lot of
chores for you that are technically not
medical device functionality, but it
would help with the workflow in general.
Uh, to summarize, we uh, discussed the
anatomy of existing surgical robots, the
new way of data-centric approach. We
talked about data and next generation
models and agentic systems in hospitals.
A few points as food for thought,
we think that future of surgical
robotics will have data-centric
architecture.
Uh,
the we think that they'll use raw data
and tokens for data exchange rather than
all different weird formats of things
that uh, that exist today. Uh, and then
we think they'll be part of an agentic
framework uh, within the clinics.
In terms of regulatory considerations,
we're seeing signs that there will be
pre-cleared foundational models that
will help you with easily or uh,
basically with an easier flow to get
clearance on a fine-tuning of that
already cleared foundational model. And
then uh, that enables potential for
subsystem clearances, which is super
important. It enables new business
opportunities by bringing
interoperability, multi-vendor systems,
and continuous feature rollout. And uh,
also as starts changing the mentality of
medical device companies to become
software defined, enables SaMD or
software as medical device marketplaces,
and potentially opens the road for
general purpose compute platform where
uh, the
uh, regulation regulation uh, burden is
going to be less.
A few links that I I suggest you watch
for if you're interested in this. This
is an video Holoscan platform which all
of what we talked about is built on top
of
the NVIDIA Isaac for healthcare which a
lot of these workflows
going from simulation training to
deployment are basically being shared
and open sourced over there. And the
NVIDIA MedTech where you will find all
these foundational models and links to
the data sets for different efforts that
we're doing in collaboration with our
partners. With that, if you have more
questions, we have few minutes left
if not we can find time after or email
us. Thank you.
>> [applause]
>> Thank you Marian and Shawn. And yeah, if
you have any questions please come up to
one of the microphones that are in the
aisles. We have about 2 minutes and if
not we can also you can also speak with
them afterward out in the hall as well.
But
any questions from come to the the
microphone here.
>> Thank for the presentation. My name is
Onur from Lam Surgical. You are using
mainly
Holoscan sensor bridge for connecting
the physical world to the GPU.
And for the medical devices you're using
IGX platform for the functional safety.
Do you have plans to also embed the
functional safety in the sensor bridge
also to allow IGX platform also to be
used in the surgical platforms?
>> That's a great question. So Holoscan
sensor bridge is an IP. So we have
customers who actually have put that
into their FPGAs and have added their
own logic. Like for example, they have
added a a soft core that they basically
implement their logic inside the FPGA
alongside of that. So it's completely
configurable. For us basically we're
adding some safety features. For
example, watermarking, CRC checking and
certain things that we need to get a
seal to IEC 61508 level certification
for HSB. That's something we're looking
into. But uh broader safety solution
will come from the medical device
manufacturer because it's specific to
your own solution. But, it is
configurable and everything, including
the RTL for the IP itself, is all open
source. So,
uh you'll be able to use it towards your
uh design.
>> Thank you.
>> Thank you.
>> I think that might have been all the
time we have for questions. We're about
to run out of time now, but if you have
any other questions, you can follow up
with them outside in the hall. But,
thank you, Mati and Sean, for the talk.
Please give them another hand.
>> Thank you.
>> [applause]
>> Thank you.
Ask follow-up questions or revisit key timestamps.
This presentation explores the transition toward data-centric architectures for next-generation surgical robots, comparing them to data centers on wheels. The speakers discuss hardware advancements like NVIDIA's IGX and the Holoscan Sensor Bridge, which enable sub-millisecond latency. They also detail the use of open-source datasets and generative AI, including Vision Language Action (VLA) models and 'World Foundation Models' for surgical simulation. Finally, the session highlights the integration of these technologies into agentic systems to improve operating room efficiency and hospital automation.
Videos recently processed by our community