GTC SJ 2026: Physical AI for Healthcare Robotics - Simulation-First Design & Accelerated Development
1116 segments
Good morning. We are going to have two
backto-back sessions for healthcare
robotics. You have seen many great
presentations at GTC. A lot of them
about physical AI. Now we are going to
take physical AI and tell you what the
leading research and companies are doing
bringing it to the most difficult
challenge and the most difficult test of
physical AI which is healthcare. So I'll
kick off by just repeating some of the
slides that you have seen right the
opportunity is massive the need is
massive
every person in this room either is a
patient or related to a patient the
health care system has a crisis of
demand and supply and robotics in our
hypothesis can help immensely. So as you
listen to the speakers across different
use cases of uh healthcare, you're going
to see how physical AI, how these
horizontal technologies are helping
already bring in a lot of innovative
solutions to make things faster, more
efficient and more accurate.
But for physical AI, everybody knows the
main challenge is data, right? And this
slide is by design supposed to be
boring. Four boxes. In an ideal case
scenario, you're going to have the data
that is curated, annotated. You know
what it what you're looking at. It is
searchable. You find the distribution of
the data that you have. You understand
what you don't have enough of or you
don't have at all. And then you go and
bridge that data gap, those longtail
data, and use synthetic data generation
and simulation to bridge the gap.
Ultimately, you go use the data to train
a model and deploy and everybody's
happy. As simple as it looks, every
single one of these boxes needs many
many PhD students to work manually to
make it really happen for a proof of
life. The data curation annotation is
very time consuming and inconsistent
between two even residency departments.
The the outcome is totally different.
When it comes to the synthetic data
generation and simulation for many years
researchers tried to solve the
challenges of healthcare robotics, when
it comes to surgical robots, clinical
robots, there are gaps and ultimately
when it comes to training the models,
right? We are dealing with a dynamic
environment in healthcare. Whether
you're looking at the anatomy or
hospital halls, everything changes in
every second. There is no prescribed
path. So you need generalizable models.
Right? So across this presentation
from the first presenter to the last,
you're going to see the new innovations
across each of these boxes and then
you're going to hopefully connect the
dots how these technologies are going to
help get us where we want to go.
basically three fundamental shifts that
we see in other industries that we are
bringing into healthcare robotics. The
word foundation models that create
physically accurate data without solving
the physics, right? That could be our
kind of a solution to address the
fidelity of simulation, right?
Generative physics. The second bucket
are those foundation models that you
have seen many of them coming out these
days from many of our partners and
Nvidia itself models like Groot Alpamo
that reasons when you go and drive and
it tells you why I'm doing what I'm
doing which is very very important when
you are thinking about healthcare
applications right explanability and the
last part is what is being happening
today with accelerating simulation and
increasing its fid fidelity. These three
pillars and the shifts in technology is
unlocking physical AI in other domains
and we think is going to do the same for
physical AI in healthcare robotics and
our task is to bring those and make them
real. So today when you look at it you
are going to deal with a lot of models
that are going to be open sourced right
and we are on the leading uh kind of
edge for these type of models but your
development starts by taking those
models and adding your own data. If you
have real world data, great. If you
don't, you have two other path. Generate
data like models like Cosmos that we
talked about to generate data and you're
going to hear from our partners that are
doing that or simulate data, solve the
partial differential equations, the
traditional way, first principal way.
Now you have a good healthy mix of three
sources of data, post- train those open
models, deploy, test. You can test in
silicone which is going to be much
faster in simulation or inward
foundation model and the flywheel keeps
going on. It is all about the data
flywheel and you see many of the leading
uh companies on this page. The same
pipeline applies for nonclinical robots,
hospital automation. And with that,
without further ado, I want to introduce
Filipo to start from the first couple of
boxes of that boring slide and then we
go from there.
>> Beautiful.
[applause]
>> Well, thank you. Thank you Mustafa for
having me here. Uh, my name is Filipo
Fukori. I am a surgeon. I am the program
director for the robotic fellowship at
Lenoxil and Northwell Health, the
largest healthcare provider in New
Northeast New States. So what we do is
we take board certified surgeons and we
tweak their skills uh and we get them
ready to go out there and and practice
in the real world. I also lead a
laboratory that does research in um
predictive uh analytics for surgery. So
looks into the performance of a surgeon
as a whole and comes up with ways and
models in order to improve the
performance of the surgeon. Um we have
seen an incredible uh innovation in the
surgical environment over the past 30
years. You know for thousands of years
people used to do surgery uh with uh in
the open fashion and then 30 years ago
we started doing surgery with like small
canulas and laparoscopy and then the
real revolution came with the robot. Now
the robot uh has more than 30% of the
share market of a minimally invasive
surgical space. Uh and the robot allows
you to collect all of the data. We're
really only leveraging a minority of the
data that is coming out of the robot and
we're all really working hard in order
to understand the data that comes out of
the robot and to make it uh actionable
so that uh we can bring out to the field
better surgeons. Uh but there are a lot
of challenges dealing with uh healthcare
data. First and foremost, it's really
hard to get data out of hospitals with
because of uh medical legal problems.
then it's very hard to curate the data.
Just surgeons and all of those clinical
people that need to give you the
annotations are just really hard to pin
down. Uh and thanks to video language
models these days like we can embed the
video language models in our pipelines
and we can actually really improve and
cut down dramatically the num the time
that is required to annotate these data
sets. And then you know whenever
[clears throat] all data is not just the
same right like when you have a video of
a surgical procedure uh the high
impactful moments are going to be two or
3% when you're going to have a bleed
when you're going to have an inadvertent
injury of a certain structure and also
we have a huge lack of ontologies and
taxonomies in surgery. We just never
asked oursel that problem and so we're
working in order to fix that problem as
well. Um and we are also working really
hard in order to enable uh going from uh
you know real real data into simulation
data and together with uh Cosmos uh the
Cosmos product uh we are going to be
able to create better simulations that
uh really take from the real world and
allow us to uh train uh better
generation of uh surgeons and then
potentially enable autonomous sections
and so on. Because the final and the
ultimate like goal that we have here is
to for a model to have full surgical
scene understanding. So for a model to
understand exactly what is going on uh
in the screen and Nvidia is helping us
out by doing that and we're taking a
page off the playbook that they use an
autonomous vehicle. So we're using uh
the same kind of um pipeline uh and with
some of their products including Cosmos
Search and Surgical Robotics. If you're
a developer, if you're working on
creating models around surgical
performance, you can search like, you
know, a specific, you know, event or you
can search a specific procedure uh and
you can use that data in order to
leverage uh and and you know, create
better models that are uh specifically
tailored to that event as well. And as
was uh telling you, we've been doing a
lot of work around surgical ontologies.
I mean, there's no way around this. Like
in order to create surgical ontologies
for procedures, you need to get uh a lot
of surgeons under the same roof.
Surgeons that are both smart and savvy
about ontologies and the ultimate use
case for computer vision model and
kinematics and autonomous surgery. So
this is what we did a couple months ago.
We got a lot of our uh members from
sages together and we created an
ontology that looks into the surgical
gestures and surgical gestures are key
uh because basically all the surgical
gesture is what you know really powers
video language action models and so on.
Uh and you know a month later we were
able to apply uh the surgical gesture
framework uh into the open age uh data
set and this was a great honor. Um and
uh obviously you know whenever you're
annotating these surgical procedures. So
for example in the top left uh videos um
you're going to see uh two things.
You're going to see certain anatomical
structure being annotated because like
those are the places in the human body
where you have large vessels and if you
get down there you could be creating
like even some lethal injuries. Uh and
also like the target anatomy uh for
repairing a hernia for example. So you
see the ligament over there that it's
being annotated and that's kind of the
landmark for us to perform a high
quality surgery and put in a mesh so
that the patient doesn't have a
recurrence of his own hernia.
Uh and you know talking about simulation
I'm a huge fan of simulation. This is
like a state-of-the-art product. It's a
sim now too from intuitive surgical and
this is me at Italian tech week a couple
of months ago. I enjoy like talking to
audiences of engineers uh and I was
simulating uh you know in real time live
uh and uh this is a great product but
it's still like a simulation where every
pixel uh is rendered pixel by pixel and
we really do think that we can do better
than that and together with uh Sean
Hoover uh and Nigel and uh Kai we're
really trying to bring to life uh the
Cosmos um realtime simulated
environment. So this is a different
approach to simulation. Instead of
rendering everything pixel by pixel, the
idea uh is that we basically uh create a
ground truth given by the first frame uh
of the surgical operation and then we
use the kinematic data in order to
simulate the interaction with the tissue
and we've created a whole suite of
models including gshian models to
simulate the interactions between the
tissue and the tool. And uh we're
working hard to bring this uh to pro uh
to production and uh you know so that
all of our surgical trainees uh can have
a better environment where to train and
learn because like believe it or not
like simulation is is hopefully
underutilized in surgical training. So
the people that are actually operating
on uh on people they sit down at a
console with me and they do actually
learn how to operate on real humans. And
this is not fair. It's not ethical. And
that's why we're trying to really work
hard to make this better. And uh last
but not least, I just wanted to show you
uh this uh video of an autonomous robot
that was pre-trained with open age and
was trained regroup. Uh I'm not going to
talk about the technicalities that it's
going to be for axle, but in the left
upper left corner you see a surgical uh
intern. So a first year surgical reset
suturing. Suturis is actually one of the
most complex things that you can do with
a robot. And you see how much the
surgical intern is hesitating. And you
know this was amazing work uh done by
Axel and the Nvidia team uh and using
video language action models which you
know reinforces the fact that you know
gesture ontologies and gesture words are
really really uh important uh the model
and the the robot was able to throw a
stitch and and tie a knot. And with that
I wanted to thank all of my team for the
incredible work that they're doing in
advancing uh you know our surgical
knowledge. Thank you. [applause]
Uh thank you Mustafa for inviting me. Uh
really uh thrilled uh to be here today
and uh tell you about our work on
autonomous robotic surgery. Um wanted to
start with quick disclosure. Um uh in
addition to being an associate professor
at Johns Hopkins, we recently founded as
semor surgical where I'm the chief robot
officer. Um
Um so uh uh yeah uh here um uh you know
we are showing this uh rising uh case
load uh so this upcoming healthcare
crisis uh where uh we have projected uh
uh in the next uh uh 10 years a doubling
of the case load. So we as engineers
really need to provide you know
technical uh you know help to cope with
this rising case load. Um uh so uh
manual and teleyoperated robotic
assisted uh uh procedures of course
heavily depend on the experience and the
skill of the operating surgeon. Uh our
vision is to augment critical parts of
this uh manual operated uh surgery with
increasing levels of autonomy. Uh this
could reduce complications uh
democratize access to expert surgery and
help with that raising uh rising case
load.
Uh traditionally uh we have um you know
solved this uh problem uh with a
modelbased approach. Um so we have uh
you know uh published a paper in science
robotic 2022.
Um where uh we uh you know uh uh first
track tissue in 3D um then planned we
had to like develop our own custom
suturing tool uh custom uh camera
system. Um and while this is
predictable, we definitely hit a ceiling
on performance about 60% as uh you know
stitch-to-stitch success rate. Um and um
uh and it it it really doesn't scale to
new procedures. So what is super
exciting about these learning based
methods and in particular imitation
learning based methods that we can build
one framework and then add more data and
learn you know new techniques and new
skills uh you know on the fly and get
better and better. Um so we started this
uh we use uh for the hardware the Da
Vinci research kit. Uh we added a uh
wrist camera and uh started with um
tackling uh some fundamental surgical
task like lifting tissue, needle pickup
and handover and not tying. Uh high
level uh really um you know uh
relatively simple. We get uh expert uh
robot demonstrations. uh we capture the
video and the associated kinematics and
then learn uh and train uh you know uh
transformer models uh to uh you know
produce the right action giving the
right video input. And with just you
know a few hundred uh expert
demonstrations we were able to show that
we can uh you know perform some of these
you know expert uh you know demonst uh
expert um you know procedures um you
know with high fidelity. So on the left
you see not tying on uh the middle uh
you know 100% success rate uh needle
pickup and handover and then lifting
tissue the simplest task you know 100%
success rate
um but uh you know what is super nice uh
this is not sequentially programmed so
it's really robust uh to disturbances so
here we knock the suture thread out of
the grasp and it automatically repairs
itself and you know completes the task
But this is all suture pad and this is
really small little surgical task right.
So how do we get it to the next level
and perform you know procedure level you
know phases of uh you know uh robotic
surgery. Uh we've been digging into uh
the colcystectomy taking out the
gallbladder. Uh we published this in
science robotic uh last summer. We're
focusing on this middle phase of
clipping and cutting the bile duct and
cystic artery. Um uh we uh obtained
about 30 uh gallbladders uh in house and
uh repeated uh this clipping and cutting
uh you know 20 uh times on each uh
sample to get about 16 hours of uh
gallbladder uh surgery data. Um and um
then uh we trained this hierarchical
policy. Uh so now we divided the whole
face of surgery into different subtasks
um and first have an uh high level uh
language policy that looks at the
history of the video and then finds out
what is the phase of the surgery that
we're currently doing and if there are
any corrections needed and that feeds
into a language conditioned um you know
low-level policy that to then execute
and uh you know find the associated
kinematics. It's very nice to also
intervene as a surgeon because the robot
tells you what it wants to do next. Uh
so you know really nice way to intervene
and uh here uh some uh an example video
then in our study fully autonomous did
not require any intervention. You see
some you know adjustments when it
initially misses on uh the bile duct and
here the clipier comes in uh places
three clips then we change into a
scissor and autonomously uh cut the bile
duct and then the same on the right with
the cystic artery. So really nice
demonstration that this hierarchical
framework can achieve you know high um
you know fidelity autonomous soft tissue
surgery um in this uh xvivo porcen model
um and here's a compilation of you know
eight uh consecutive uh you know study
uh videos uh in our Xvivo model.
Um we since uh deployed uh this
framework the SRT framework on different
robotic systems for different
application. Uh top left you see
endoscopic uh guidance um and then uh
for endoluminal guidance uh under X-ray
uh you see uh the suture bot that we
published in Nurops for end to- end
suturing suturing. uh on the bottom uh
uh left is uh for trachea uh tumor
resection and bottom right is an example
for partial nefrectomy so cutting out a
portion of a kidney cancer. So all these
examples are in-house data set uh and
policies that we trained with in-house
you know bespoke you know policies. So
really to take this to the next level,
super excited to have collaborated with
Nvidia and so many partners to create a
much larger data set so we can you know
create much more robust uh you know uh
policies and uh uh uh uh really tackle
you know cross embodiment task and and
more uh you know uh more applications.
Um so we collected uh over 150,000 uh
you know uh trajectories with
kinematics. That's the largest um
surgical robotic data set uh published
uh put it out on on on on Monday and
already have a thousand downloads of the
you know one terabyte of data. So I
think it really creates a lot of uh
excitement in our community. Uh what can
we do with it? Uh we can train uh Groot
to uh perform suturing. So this is
trained on uh the uh full uh uh open age
data set and then fine-tune on
suturebot. And here's an example of end
toend suturing. Uh so super exciting
early results. Uh hopefully we'll follow
this up with a uh paper in the next
months uh with a more detailed study. Uh
but already uh in this initial you know
few weeks of testing uh we showed that
we can outperform the state-of-the-art
in suturing with this crude age model.
Um and what is also super exciting that
it does require a lot less uh suture
data to get proficient. So on only 33%
of the data set we already achieve uh
really high performance.
Another collaboration is with uh Nvidia
on Cosmos um where uh we uh trained a
world model trained on open age and you
can see always the video pairs. Uh the
left is the original video, the right
one is the AI generated one. uh and uh
so super exciting that we can you know
make this so realistic and that helps us
in policy testing.
Um so uh where do I see the the future
of autonomous surgery? It's going to
take a little while before procedure
level autonomy uh you know becomes uh
you know a product uh but I think these
autonomous functions in laparoscopic
surgery will come over the next few
years. Uh super exciting that Moon
Surgical already has uh the you know uh
camera control uh you know other uh
lowerhanging fruits that are a little
less risky would be suction holding and
tissue traction and port placement. Um
uh yeah with that uh let me thank uh my
uh team and all the great collaborators
and uh thank you so much for your
attention.
[applause]
All right. Hello everyone. Um, Mustafa
Jay, thanks for having me here to talk a
little bit about how we're looking to
close the data gap to enable physical AI
on Maestro. And so, um, thanks Axel also
for the shout out on our our scope pilot
feature. So, for those of you who aren't
familiar with Maestro, want to quickly
introduce it. It's a two armed cartbased
system that collects to off-the-shelf
laparoscopic instrumentation. And we
believe that this is enabling a new
category of minimally invasive surgery.
One that allows for more procedures with
fewer resources. Uh enhances procedure
consistency for increased patient safety
and and or staff quality of life. And
crucially for for this conference and
audience, we're equipping the O with a
closed loop feedback mechanism to
deliver AI based continuous
improvements. And we're available in the
US and the EU and we have a broad
indication for holding and positioning
uh laparoscopic instruments. So the
system itself at its at its heart is a
collaborative surgical robot. So the
user is able to grab or hold of his or
her tools and move them anywhere in the
workspace. Um when they let go of the
tools, they'll be locked in place. Um
and so the system sort of int infers the
intent of of the of the user. In this
particular part of the clip here, you
can see a setup joint being
repositioned. And this is key to the
sort of later part of my talk where I'm
talking about how we configure the
system at the beginning of a surgery is
is a is a a problem statement that we're
trying to solve or make more efficient.
And simulation uh can really help. We
can also port top as you see here moving
from one port to another without any
re-registration. Um, and as as Axel just
alluded to the first sort of advanced
feature we've deployed, we call it scope
pallet where here as the surgeon is
moving the instrument with manually with
their right hand, the camera is
automatically following it. So, we're
giving the surgeon the ability to
control three instruments with just two
hands.
Um, Maestro adapts to any laparoscopy.
So, I showing you here just four
different surgeries. And the uh the
reason for showing this slide is that
you'll notice that the arm
configurations in every uh surgery here
is very different. And so one of the
things that we're trying to do to make
surgery more efficient is allow them to
set up the system quicker. Um we have
done over nearly at this point nearly
3,000 procedures, but you'll appreciate
different rooms, different setups,
different beds, different surgeon
heights. Uh very much uh each case is
slightly uh is slightly different. And
you either do thousands and thousands
more procedures to optimize this or you
simulate to try and make things faster.
And that's exactly what we're doing and
I'll I'll talk about in a moment. Um
before jumping to that I just want to
touch on the fact that Maestra was
designed from the beginning as an edge
compute powerhouse to be at the bedside.
So moving around this graphic from left
to right we pull in the laparoscopic
video feed so we can see inside the
patient. We have two 3D depth cameras
doing RGB and D imaging on both on the
top of the uh what we call this the
lighthouse and also down below lower in
the system so we can see outside the
patient. And I'll be referring to this
lighthouse camera a little bit later. So
just just want to point your draw your
attention to it. The robot arms
themselves obviously are are sensors in
our case high resolution force sensors.
Um the system is connected uh outbound
only at this point but it's it's
connected to Wi-Fi or 5G and all of that
is interfaced to the IGX or uh with an
A6000 GPU and an Azure capture card. So
that's the the product as it is today.
But obviously this is a session on
simulation. So I do want to touch on how
we are using simulation to close the
data gap. Um, and our first target
application is optimized arm deployment.
And and I love the suturebot from from
Axel, but in a regulated environment,
there's a crawl, walk, run element to
this. And the first place that we're
targeting is something that is a lower
patient risk. It is set up. I fully
believe we're going to keep moving
through that the risk profile to be more
and more advanced things as as Axel
touched on. Um, but where we're starting
is in optimized deployment. And this is
not currently a feature, but I wanted to
share as to how we're thinking about the
problem and what we're doing. So if you
take the you know the three computer um
concept uh on the top left we're
simulating uh in Isaac for healthcare
with the maestro embedded within uh the
simulation environment. We're then
generating synthetic and using real
world data and training um to gen
simulate lots and lots of different
configurations. Again you know we don't
want to do we don't have to get to the
point where we've done 50,000 surgeries
to simulate 50,000 surgeries and
thankfully we now have the tools to do
this. And then lastly in the bottom we
have the deploy side where the IGX at
the bedside and this video clip just
shows how an example how the arms might
deploy. So going through this the first
up you have the simulation and what's
cool about this simulation is that it's
running our full software stack. So
that's our user interface and you can
touch on the user interface uh as if you
were a user and the system let me is
that not playing?
No, it is not. Okay, no worries. Um let
me go back up. What you would see here
is that uh as you go through the user
interface, you touch on the uh the guey,
the arms will deploy, the system will
move, and you're basically getting a
full simulation within it. But you'll
appreciate we're able to move the system
with relative to the bed, the angle, the
height of the bed, the color of the
drape, all of the different
configurations that we would see in a
clinical setting. Um, and that it's been
very powerful. It is a great debug tool
if you're if you're looking for debug
tools as well because the software stack
is integrated. um and we're using it to
render uh different configurations, but
there is a need to enrich your data set.
And so what we're doing here is we're
taking um real world video from our that
top camera I referred to and we're
generating synthetic versions of that
video using Cosmos transfer. So the
video on the top is is a real video uh
as recorded, but the lower one is a
synthetic video that's been generated uh
fine-tuned Cosmos transfer. And so
things that are different, there's no
blood on the patient's skin, the skin
has changed, the profile, the drape is
different, the light cable is a
different illumination. And so this way
we're able to generate different
clinical scenarios without having to
actually uh uh do the procedures
themselves.
We're also going the other way where
we're taking the Isaac SIM uh sorry
visualization from that top camera uh
and pro applying cosmos transfer to
generate realistic or the goal is to
generate realistic uh imagery. So we
again don't have to uh create those
images even from a real setting. But
what you'll see if you look look closely
at the the hands at that top image
there's actually Cosmos transfer put
hands on the robot. Now our robot
doesn't have hands but you know I guess
it that's where it's its background is
coming from the world foundational model
and so there is a need to fine-tune that
with real world data. So this is still a
work in progress. It's it's it's
definitely not solved yet. The the best
result we've gotten yet is in the lower
video. It has improved the drape um some
of the elements of the skin but it as I
said it still has some way to go. uh but
I do hope to be back here next year and
showing results from that. Um we then
take the training data sets the
synthetic and the real and we start to
apply training and so there's there's
two or apply them to training uh the
there's two different sort of problem
statements here. One is the on the left
you got the global environment where you
got the bed the patient the the robot
that's one set we're doing and the other
side of it is actual perspective from
the camera uh because you needed to
teach the system to be able to perceive
the scene and then decide what to do
next. And our goal here is obviously to
identify an optimum pose for a given
clinical scenario.
Um from a deploy perspective, as I said,
you know, regulated environment, we do
have to take things slowly. What we'll
start with is um a workspace
optimization button in our guey. And so
if the system is able to perceive a
scene that meets certain criteria and
does have has identified an appropriate
pose, this button will appear and you'll
press it and the system would then would
then deploy to that location. Um but
where we're going next and you know as
Axel and Filipo touched on is is
training vision language action models
with this uh with this data sets uh and
deploy them directly on Maestro. And so
we we we've we have just gotten just
recently um a trained version on group
ed 1.6 and Sean told me yesterday it's
already obsolete so we have to uh talk
to the guys about that and things are
moving pretty fast. Um but that's it's
exciting to think of what we will be
able to do and again I hope I'm back
here next year to talk about the results
on that. So, so lastly, you know, where
are we going? And you know, I listed it
as VA driven capabilities and
intentionally slightly vague there
because I think there is definitely a
hierarchy of uh tasks that we'll tackle
based off the risk to the patient and
the surgeon. But we do have this data
set of cadaavver data where we can do
very controlled experiments, synthetic
data, real world data um curating that
in the LR robot format as the open H
initiative is harmonized around. We're
training both said we we did contribute
to the open H data set as well in group
and group N1.6 six and then lastly the
the evaluation and deployment and I draw
your attention to that image on the left
of the valent deploys. I took it from a
presentation I gave a couple of weeks
ago at ORC and the at the ORC surgical
AI day and wanted to call out that in a
regulated environment we will absolutely
need to consider the safety and system
architecture to deploy these things and
the AI will sit in the middle of it but
there will be nested loops uh that
you're making sure that the system is
doing exactly what it needs to do to be
safe for the surgeon and the patient. Uh
and that's it. Thank you very much.
>> [applause]
>> Hi. Hello everybody. I'm Yossi, CEO and
founder of Lamb Surgical. Lamb is a
Swiss company. Our software director was
supposed to be here today. He hurt his
back. Couldn't sustain the flight from
Switzerland. So, uh, you're stuck with
me.
We are
we are FDA cleared and uh small
disclaimer some of the things that I
will show our future development.
This is uh say our mission reason we
wake up in the morning. Um as you can
see where we are now. So red bar is
elderly population in the US above the
age of 65. The blue bar is available
surgeons per
population. So as it appears in 2025
we're okay only small number to
remember. Today 40% 40% of US orthopedic
surgeons are above the age of 60. So
take it four or five years down the road
and then you see the perfect storm
coming few years from now. there will
not be enough surgeons in our case
orthopedic surgeons to treat the elderly
population and something need to be done
this session is about physical air some
something need to be done which is not
incremental software improvement
something need to be done in the
physical world to really accommodate for
this uh coming problem I think also
Mustafa mentioned this
and many are talking about it this is
quick uh I think all of the sessions
until now were in the soft tissue area
and really the last two decade the soft
tissue robotics were the most dominant
we are in the hard tissue robotics hard
tissue spine long bones joints and there
in the hard tissue the robotic
architecture is very different from what
you're accustomed to see in soft tissue
robotics is most of the time remote
manipulation moon is a is an interesting
I don't know to call it interesting case
but in a a a large form factor of um
surgical robotics. In soft tissue, it's
two to four robotic arms. The surgeon is
not even touching the patient sitting in
a console and it's a remote
manipulation. It's a a teleoperated
procedure. In hard tissue, it's a bit
different, bigger, stronger robotic
arms.
And the the most dominant form factor
and there's only few examples here is
usually it's one robotic arms arm coming
from the bedside merged with a what is
called a navigation camera most of the
time in infrared camera and the basic
idea is there it's concept invented 30
years ago. The idea is one bone marker
attached to the bony anatomy, one bone
marker in the end defector of the robot.
And what you see here is this uh you
don't really see it on your screen.
Never mind. You see this stationary
camera. It's usually an infrared camera.
So this camera looks at at these markers
and determine where the end effect or
the tip of the robot is in relation to
the bone. And then this robot can assist
the surgeon in hard tissue. usually is
to drill to apply saw to cut planes and
this is how it is done today. The
equivalent for this is is if I would
tell you imagine yourself trying to
hammer a nail to the wall with one hand
because this is what you have today. One
robotic arm and you're trying to hammer
a nail to the wall or to tie your
shoelaces. Never mind the analogy. One
arm one robotic arm is very limiting.
It's a very limiting architecture. That
that is why
While you can see today in intuitive and
soft tissue robotics really
um proliferated significantly in hard
tissue robotics it's a different
picture because these robots are
relatively limited can do only simple
tasks.
So what do we do in LM? We bring a first
of its kind upper torso humanoid
architecture meaning two robotic arms
synchronized with a third arm which is
the vision and then you can have the
human architecture that you are so
familiar with and use use it every day
without noticing it your uh something so
simple like to tie your shoelaces again
try to do it with one hand and this
architecture is missing in heart tissue
so two operating arms
dynamic vision that is synchronized with
the arms. Don't have enough time to
explain why is it important. Uh one of
the key elements is instrument and
implant agnostic. What does it mean? The
ability to operate any tool and not only
proprietary tool that we know um the
equivalent is you if I give you 30
different markers or pens, you can
operate all of them. No problem at all.
You take it and within a second you
calibrate it and you operate it. In
robotics,
robotics, heart tissue robotics, this
capability doesn't exist. You can
operate with the million-doll robot that
you bought only my proprietary implants
or instruments and this is a very
limiting thing and we are also doing
this where the idea is last bullet. We
want to have the pathway towards
supervised autonomous surgery. We want
to raise the bar so the robot can do
more sophisticated tasks and to automate
parts of the procedure.
Um how does it look? So let me operate
like this. [snorts]
So now you have this upper torso
humanoid architect the two operating
arms. So you have the two arms one in
this case that you see one can stabilize
a specific bone monitor it have
plurality of cameras and sensors. So
while the other arm doing what you see
here as as a bone milling the specific
bony element is controlled stabilized
and then you can have predictable
results. This is the reason why current
robotics cannot do it. Uh you need to
have predictable results and that's what
we're doing. Let me jump this forward.
Now where's the where's the challenge
and uh where we are cooperating with
Mustafa and the guys in Nvidia. One of
the things that at least occurred to me
at some point that a simple robot can do
simple tasks but also can do simple
mistakes. That's relatively easy for the
surgeon you have one arm come stand here
do this very simple. Now you have this
sophisticated robot that can do multiple
things and and the equivalent as as a
father to small girls, the equivalent
you you give your your your kids a fork
and knife and you you put it in the hand
in your in their hands. You say eat
nicely. What does it mean? So you say,
"No, no, you see, so you tell them this
is how you hold the fork. This is how
you hold the knife and eat your pasta."
Okay? So they learn to do it and but now
it's not pasta, now it's rice. So it's a
complete different action and now it's
different food. So and I'm not even
talking about more complicated thing
than a fork a fork and knife. So this
training is is a challenge and it's an
iterative
process that needs to be done and uh and
again I go back to our situation
when you have this three arms system and
you want to deploy it in the field
this system needs to be had to needs to
have a basic training of how to behave
how to move and one thing that
differentiate
hard tissue significantly from soft
issue is the the fact that it is a
collaborative robot. There is people
always there between five to five to 10
people working together with the robot.
So I will do this uh part uh fast
because I see I'm running out of time.
So uh I I will jump this forward. But
the idea is you have this complicated or
or very capable robot cooperating with
several people around the table and you
need to train it. One second. Sorry. I
will skip this quickly.
There you go. So, uh, one of the things
that at least in our list we first to
solve before we go to the clinical
application and to the bone and to the
tissue, how this robot can be can
collaborate with the people around.
There are people like the one that you
see on the right as sterile person.
Okay. So, the robot can come close even
touch it. But the the other guy is
non-sterile. You cannot come close to it
when you're sterile. There are
instruments, sharp instruments. Again,
one thing very common to a heart tissue
surgery, orthopedic surgery, hundreds of
different instrument, not your
instruments, other vendors instruments.
So, this robot needs to know how to
collaborate with all these people going
around you, the instruments that are
involved always in the surgery and then
the clinical application. So with
current
decision matrix we cannot this the
number of variable do the math is is
not scalable this solution and here AI
is something that we see with a high
value how to create what the guys can
explain or explain better than me to
create this uh and I will skip the
flowchart but to create this simulator
create this synthetic data and now train
the robot in a scale scalable way. So in
the first time that we deploy it in and
I'm finishing here. The first time that
we deploy it in the O, it has the basic
ability, the basic knowledge of how to
behave so the surgeon and the surgical
team can uh collaborate with it
successfully. Thank you very much.
[applause]
Just uh a couple of quick questions if
there is any. The amazing speakers are
here. Uh Jade is going to pass the
microphone and you guys can ask
questions. Just two short questions
please if there is any.
>> This is all
that should be on. Okay.
>> I'm Jonathan Strong. I'm actually a home
and community based provider up in
Alaska. Are there any robotic stuff
going on in home health care? Cuz that's
where I'm at. But anything that you guys
have heard of in home health care to
help like elderly and disabled people in
in their own homes? Anything like that
you've heard of?
>> Yes, there there are companies that are
working on those. Yes. Uh across the
board you can think of like from even
infection control to delivery within the
hospitals and beyond, right? And then
you can even think about exoskeleton
type of robots that are going to help
you and that is not within the
confinements of the hospital. But yes,
there are the robotic uh companies
working on that and the tools and stack
is the same, right? You still need data,
you still need to train.
>> One more question.
>> Hey there. Um I work with a lot of
startup companies and uh leveraging the
Jetson Thor platform. I'm curious on two
ways. is when I help out with smaller
companies who don't have simulation data
and are trying to leverage that to train
models uh are they able to access like
the open h or different things to have a
data set to work off of that's public
and then secondly you know when you guys
are using the Jetson for the multi-ensor
input processing is that running the
full um like learning process on top of
those as well or do you have to have two
machines communicating
>> two two machines communicating but
that's a great question because our next
session and you see Jay here uh please
be seated uh or just take a quick break
and come back. These are the same
sessions back to back. We're going to go
in depth about the runtime computer and
kind of differentiating between what you
need for the runtime versus what we
would need in the training.
That's all the questions we have uh
right now. What we're going to do is
we're going to get these guys off, get
uh next speakers miked up. everyone can
kind of stand up, stretch their legs,
and uh you know, feel free to mingle,
mix and mingle between sessions. And um
we'll see you back here in at the top of
the hour. Thank you so much.
Ask follow-up questions or revisit key timestamps.
The session introduces healthcare robotics as a critical application for physical AI, addressing a massive demand and supply crisis. The primary challenge identified is data (curation, annotation, generation), which can be overcome through foundational models, generative physics, and accelerated simulation. Speakers from Northwell Health, Johns Hopkins/Asensus Surgical, Moon Surgical, and Lamb Surgical then detail their innovative approaches. Filipo Fukori highlights the use of robotic data for surgeon training and the challenges of data access and curation, leveraging video language models and Cosmos for better simulations and surgical scene understanding. Axel Krieger from Asensus Surgical discusses moving from model-based to learning-based autonomous surgery, showcasing high-fidelity performance on fundamental tasks and procedure-level autonomy using hierarchical policies and large datasets like OpenAGE. Jay from Moon Surgical presents the Maestro system, a collaborative robot leveraging edge compute and simulation for optimized arm deployment, emphasizing safety in a regulated environment. Yossi from Lamb Surgical introduces a novel upper torso humanoid architecture for hard tissue robotics, aiming to overcome the limitations of single-arm systems and enable supervised autonomous surgery, focusing on scalable training for complex human-robot collaboration. The session concludes with a brief Q&A, touching on home healthcare robotics and data access for startups.
Videos recently processed by our community