HomeVideos

Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work

Now Playing

Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work

Transcript

1348 segments

0:07

Good afternoon, everyone.

0:09

It's great to be here.

0:11

Uh we talk a lot about AI

0:14

that lives on your screen,

0:17

uh lives in the digital world.

0:19

And today, I'd like to talk to you about

0:21

a different kind of AI that we've been

0:23

building at Waymo. AI that lives in the

0:27

real physical world.

0:31

How many of you, by the way,

0:33

have been in a Waymo? Just raise your

0:35

arms.

0:36

Wow, okay. That is impressive.

0:38

Especially, I understand many of you are

0:40

out of town. Uh the folks who are

0:43

visiting and have not had a chance to

0:45

check out Waymo, I hope you while you're

0:47

here in the Bay Area, give it a try.

0:49

Uh so, this being

0:52

a startup school,

0:53

I structured this presentation as a

0:55

sequence of lessons.

0:57

Seven lessons that we've learned over

0:59

the years at Waymo

1:01

around what it takes to build and safely

1:06

ship today's most mature application of

1:09

AI in the physical world,

1:11

the Waymo driver.

1:14

Uh let me start with

1:17

a short video.

1:19

Uh this is a clip from a ride that I

1:22

recently took in a Waymo

1:24

with my kids.

1:26

Uh so, as you see here, you know, we're

1:27

moving uh forward or proceeding through

1:30

an intersection, and a couple of human

1:32

drivers just decide to cut in right in

1:34

front of us.

1:36

And the Waymo driver

1:37

reacted safely, reacted smoothly. In

1:40

fact, so much so that the kids, my kids,

1:42

were preoccupied in the backseat, they

1:43

didn't even notice that anything

1:45

happened.

1:46

And to me, this was a

1:47

pretty powerful moment. I've been

1:49

working on this technology and this

1:51

product for close to two decades,

1:53

and, you know, it just did something

1:56

fairly important. It

1:58

acted safely. It kept my kids safe. It

2:00

kept everybody safe.

2:02

And nobody noticed.

2:04

And that I think will be a bit of a

2:05

theme in general when it comes to

2:09

physical AI.

2:11

That the best AI moments

2:13

will look like nothing happened. It's

2:14

just the task got done safely and

2:18

smoothly.

2:21

And these sort of moments where the

2:23

Waymo driver

2:25

kept everyone safe are happening daily

2:28

across our fleet.

2:30

Today, the Waymo driver is serving

2:32

around 500 trips per week and driving

2:36

over 4 million fully autonomous miles

2:38

every week in 15 cities across the

2:40

United States.

2:42

Just for our comparison, that's over 300

2:45

years every week of an average American

2:48

driver per year.

2:51

And the Waymo driver is accomplishing

2:53

that with a superhuman safety record.

2:57

So, what what does it take to build and

3:00

deploy an AI agent in the physical world

3:03

at scale?

3:07

Now, in Silicon Valley,

3:09

there's a common mantra

3:11

to move fast and break things.

3:15

However, when you're dealing with atoms

3:16

instead of bits,

3:18

breaking things is not really okay.

3:21

So, the thing you have to do

3:23

is to move fast

3:26

and ship safely.

3:29

And that's a much more difficult thing

3:30

to do.

3:31

You have to build systems that are

3:33

robust from day one.

3:35

You have to build AI models and you have

3:37

to build training recipes where safety

3:39

is the foundation and not an

3:41

afterthought, not an add-on.

3:45

And by the way, the problem itself of

3:46

physical AI is different from digital

3:48

AI.

3:50

There

3:51

four main gaps that you have to contend

3:55

with if you're building AI for the

3:56

physical world versus the digital world.

3:59

First,

4:00

there is the cost of air gaps.

4:03

I have a language model or a chatbot or

4:05

a co-pilot and it makes a mistake, you

4:08

know, usually it costs you a retry.

4:11

In the physical world,

4:13

the cost of a mistake can be measured in

4:15

human lives,

4:16

not tokens. There's simply not an undo

4:20

and a retry button.

4:23

Secondly, you have the latency gap. And

4:25

typically, where you're running a VLM

4:28

uh or you know, uh a digital assistant,

4:31

it can take many seconds, sometimes

4:34

minutes to come back with an answer to

4:36

you.

4:38

A car traveling at freeway speeds moves

4:41

about 100 ft in 1 second. So, there

4:44

milliseconds really matter. And you have

4:46

to run all of your inference, make all

4:50

of your decisions on board a computer

4:51

that fits in a trunk of your car.

4:55

Next, there's the data gap. Uh digital

4:58

AI had the internet.

5:00

A this wonderful

5:03

immense cache of pre-labeled human

5:07

knowledge and human thought that we've

5:08

ever assembled.

5:10

There's no digitized version of the

5:13

internet for the physical world.

5:17

And lastly, there's the validation gap.

5:19

In digital AI,

5:21

often you can ship something that's you

5:24

know, good enough.

5:25

And then you let your users

5:27

you uh use your product, they find the

5:30

edge cases, and that allows you to

5:32

deploy on day one practically at

5:35

unlimited scale. And then you can just

5:37

iterate and hill climb on quality from

5:38

there.

5:40

In physical AI, the situation is

5:42

different.

5:43

Given the high cost of errors,

5:46

you need to have a very high level of

5:49

safety and a very high level of

5:51

confidence on day one before you deploy

5:54

your first robot, before you drive your

5:56

your first autonomous mile.

5:59

Now, at the same time,

6:01

when you're dealing with physical AI,

6:04

uh the actual experience of having your

6:07

agent in the real world is invaluable

6:10

and it's irreplaceable.

6:12

Uh these systems are not just something

6:14

that you can build in the lab, you know,

6:16

get it

6:17

perfect, and then deploy at full scale

6:20

overnight.

6:21

So, given those two factors, you really

6:24

need

6:25

to super clearly and super crisply

6:27

define the operating conditions and the

6:30

deployment parameters of your agent, and

6:33

then build a rigorous

6:34

framework to guide your deployment so

6:37

that you can scale in a responsible

6:38

manner.

6:39

And this is

6:42

absolutely critical. Uh this is how you

6:44

earn trust from your customers, from the

6:48

communities, from the regulators, and

6:49

yourself.

6:52

Uh so, at Waymo, we see these gaps, of

6:54

course, in the context of autonomous

6:55

vehicles, uh but these gaps uh will show

6:58

up in practically

7:01

any sort of non-trivial physical agent

7:02

that we will deploy uh in some shape or

7:04

form.

7:06

And driving is simply the first domain

7:08

where AI has crossed these four gaps at

7:11

scale with the public interacting with

7:13

our product.

7:15

So, let's dive into those lessons that

7:16

we've learned over the years at Waymo

7:18

from uh working on this problem, and uh

7:21

talk about how we address those gaps. Uh

7:23

I have seven lessons in this talk. Um I

7:26

They're all technical. There's a lot

7:28

more that goes into building a company

7:30

and building a product, uh but today

7:31

I'll just focus on the technical aspects

7:33

of building AI for the physical world.

7:36

Uh and

7:37

each one of those lessons, I think, by

7:39

itself will not be exactly

7:41

earth-shattering.

7:42

Uh you know, a lot of it will overlap

7:44

with likely things you've heard

7:45

elsewhere. But, I hope that the

7:48

grounding of these lessons in our

7:50

experience and some of the nuance that I

7:53

can add about how they showed up in our

7:56

experience of deploying a physical agent

7:58

in the and scaling it safely will be

8:01

interesting and useful for many of you

8:03

who are in the space as you build your

8:05

product, as you build your startup.

8:08

Uh so, let's dive in. The first lesson

8:10

has to do with this

8:12

uh massive

8:14

frustrating

8:16

sometimes

8:18

soul-crushing difference between a demo

8:19

and a real product.

8:21

And a working demo is 1% at best of the

8:24

work that you have to do. The many nines

8:27

of performance, the many nines of

8:28

reliability that follow, that's where

8:30

the real work happens.

8:31

And if you're a founder in the room, uh

8:33

chances are you are focused on getting

8:35

that first prototype, that first demo

8:38

off the ground. And when you hit that

8:40

first version of a system that works,

8:43

that first 90%, when the demo actually

8:45

works, it feels incredible. You feel

8:47

like you solved it, the sky's the limit,

8:49

you're extrapolating forward. And in our

8:51

world, we hit that that first milestone,

8:54

that first 90% back around 2010.

8:58

So, when this project started uh in

9:01

2009,

9:03

before we

9:05

started building the system, we set a

9:07

couple of pretty ambitious goals for

9:08

ourselves. One was to drive 100

9:11

autonomous 100,000

9:14

miles in autonomous mode.

9:16

The second goal was to drive 10 routes,

9:20

each one was 100 miles long, uh chosen

9:23

to cover a wide variety of conditions

9:24

across the Bay Area, and we had to do

9:27

each one from beginning to end without a

9:29

human intervention.

9:30

We had at the time a team of about a

9:33

dozen engineers, and we accomplished

9:35

both of these goals in about a year and

9:37

a half. And keep in mind, this was a

9:39

well before any of the AI breakthroughs,

9:42

before ConvNets, before Transformers,

9:44

before BLMs, before any of the stuff

9:45

that we talk about today.

9:47

Uh and yet, you know, we got it done.

9:50

And kind of by demo standards,

9:52

we driving autonomous driving was solved

9:55

in 2010.

9:56

Right? We handled everything. We handled

9:59

We could drive during the day, during

10:00

the night. We handled traffic,

10:02

pedestrians, cyclists, traffic lights,

10:04

construction zones, on freeways, on

10:06

surface streets. So, we were {quote} and

10:08

{unquote} capability complete.

10:10

And, you know, at the time we felt like

10:11

we're on top of the world.

10:13

But then we quickly ran, as we started

10:15

building towards the product, we quickly

10:17

ran into a brutal reality that there's a

10:19

massive difference between

10:22

doing something once, or driving 10

10:23

routes once, and building a scalable

10:25

service with nobody behind the wheel.

10:30

It took us about 10 more years

10:34

to begin providing a service, and then 5

10:36

more years to scale to half a million

10:39

trips per week. So, the demo took 18

10:41

months, the product took about 15 years.

10:44

But now we're scaling exponentially. To

10:46

date, we've served well over 20 million

10:49

fully autonomous trips,

10:51

and we've driven well over 200 million

10:53

fully autonomous miles.

10:55

And we have rider-only vehicles

10:57

operating in 15 cities across the United

10:58

States.

11:01

And we're scaling exponentially. It took

11:02

us

11:03

uh

11:04

15 years to get to that first 100

11:08

million miles,

11:09

and about 7 months

11:11

to drive the next 100 million.

11:14

It took us about 8 years to go from the

11:16

time

11:18

when we started our initial rider-only

11:20

operation

11:21

to the time when we had uh when we were

11:24

serving riders

11:25

uh in four cities.

11:27

Earlier this year, we launched four

11:29

cities in just 1 day.

11:33

So, why does bridging that gap from demo

11:37

to product takes a long?

11:39

Uh, because there's this harsh

11:41

engineering reality

11:43

that you can't really cheat, that

11:44

reliability and performance lives on

11:46

this exponential ladder of nines. So,

11:48

getting to that first 90% or 99%, that's

11:51

the easy part.

11:53

But then every next nine that you want

11:55

to add, that takes about 10 times more

11:58

effort. So, you need to know up front

12:00

exactly how many nines your product

12:02

actually needs.

12:05

So, demo

12:07

might need, you know, one nine, an

12:08

assist product or, you know, co-pilot

12:11

might need a few, but a fully autonomous

12:13

AI agent that we're going to be putting

12:15

out in the physical world that engages

12:17

with the public, you know, with kids

12:18

running around, that needs a whole stack

12:20

of them.

12:21

And at scale,

12:23

the long tail is the problem space, is

12:26

your entire problem statement. When you

12:28

drive millions of miles per week,

12:31

a rare event that might happen once in a

12:33

million miles, that just becomes your

12:35

daily reality.

12:37

And getting those next nines

12:39

means doing something different

12:42

every time. So, you don't get to say six

12:45

nines of performance or reliability by

12:47

doing the same thing that you did for,

12:49

you know, to achieve the first two, but

12:51

longer. You have to do fundamentally

12:53

different things. You requires a

12:55

fundamentally different approach. For

12:57

example, we can take reliability. You

12:58

can, you know, get to the first couple

13:00

of nines by just doing proper

13:01

engineering and uh

13:03

doing, you know, uh some bug fixes.

13:06

But to get to the next few, you need to

13:07

have You need to invest in fundamentally

13:10

different approaches. You need to build

13:12

fully redundant systems,

13:14

uh have tiered full backup

13:16

architectures, and so forth and so on.

13:18

And the same thing holds for the

13:20

performance of AI models.

13:23

So, what that actually means is that in

13:25

this space, it's incredibly easy to get

13:28

started,

13:30

but it can be excruciatingly difficult

13:32

to get to the real product.

13:34

And that effect is only amplified with

13:38

every wave of technological

13:40

breakthroughs.

13:41

And that naturally leads to hype cycles.

13:44

So, every AI breakthrough from, you

13:46

know, deep learning to conv nets to

13:48

transformers, VL aims, you name it. It

13:51

makes it that much easier to get

13:53

started. Your demos, your prototypes,

13:54

they get a 100 times easier. But the

13:56

tail, that's where the hard problems are

13:59

that moves much less. It moves, but the

14:01

effect is muted. And that's why every

14:03

hype cycle produces a wave of absolutely

14:05

spectacular demos and very few real

14:07

products.

14:09

And the recurring mistake of every cycle

14:10

is spending on the demo when you should

14:13

be saving for the nines.

14:16

Now, you know, this being a startup

14:18

school, the last thing I want to do is

14:19

throw too much cold water on the magic

14:22

and the excitement of those early days.

14:24

This this time is absolutely magical.

14:26

It's amazing.

14:27

Cherish it. Leverage it. But the key is

14:30

to remain honest about the product that

14:32

you're building, the number of nines and

14:34

performance and reliability that that

14:36

product demands, and not cutting corners

14:39

to get there. Uh otherwise, you might be

14:41

in for a pretty rude uh awakening later.

14:43

So, count your nines before you count

14:46

your demo views.

14:48

Uh and this brings us to the second

14:50

lesson.

14:51

Once you know how many nines your

14:53

product actually needs, it fundamentally

14:55

dictates the architecture and the core

14:58

technical approach that you need to

14:59

pursue.

15:01

Now, every technology has a performance

15:05

versus effort curve, right? They all

15:07

tend to start fairly steep and go up,

15:09

and then they flatten out.

15:11

And you know, as I just mentioned, every

15:13

other nine gets an order of magnitude

15:15

more difficult.

15:16

So, common failure mode is picking the

15:19

tech that gives you the fastest early

15:22

ramp,

15:23

riding that steep curve, feeling like

15:26

you're winning, projecting that, you

15:27

know, steep slope into the future and

15:29

feeling like the sky is the limit, and

15:31

then hitting the plateau, and

15:32

discovering that the technology path

15:35

that you picked actually flattens out

15:37

way before the performance that is

15:41

required by your product.

15:43

Now, you might still choose to be, at

15:45

least for a while, on that steep curve

15:48

for a variety of practical reasons. You

15:50

know, maybe you want to prototype

15:51

something or demo something or build

15:53

something in service of learning, but be

15:55

honest with yourself where you're

15:57

building for the purpose of a demo, for

15:59

the purpose of learning, or towards an

16:01

actual product.

16:04

Uh so, let's take

16:05

uh an example from our domain,

16:08

autonomous vehicle and sensing.

16:10

Uh there's been a long-standing debate

16:12

about what kind of sensors do you

16:14

actually need for autonomous driving?

16:16

Naturally, more sensors means higher

16:19

performance, but also means higher

16:21

complexity. So, humans

16:23

of course can drive with just eyes, so

16:25

there's that proof of existence. Now,

16:28

and if the goal were to just

16:30

approximately match human performance or

16:33

to build an assist product, that's a

16:35

very reasonable way to go.

16:37

However, if you are targeting full

16:38

autonomy, and you're targeting

16:41

superhuman, strongly superhuman

16:42

performance, you find that weak sensing

16:45

just leads to a safety curve that

16:46

flattens out way too early.

16:49

So, at Waymo, we've taken an approach

16:51

where we use multiple sensing

16:52

modalities. We use cameras, lidars, and

16:55

radars,

16:56

and they all complement each other.

16:57

Cameras give you high resolution and

16:59

color,

17:00

uh but they're passive and they degrade

17:03

in darkness and glare.

17:05

Lidar gives you a direct measurement of

17:08

the 3D structure of the world uh around

17:10

you, and radar is very good at punching

17:14

through environmental conditions and

17:15

weather like fog or rain or snow and it

17:18

directly can measure velocity using

17:20

Doppler.

17:21

Uh lidar and radar are active sensors.

17:24

Uh so that means they see just as well

17:27

in pitch darkness or

17:29

for example, when driving into a

17:31

blinding sunset.

17:32

And these different sensing modalities,

17:34

of course, they're not backups to each

17:35

other.

17:36

Uh in our stack

17:38

each modality has an encoder and the

17:41

information from all of those sensors

17:42

get fused into a single view of the

17:45

world around us that is much more

17:47

precise and um generally vastly superior

17:51

to what you get with any one sensor.

17:53

So let me show you a few examples.

17:55

Uh here's the scene where a Waymo is

17:58

driving in a dust storm in Phoenix.

18:01

So what you see here is what the scene

18:03

looks like to our fairly advanced

18:07

high-resolution and high-dynamic-range

18:08

camera.

18:09

It's very close to what a cat, you know,

18:11

a human would see in the same

18:12

conditions.

18:14

Not much. And here on the right is what

18:17

the lidar sees for the exact same frame.

18:20

And you can much more clearly see that

18:21

there's a pedestrian standing on the

18:23

side of the road.

18:24

So if they were to step

18:27

onto the road

18:28

that early detection can make a really

18:31

big difference in how the situation

18:32

plays out and the safety of everyone

18:34

involved.

18:36

Here's another example. At night,

18:38

driving along and there are a couple of

18:39

pedestrians who are about to jump onto

18:42

the road over a concrete construction

18:44

barrier. Again, at the bottom you see

18:46

the camera, really can't see much, and

18:48

the lidar view at the top.

18:50

Again

18:51

lidar versus camera.

18:54

Here's another example.

18:59

A couple of dogs chasing a ball and a

19:01

couple of kids chasing the dogs.

19:04

And

19:05

big difference. Here's what it looks

19:07

like to the camera. Here's the lidar and

19:09

the early detection of the uh the kids

19:12

is off to the side and there are no

19:13

headlights, there are no lamps there,

19:14

it's complete darkness. So, it makes a

19:16

big difference.

19:17

Uh or think about what happens when

19:19

something physically obstructs the view

19:21

of your sensors.

19:24

I

19:25

Yeah, if you don't have redundancy in

19:26

sensing, you can have, you know, a

19:28

single leaf land on your sensors and

19:31

bring your robot to a full stop.

19:35

Uh now

19:36

So, you need redundancy.

19:38

Redundancy, of course, does not

19:40

necessarily mean multiple sensing

19:42

modalities, but if you need redundancy

19:44

anyway, you might as well

19:47

uh benefit from the complementary

19:49

physics of the different sensing

19:50

modalities in the nominal case.

19:53

So, here's a a video of

19:56

uh one of our cars that picked up a leaf

19:58

or actually, I think a full uh branch of

20:00

a tree

20:02

uh that our wipers were unable to shake

20:05

and the car detected that and because we

20:06

have sensing redundancy, it safely was

20:09

able to get back to the depot for proper

20:11

cleaning.

20:13

Uh

20:14

so, specifically, when it comes to

20:17

hardware,

20:18

uh

20:19

do not anchor to today's components

20:23

prices.

20:24

Uh

20:26

We're on the sixth generation of the

20:28

Waymo driver, the Waymo hardware suite

20:30

today and with every generation, the

20:34

hardware not only delivered amazing

20:35

capability, but we're able to

20:37

drastically simplify and radically

20:39

reduce the cost of the hardware as well.

20:41

So, betting your company, betting your

20:43

approach on today's hardware prices is

20:46

just betting your company on a number

20:48

that has a fairly short shelf life and

20:51

is going to expire.

20:53

Uh so, hardware will change.

20:55

Uh many components will get commoditized

20:57

and drop in price. So, design for that

21:00

future and be ready to upgrade.

21:05

And then brings us to the next lesson,

21:07

lesson number three.

21:08

Uh technology m- moves incredibly fast,

21:12

especially nowadays.

21:13

So, you need to be ready to ride those

21:16

tech waves and do that repeatedly.

21:18

And when you do,

21:19

as you not only think about the wins and

21:23

performance and the wins and capability,

21:25

you have to be very mindful about uh

21:28

unification and simplification.

21:31

Uh over the years, we've seen a number

21:34

of major breakthroughs in technology, a

21:36

lot of them around AI. And with every

21:39

wave of innovation, we pretty much

21:40

rebuild the WeRide driver around that

21:43

major wave of AI breakthroughs. And we

21:45

often push the state of the art in those

21:46

areas forward ourselves.

21:48

Uh we leverage confidence around 2013

21:52

for computer vision and perception.

21:53

Then, when transformers came about

21:56

around 2017, we bet big on them both for

21:59

perception and for the task of behavior

22:01

prediction and decision-making and

22:03

planning. Turns out, the task of driving

22:06

is not that dissimilar from the task of

22:08

modeling language uh because of the

22:11

social aspects of driving. You're kind

22:12

of having a conversation with other

22:14

dynamic actors in the world, but you're

22:15

doing that in the space kind of body

22:17

language of your agent, your car, as

22:20

opposed to just the language of words.

22:23

Uh and you operate in sequences and

22:25

local continuity matters, uh but so does

22:27

global context. Uh and today we're

22:29

leveraging uh the latest in VLMs and

22:32

frontier world models.

22:35

Now,

22:36

using the latest tech for capability and

22:40

performance wins, I don't want to say

22:42

it's easy, but it can be, you know,

22:44

reasonably straightforward doing applied

22:45

research in isolation or starting a

22:47

tiger team to, you know, prototype uh

22:49

some new technology

22:51

uh is not the most difficult part.

22:53

There's many companies, many teams that

22:54

are excellent in this.

22:56

The much harder muscle to build is to

23:00

carry that bleeding edge research into

23:03

production and deploy it in a safety

23:05

critical environment without

23:06

regressions.

23:08

And do it without breaking stride on the

23:10

scaling of your product.

23:13

And adding capability, again, is not the

23:16

hardest part, but adding capability

23:18

while at the same time reducing

23:19

fragmentation and reducing complexity,

23:22

that is really important. And finally,

23:24

the hard muscle to build as a company is

23:26

to be able to do that repeatedly through

23:28

multiple waves of technical innovation

23:30

or technical breakthroughs.

23:32

So, on this front, I have two bits of

23:34

advice.

23:35

Uh the first one, when a technology, a

23:37

new technology shows up,

23:39

yeah, it can be very exciting, very

23:41

tempting to kick off uh a new effort, a

23:44

tiger team to pursue it. And that's

23:46

great, you should absolutely do that.

23:48

However, when you do,

23:50

it's very important that you consider

23:52

what you would do after. Under a success

23:54

scenario, let's say that effort

23:55

succeeds,

23:57

you should be very clear on what the

23:58

path of that new innovation is for your

24:01

company, for your entire product, for

24:02

your entire system. Uh often times I've

24:04

seen a failure mode where, you know, a

24:06

project, a very difficult technical

24:08

project succeeds,

24:10

and then there's a dead end. That can be

24:12

very wasteful, that can be completely,

24:14

you know, deflating.

24:16

Uh the second bit of advice I have here

24:18

is uh when pursuing new tech,

24:22

again, don't just ask what does this new

24:23

tech give me in terms of capability and

24:25

performance, also ask has it simplified

24:27

my stack and has it led to fragmentation

24:30

or unification. So, set your launch bar

24:32

to demand both breakthrough performance

24:34

and at the same time radical

24:37

simplification and unification.

24:40

And this exact philosophy and this

24:42

muscle that we've built uh at Waymo over

24:44

the years is what produced our latest

24:47

core technology.

24:48

Uh and this

24:50

uh the heart of it is the Waymo

24:51

foundation model.

24:52

Now, the Waymo foundation model is a

24:55

multimodal

24:56

world action language model. It's kind

24:58

of a mouthful, so let me unpack the

25:00

ingredients. It's a multimodal model

25:03

because it is able to process these

25:05

multimodal sensor inputs, cameras,

25:08

lidars, and radar.

25:09

Uh it's a world model cuz it inherently

25:11

understands how the world works, the

25:14

physics, the dynamics, as well as the

25:16

social and semantic aspect of it.

25:18

It's an action model because we are not

25:20

just passively observing how the world

25:23

evolves. Uh we're an active participant.

25:26

So, the model needs to understand the

25:28

effects of our the actions of our agent

25:31

on the world and be able to tell the

25:33

good ones from bad ones.

25:35

And finally, it's aligned with language

25:38

and that allows us to unlock general

25:40

world knowledge from visual language

25:43

models. And that's incredibly useful in

25:45

the long tail of rare semantic uh

25:47

situations.

25:50

So, more specifically, uh this is what

25:52

the architecture looks like. It's kind

25:54

of your typical encoder-decoder

25:56

architecture.

25:57

The encoder part takes in the multimodal

26:00

sensing and compresses it or encodes it

26:02

into a efficient into an efficient

26:04

representation uh that

26:07

retains all of the relevant data, all of

26:10

the relevant information for the

26:12

generative part or the decoder.

26:15

It's an end-to-end model um which has a

26:18

couple of nice properties. It allows us

26:20

to effectively backpropagate the

26:21

gradient from the task that we actually

26:23

care about all the way to the early

26:25

layers of the model.

26:27

And it allows the encoder to reach you

26:30

kind of learn the right rich

26:31

representations

26:33

uh for what the generative part needs to

26:35

solve the task. Uh it uses a system one

26:38

system two think fast think slow

26:40

architecture and it leverages the

26:42

general world knowledge of VLMs for

26:45

efficient learning of semantic tasks.

26:47

So, let's dive uh deeper.

26:49

Uh first, the

26:50

think fast path. Uh that part fuses the

26:53

raw data from our cameras, our lidars,

26:55

our radars, and that allows for

26:58

split-second safety-critical decisions.

27:00

So, you can think of it as kind of your

27:01

driving instincts. Uh this is what

27:03

allows the car to brake instantly if,

27:06

let's say, a pedestrian runs into the

27:08

road or cyclist that's nearby swerves

27:10

into your path. Uh this is like the the

27:13

if you will the lizard brain of your

27:16

agent that deals with a lot of geometric

27:19

tasks and can react in milliseconds.

27:22

Uh second is the slow path. Uh that's

27:25

the part that's responsible for the more

27:27

complex

27:28

uh semantic and scene-level

27:30

understanding type tasks. And these sort

27:33

of tasks, these things don't typically

27:35

change in milliseconds. So, there you

27:38

can afford a bit more latency, and you

27:41

can trade that off

27:42

for higher capability and uh higher uh

27:46

levels of uh of of reasoning.

27:49

So, for example, if the Waymo driver

27:51

encounters a situation when there's, you

27:54

know, a vehicle, let's say it's on fire

27:55

on the side of the road, the fast path

27:58

might just see it as a generic generic

28:01

obstacle and, you know, reason that the

28:03

path ahead of us is clear.

28:05

And this is where the slow path comes

28:06

in, and that path can use deep semantic

28:09

reasoning to understand the semantics of

28:11

that object, the car being on fire, and

28:13

the broader scene context, and that

28:15

allows our driver to decide to take a

28:18

very, you know, different action and our

28:20

different route entirely. Even if

28:23

geometrically the path ahead of us is is

28:25

clear.

28:27

Uh and finally, there's the generate

28:28

component. That's the decoder. Uh that's

28:31

the component that understands and can

28:32

produce behavior. Uh it understands how

28:35

other actors behave, uh and it allows us

28:37

to make predictions and plan our own

28:40

driving decisions.

28:43

And our Waymo foundation model powers

28:45

the Waymo my

28:46

that runs on different generations of

28:49

hardware and runs on different vehicle

28:52

platforms. You have our fifth

28:53

generation, the sixth generation, the

28:54

JLR I-Pace, the Olli and and the Hyundai

28:58

Ioniq. And in the future will power

29:00

different products in different

29:02

commercial applications like trucking

29:04

and personally owned vehicles.

29:06

So by uh

29:09

leveraging the strategy of focusing on

29:12

the high capacity foundation of uh of

29:15

our model

29:16

we're able to move a lot of complexity

29:18

upstream to that large shared foundation

29:21

and that allows us to make that

29:23

specialization layer

29:25

uh that's running on the car

29:26

uh

29:27

uh pretty lightweight.

29:29

And that in turn allows us to speed up

29:30

the development process.

29:32

So the most important muscle in this

29:34

lesson uh is

29:36

for your company to not just leverage

29:39

the

29:40

tech of the day

29:42

but have the ability and build that

29:44

muscle to repeatedly ride those tech

29:46

waves and pull in the results of that

29:48

innovation into production without

29:50

regression, without breaking stride in

29:52

deployment and scaling, and without

29:53

drowning in complexity.

29:57

So let's move to the next lesson.

29:59

Uh

30:00

There's a well-known lesson in the AI

30:04

community

30:05

that

30:06

general methods that leverage massive

30:09

compute and massive data will always

30:13

beat

30:14

uh methods that rely on handcrafted

30:17

engineered human knowledge.

30:20

That's the so-called uh bitter lesson

30:22

that Richard Sutton uh published and

30:25

formulated in 2019.

30:27

And we have lived this and we have seen

30:29

this in every wave of technical

30:31

breakthroughs. Each time the bitter

30:33

lesson holds, methods that scale best

30:35

with compute, with data, they always win

30:38

out.

30:39

And by the way, uh this is uh you know,

30:41

one of the reasons why we bet on the

30:43

approach of building the foundation

30:45

model. Um,

30:47

there is a well-known property that if

30:50

you bet on high-capacity model and you

30:52

use your data and your compute on that,

30:55

you just get better scaling laws and

30:57

then you distill into smaller, more

30:59

efficient models that are running on

31:01

your agent in real time, you just get

31:03

better scaling laws as opposed to just

31:04

focusing on the smaller models directly.

31:08

Uh, so one

31:10

nuance area where uh, this lesson shows

31:13

up is the use of structure in your

31:16

models.

31:18

And depending on how you use your

31:19

structure,

31:21

you can end up on either side of the

31:23

bitter lesson.

31:25

Essentially, structure that fights scale

31:28

will always lose.

31:29

And structure that channels scale always

31:31

wins.

31:33

And in particular, this comes up around

31:34

the discussion of end-to-end models. As

31:37

I mentioned, an end-to-end model is, you

31:39

know,

31:40

uh, has some very nice properties. You

31:41

back propagate a gradient from the final

31:43

task all the way through the model and

31:45

it allows the API between the encoder

31:48

and the decoder to learn to use rich

31:51

learned representations.

31:53

And, you know, those are the easiest

31:54

models to build and train. Um, you know,

31:57

you can start uh, the architectures are

31:59

known. You can start with doing some

32:00

imitation learning in a

32:02

uh, kind of a black box end-to-end model

32:04

will give you very rapid progress and

32:06

you will ride that very, you know,

32:07

initial steep part of the curve. Uh, and

32:10

for some products, that's enough.

32:11

But if you need to reach superhuman

32:14

levels of performance in a fully

32:15

autonomous agent in a safety-critical

32:18

environment, uh, just doing kind of that

32:20

basic vanilla end-to-end is not enough.

32:22

And this is where structure comes in.

32:24

And the key question here is,

32:27

does the structure boost scale or does

32:29

it fight it? Does it limit and constrain

32:32

your solution space?

32:34

Or does it help you scale without loss

32:37

of generality?

32:39

So, let me give you

32:40

an example. Let me illustrate this point

32:42

with kind of a simple thought exercise

32:44

and a toy problem.

32:46

Imagine

32:47

you're building a robot

32:50

that will play the game of Go.

32:52

And it wants you know you want it to

32:53

play the game in the physical world,

32:55

right? So, you have a camera that's

32:58

observing the board and you have an

32:59

actuator that will actually move the

33:01

pieces around. Now, one way you can

33:03

build such a robot is to you know have

33:06

an end-to-end system that goes directly

33:07

from pixels to actuation and maybe you

33:09

train it by giving it some videos of you

33:11

know how humans play the game.

33:13

And that could be a very interesting

33:15

research exercise. However, if your goal

33:17

was to build the world's best playing Go

33:20

robot, that's probably not the most

33:21

efficient way to go.

33:23

And

33:24

the reason for that is that there is a

33:27

very simple

33:30

intermediate representation that

33:31

captures completely the state of the

33:35

game, the state of the the task that

33:37

you're trying to solve as the 19 by 19

33:39

board.

33:40

And that gives you a fully observable

33:42

and complete state of the world that you

33:43

care about at least for the you know

33:45

game playing length part. So, leveraging

33:47

that structure it doesn't limit your

33:49

model. It doesn't constrain your

33:50

solution space, but it gives you a very

33:52

helpful way to scale.

33:54

Now, that of course was a toy example.

33:57

Anything that's not trivial that you're

33:59

trying to deploy in the physical world

34:01

will not have that property. And the

34:03

fact that such a simple clean engineered

34:05

representation doesn't exist in the

34:06

physical world is the whole reason why

34:08

we need end-to-end systems and learned

34:12

representations and learned embeddings.

34:14

But in the physical world, there does

34:15

exist structure.

34:17

You have laws of physics, you have rules

34:19

of the road, you have objects that

34:20

behave in reasonably predictable ways.

34:22

And you can use that structure

34:25

in addition to the learned

34:26

representations to boost your

34:27

performance,

34:28

simplify validation, and And the end of

34:30

the day just get better scaling laws.

34:33

Uh and this is the approach that we are

34:35

pursuing at Waymo, which we call

34:37

structure-augmented-end-to-end.

34:39

So, we go beyond the basic vanilla

34:40

end-to-end by augmenting the learned

34:43

embeddings with materialized structure

34:45

representations.

34:47

And that gives us a few very important

34:48

advantages.

34:50

So, first

34:51

is validation at inference time.

34:54

Now, because the model isn't just a

34:55

black box where sensors go in and you

34:58

know,

34:58

actuation commands go out,

35:00

we can create a very powerful

35:03

correctness and safety validation layer

35:05

that you can run in real time in

35:09

when the agent is deployed on our

35:10

vehicles.

35:12

And this is really important for any

35:13

agent that's operating in the physical

35:14

world.

35:17

Uh secondly,

35:18

we get great wins in efficiency when it

35:21

comes to large-scale training and

35:23

evaluation of the generative part of the

35:25

model, the decoder. Uh if all you have

35:29

is a black box end-to-end system,

35:31

you are forced to do all of your

35:33

evaluation and all of your training in

35:35

the end-to-end setup all the way from

35:37

sensors to decisions to actuation.

35:39

Having that intermediate structured

35:41

representation allows you to kind of mix

35:43

and match. You can do some training at

35:45

larger scale

35:47

and some evaluation in the space of

35:48

those compact structure representations

35:51

and some in the full space of end-to-end

35:53

from sensors to decisions.

35:55

Uh and finally, we get strong verifiable

35:59

feedback signals for both evaluation and

36:02

for training

36:03

training recipes to support things like

36:05

reinforcement learning. That additional

36:07

materialized structure just gives you

36:09

much more powerful tools for evaluation,

36:12

for metrics, as well as crafting your

36:15

loss function or

36:16

reinforcement learning recipes.

36:18

So, the lesson here is to bet on a

36:20

system that's maximally learned and

36:22

minimally constrained

36:24

and leverage structure intentionally to

36:26

boost performance and scaling laws both

36:29

in training and in evaluation.

36:33

Now, that raises the question of how do

36:35

you actually train and evaluate your

36:37

physical AI agent?

36:39

And that brings us to the next lesson.

36:41

To build

36:43

and safely deploy an agent in the

36:45

physical world, it is absolutely

36:47

critical to have a good large-scale

36:50

realistic high-fidelity simulator.

36:53

Now, there's two ways you can do

36:55

training and evaluation. You can do open

36:57

loop and you can do closed loop.

37:00

Uh in open loop, kind of passively

37:02

observing uh input to output pairs, uh

37:05

and you can use that, you know, for

37:07

evaluation, for training imitation

37:08

learning works like that. Evaluation

37:10

takes usually the shape of, you know, if

37:12

you are find yourself in this situation,

37:13

what would you do? And then you you

37:15

score that. Uh and that's in contrast

37:17

with closed loop, where in closed loop,

37:19

you take an action,

37:21

uh you see the effect that that action

37:23

has on the world,

37:25

then uh you update your uh through your

37:28

sensors the view of the world, you take

37:30

another action, and so forth and so on,

37:32

and you evaluate and you train on those

37:34

sequence of actions and sequence of

37:35

world evolutions. Now, the ability to

37:38

take an action and evaluate that

37:40

counterfactual

37:42

is absolutely vital for building and

37:45

deploying safety-critical agents in the

37:47

physical world.

37:50

Uh so,

37:52

a real simulator

37:54

is how you do that. And a real simulator

37:56

isn't just some lightweight tooling that

37:58

sits next to your AI.

38:00

Uh it is

38:03

an uh uh a big AI model in of itself.

38:06

And the problem of building a good

38:09

realistic simulator is just as hard as

38:11

building the agent itself.

38:13

Uh so, the AI behind the simulator

38:15

really needs to understand how the world

38:17

works, the physics, the semantics, you

38:19

know, the traffic, the weather, and so

38:20

on and so forth. And the quality of that

38:23

simulator has to be high enough so that

38:25

it doesn't only look good, but it's

38:27

sufficient to train and evaluate with

38:30

high confidence an agent that you're

38:32

going to be putting in the world in a

38:33

safety-critical environment. So, in

38:35

other words, you have to build

38:38

a highly accurate generative world

38:40

model.

38:42

And at Waymo for years, we've been

38:44

building what we called behavioral world

38:47

models. And we're doing that way before

38:49

the term models even became popular.

38:53

And now in the era of end-to-end models,

38:57

you also need on top of behavioral

38:59

realism, you need sensor realism as

39:00

well.

39:01

And in fact,

39:03

building an end-to-end model has been

39:06

fairly easy for quite a while now, but

39:09

evaluating it in closed-loop,

39:12

that was the hard part of the problem.

39:14

So, we've moved on to building sensing

39:17

world models. And because we're using

39:19

that structured cemented representation

39:21

in our models,

39:22

we can also leverage that structure in

39:24

our simulation.

39:26

Our behavioral world model operates in

39:28

the space of structured intermediate

39:30

representations.

39:31

And the tightly coupled sensor world

39:34

model then producing produces realistic

39:38

sensor simulations.

39:40

Our world model leverages

39:43

the great work of Google DeepMind's on

39:46

Genie 3. And that gives us the ability

39:48

to produce controllable and highly

39:50

realistic scenarios both in the

39:52

behavioral as well as sensing aspects.

39:56

And that in turn allows us to not just

39:59

evaluate our agent and train new

40:03

versions of our agent in situations that

40:06

we've previously encountered, but allows

40:08

us to train and evaluate in purely

40:10

synthetic rare scenarios that we've

40:13

never seen in the real world.

40:16

So, what you're seeing here is not just

40:19

a generative video, it's a generative a

40:21

full generative simulation of the Waymo

40:23

driver

40:24

in

40:25

operating closed loop. So, here we're

40:27

simulating what would happen if it came

40:29

across a car that was stopped in the

40:30

lane on the freeway.

40:32

And you can go further than that. Here's

40:34

a plane

40:36

that's landing on a freeway in front of

40:37

us.

40:39

Or you can simulate an elephant on the

40:41

loose walking through the intersection.

40:45

Snow on the Golden Gate Bridge.

40:48

Or a dinosaur walking around. So, the

40:50

lesson here is that closed loop

40:52

simulation is absolutely required for

40:55

evaluation and is extremely valuable for

40:58

training of your physical AI agents.

41:01

So, you need highly realistic

41:04

large-scale

41:05

simulation to trail train and evaluate.

41:08

And this brings us to lesson number

41:10

six.

41:12

When you're dealing

41:15

with a problem of that complexity, you

41:18

can't just build a model and call it a

41:20

day.

41:22

You have to build an entire ecosystem.

41:24

And then you also need a flywheel that

41:26

powers it.

41:28

Because to make this work at scale, you

41:30

can't just build a agent you and one AI,

41:33

you need to build three.

41:35

You're building the agent.

41:37

For us, that's the driver that drives

41:39

the car.

41:40

You also have the simulator.

41:42

Which is that virtual playground for the

41:45

agent to learn in.

41:46

And then you have the critic.

41:48

And the critic is what rigorously

41:50

evaluates and judges the performance of

41:52

the agent and tells it how to improve.

41:55

And the good news is that

41:57

the fundamental reasoning and the

41:59

generative capabilities of all three of

42:01

those are shared and that's why in our

42:03

case they're based on the same

42:05

foundation world model.

42:11

Now, once you have these three pillars,

42:14

you can create an incredibly powerful

42:16

flywheel to accelerate your progress.

42:20

So, a deployment of your agent uh in the

42:22

real world generates data.

42:24

That data then grounds the simulator and

42:27

makes it more realistic.

42:29

The simulator generates harder edge

42:32

cases for the critic to score and for

42:35

the agent to learn from. So, the agent

42:36

gets smarter, gets deployed in the

42:38

physical world, generate more data, and

42:40

that powers the flywheel and accelerates

42:42

progress.

42:44

But, a flywheel, of course, will, you

42:47

know, spin in

42:49

any direction or in place.

42:51

So, in order to make it go in the

42:53

direction you want, you need to guide it

42:55

by metrics.

42:56

And that brings us to the final lesson,

42:58

that your model is really table stakes,

43:01

but eval and metrics, that's your most

43:04

important. That's your strategic moat.

43:08

So, build your eval

43:11

before you build your technology. Build

43:13

your eval and your metrics before you

43:15

build your product.

43:16

If you can't

43:18

quantitatively define what good enough

43:20

means, you're not really building a

43:22

product, you're just iterating on your

43:24

demo.

43:25

So, nowadays,

43:27

the best model architectures are fairly

43:29

well known, and new ideas tend to uh

43:33

proliferate fairly quickly.

43:35

Data

43:37

is incredibly important,

43:39

but without good metrics, you're just

43:41

flying blind. You aren't leveraging the

43:43

best data, and you can't really evaluate

43:45

the ROI on making changes to it. So,

43:48

really, eval and metrics, that that's

43:51

your foundation, and that's what steers

43:54

your whole tech stack.

43:57

Uh but, for physical AI agents,

44:01

model-level evaluation is not enough.

44:04

When you're putting an AI agent into the

44:06

physical world, your eval and your

44:09

validation needs to go much deeper and

44:11

much broader.

44:12

You need to evaluate and validate every

44:15

component of your system

44:17

from the physical layer to the

44:19

behavioral layer uh that's running on

44:21

board in the physical world, as well as

44:23

the off-board components,

44:25

and uh all of the operational processes

44:27

around it. So, for us, we call that the

44:29

safety and readiness framework, and we

44:33

spent years building and refining it,

44:35

and uh that's what guides our

44:37

development and our deployment and our

44:39

scaling, and I consider that to be one

44:41

of our most important uh assets.

44:46

Uh bec- And this again,

44:49

the reason it's important is because in

44:51

the physical world, trust is everything.

44:54

And eval and metrics is how you go about

44:56

earning that trust.

44:58

You don't just win

45:00

trust by talking about, you know, the

45:02

clever technical solution or the clever

45:05

state-of-the-art architecture of your

45:06

models,

45:08

or doing, you know, flashy demo. You

45:10

earn it gradually, day by day, by uh in

45:14

the field, by relentlessly proving that

45:17

your system is safe and that your system

45:19

works.

45:20

And of course, you can't just prove that

45:22

to yourself behind closed doors, and

45:24

this is exactly why we openly publish

45:26

our safety data and our uh safety

45:30

ongoing safety research.

45:32

So, then that earned trust becomes your

45:35

ultimate business advantage. All right?

45:37

Your your models

45:39

can be leaked, algorithms can be

45:41

replicated, but hundreds of millions of

45:44

miles of fully autonomous operations in

45:46

the real world, backed by evidence-grade

45:49

evaluation and publicly audited proof,

45:51

that is much, much more difficult to

45:52

replicate.

45:55

So, when you zoom out and look at this

45:57

playbook as a whole, you realize that

45:58

none of these lessons works alone.

46:01

So, the nine set your bar and ensure

46:03

that you pick the right technology and

46:05

the right technical approach so that you

46:06

don't get stuck in the local minimum.

46:09

Uh then intentional use of structure to

46:11

boost scaling

46:13

uh and the ability to ride technical

46:15

waves of innovation and gets helps you

46:17

get to the right level of nines.

46:19

Uh and your AI ecosystem with the agent,

46:22

the simulator, and the critic guided by

46:24

your eval and metrics, that's what

46:26

allows you to build that powerful

46:27

flywheel and that's how all of these

46:30

effects uh compound.

46:32

And it's this playbook that we've been

46:35

refining uh over the years is what

46:37

allows us to achieve the strongly

46:40

superhuman safety performance of the

46:41

Waymo driver.

46:43

Uh this is a snapshot of the latest

46:45

safety data we've released is based on

46:47

over 220 million fully autonomous miles

46:49

and we're seeing there that in the areas

46:51

where we operate,

46:53

the Waymo driver is about 17 times

46:55

better than human drivers when it comes

46:58

to crashes

47:00

uh that cause serious injury.

47:03

And that really matters

47:05

uh because today

47:08

somewhere in the world, every 26 seconds

47:11

someone loses their life on a road to a

47:15

crash event.

47:16

And at the current scale, what that

47:17

means is that Waymo is preventing a

47:19

serious injury every 8 days. And this

47:22

isn't just a metric on a dashboard, that

47:23

means that someone's loved one got to

47:27

walk through the front of door at the

47:28

end of the day safe and unharmed.

47:32

So, these are just the early safety

47:35

uh benefits of AI in the physical world

47:37

and they will only grow from there.

47:39

If you look at the broader landscape,

47:41

the opportunity here is absolutely

47:43

massive.

47:44

Uh physical AI right now is where

47:46

digital AI was a few years ago and we

47:48

have all of the right ingredients to go

47:49

after it.

47:51

Uh we have generative world models, we

47:52

have the architectures, we have

47:54

affordable compute and sensing, we have

47:56

proven scaling laws, and we have a real

47:58

product operating at scale.

48:01

And the last decade of AI happened in

48:03

the digital world, and the next decade

48:06

will also happen in the physical world.

48:07

And for those of you who decide to build

48:10

in this space,

48:11

uh good luck.

48:13

Have fun.

48:14

And remember who you're building for.

48:17

Your mission and your customers, that's

48:19

what matters. Uh otherwise, tech is just

48:21

a science project. And at the end of the

48:23

day, as exciting as exhilarating the

48:26

tech is, nothing really beats the joy of

48:29

making a difference in people's lives.

48:32

>> [applause]

48:38

>> Wait, what are we doing?

48:39

>> We're in our first ever Waymo.

48:40

>> And what does it mean when we're in a

48:42

Waymo?

48:42

>> It means that there is nobody

48:44

>> Nobody driving this thing!

48:46

>> And uh this is a fully autonomous Waymo

48:49

ride.

48:49

>> I cannot believe this. The car did a

48:52

better job than the if somebody was

48:54

driving.

48:57

>> The truck was over the yellow line, so

49:00

the Waymo braked and moved to the side.

49:06

>> It knew how to pronounce my name.

49:08

>> Oh my god. You can do this, man.

49:09

>> Oh, this is the I say [music] Waymo.

49:11

This is new.

49:15

>> This is so cool.

49:18

>> I'll never forget this.

49:20

Never.

Interactive Summary

The video outlines seven lessons learned from developing Waymo's autonomous driving technology. It highlights the vast difference between creating a working demo and building a safe, scalable physical AI product. Key technical strategies include utilizing redundant sensor modalities, adopting structure-augmented end-to-end models, and building an ecosystem that includes a high-fidelity simulator and a critic. The speaker emphasizes that safety, metrics, and evaluation are the core pillars for building trust and achieving long-term success in the physical AI domain.

Suggested questions

4 ready-made prompts