HomeVideos

Jeff Dean: The 1% Rule for Building in AI

Now Playing

Jeff Dean: The 1% Rule for Building in AI

Transcript

1435 segments

0:07

All right. Should we go Should we get

0:09

started, Jeff?

0:10

>> Sure. Sounds great.

0:11

>> All right. Jeff, welcome. And again,

0:13

thank you so much for being here.

0:14

Especially I just got a cold and thank

0:17

you for being here.

0:17

>> Yeah, I'm afraid I've lost my voice. I

0:19

don't normally sound quite like this,

0:20

but we'll we'll do what we can.

0:24

>> So, um, you built map reduce, big table,

0:27

tensorflow, the TPU, Gemini. We could

0:30

spend a whole hour on all the things

0:32

you've done, but what I love is that

0:34

you're still making bold predictions in

0:36

public. Last year, yes, last year in May

0:42

2025 at AI Ascent, you said that AI is

0:46

at the level of a junior engineer.

0:50

That was about a year ago. It's been How

0:55

close are we to that prediction? Yeah, I

0:58

mean I feel like uh the models have been

1:00

getting a lot better at sort of

1:01

agent-based longer running coding tasks

1:04

and it seems pretty clear that they are

1:06

now actually pretty capable and

1:08

depending on exactly your definition of

1:10

of junior engineer it seems pretty

1:12

spot-on I would say.

1:15

>> What did you underestimate from that

1:18

prediction?

1:19

Um I mean I think the

1:24

the ability to do more and more complex

1:26

tasks has been growing faster than I

1:28

thought. Um and I also think uh outside

1:32

of coding these these agent-based

1:34

systems are are really starting to shine

1:36

in other domains and I I think uh you

1:39

know uh that's that's going to be an

1:41

important trend in the future.

1:44

>> So [snorts] give us another bold

1:47

prediction. What do you think is going

1:48

to be the 2027 edition?

1:51

>> Uh I think you will see a lot more

1:54

automation of uh ML systems themselves.

1:58

um basically getting ML systems to

2:00

improve their capabilities by running

2:03

lots of experiments, breaking things

2:04

down into subpros, you know, running

2:07

those subpros in a tight automatic

2:09

experimentation loop, putting the

2:11

results together and being able to then

2:14

uh you know, get some improved system uh

2:17

out from that uh sort of fully automated

2:20

problem decomposition and and automated

2:23

experimentation that I think that's

2:24

going to be really exciting. M

2:25

>> I think that also applies not just to ML

2:28

but also to other fields of science and

2:30

engineering. Um basically anything where

2:32

you can have a measurable objective uh I

2:35

I think you can uh actually make a lot

2:37

of progress these days.

2:40

>> Now let's go back to a little bit in

2:42

history. Back in way back in 2001 Google

2:47

search used to run on hard drives.

2:49

>> Yep. And you and Sanjay did the math and

2:53

realized that at some point the whole

2:55

search index would finally fit in all of

2:58

the RAM of all the computers you had

3:02

running

3:03

and you made that radical realization

3:07

and you basically in few days with

3:09

Sanjay shipped in production a whole new

3:12

search version that worked in RAM rather

3:14

than hard drive and that was the thing

3:17

that got Google to be so fast. Google

3:19

searches.

3:21

>> So history tends to remix.

3:25

What is the it fits the memory moment

3:28

right now in 2026 that everyone in this

3:31

room is still

3:34

should be thinking about and designing?

3:36

>> Yeah. Yeah, I mean it's a little

3:38

different, but I think uh you're going

3:40

to see more and more

3:43

uh uh high performance and um low energy

3:48

uh inference hardware systems because I

3:51

think everyone is now realizing that

3:53

inference is the key to making you know

3:55

these agent-based systems be available

3:58

to more and more people and that latency

4:01

is really important and that

4:03

specialization of the hardware is a

4:05

really key way you can make uh things

4:07

that are more energy efficient and lower

4:10

latency than more general purpose uh

4:12

computational devices like say GPUs or

4:14

TPUs

4:16

>> because I think all everyone here is

4:18

used to waiting for responses on on

4:21

models. [clears throat]

4:22

So

4:22

>> waiting is no fun

4:25

>> master speed.

4:27

So you're saying what if we don't have

4:29

to wait anymore?

4:30

>> Yeah. I mean, I think we'll imagine what

4:32

you could do with something where the

4:34

latency is, you know, 50x better.

4:38

>> Interesting thought.

4:40

Now, what's one assumption that perhaps

4:43

6,000 people in this room hold that's

4:46

already false about AI?

4:50

>> Yeah. Uh, that's that's a good question.

4:51

I mean I think um

4:55

probably one thing is people don't quite

4:57

realize how possible it is to have you

5:00

know agent-based systems that can run

5:02

not just for an hour or two hours on a

5:05

problem you care about but for some

5:08

problem domains and with highly capable

5:09

models underlying them you can get them

5:12

to run for days or weeks and do really

5:15

really complicated tasks and I think

5:17

that's you know starting some people are

5:20

starting to see inklings of this but I

5:22

don't think everyone has really

5:23

internalized this and that's going to be

5:25

really uh a pretty big deal.

5:28

>> What's a particular task that you have

5:29

run that has run for weeks? What what

5:32

was it? What did the tell what did you

5:34

tell the agents to solve? Yeah, I mean I

5:36

think uh you can tell agents to uh go

5:40

off and implement

5:43

um you know completely new versions of

5:46

software in different programming

5:47

languages that might be you know have

5:49

better safety properties or better

5:51

performance properties uh that and then

5:53

then they can go off and and actually do

5:55

that in a you know pretty serious way.

5:58

>> That's pretty cool.

6:00

[clears throat] Now, one thing that

6:02

you've been very well known for is

6:04

you're really good at napkin math.

6:07

Sounds funny. So, one of the stories

6:10

about you is that back in uh 2013 when

6:14

speech recognition started to work at

6:16

Google, you did the nap napkin math

6:19

where if every Google user used their

6:23

phone and talked to it and used the

6:24

speech recognition system for three

6:27

minute just three minutes a day, you

6:29

found that the system requires a Google

6:32

server. you would have to double the

6:33

fleet which would be really really

6:36

expensive just to do speech translation.

6:39

>> Yeah.

6:40

>> And instead you basically built a custom

6:44

ship and that was the origin story of

6:47

the TPU.

6:48

>> Yeah. Yeah. I mean I I sort of had done

6:51

you know we were starting to see really

6:53

good uh quality results on this the sort

6:57

of deep learning based speech systems uh

6:59

speech models we were training. um but

7:02

they were computationally expensive

7:03

compared to the old speech system but

7:05

they haved the error rate. So that was

7:07

like the equivalent of 20 years of

7:09

advances in speech recognition in just a

7:11

few months of like fiddling with the

7:13

model and getting scaling it up a bit

7:15

and getting better data. And so we

7:19

started to get worried that if speech

7:20

worked a lot better, people would use it

7:22

more. And so that that back of the

7:24

envelope calculation was really about

7:26

that like well what if people start

7:29

start to use speech recognition more to

7:31

dictate emails or to talk to their phone

7:33

or whatever. Um and yeah it turned out

7:37

that um we realized that we needed some

7:41

better solution than running on CPUs at

7:44

the time. And so we came up with TPUs

7:47

which are sort of very specialized for

7:50

essentially low precision dense linear

7:52

algebra which is at the heart of nearly

7:54

all of the modern machine learning

7:56

algorithms we we use today. And um if

8:00

you build a specialized chip for low

8:02

precision dense linear algebra and can't

8:04

do anything else that turns out to be

8:06

really useful for machine learning

8:07

inference uh even though it can't run

8:09

Chrome or Word or whatever. Uh, and so

8:12

that system produced a chip a couple

8:14

years later that was uh 30 to 80 times

8:19

more energy efficient than CPUs and GPUs

8:22

of the day and also much much lower

8:25

latency like 20 to 30x lower latency

8:28

>> which is incredible what the foundation

8:30

that TPU has become today. No way you

8:32

would have predicted that TPU would be

8:34

so foundational now with transformer

8:36

architecture which was invented way

8:38

later before you actually invented the

8:40

TPU. Yeah, I mean that's sort of why we

8:43

built a general purpose linear algebra

8:46

system, which is what a TPU is really.

8:48

Um, because we knew ML algorithms were

8:51

still evolving and you didn't want to

8:53

over specialize, but you wanted to

8:55

specialize enough that you got the

8:57

dramatic performance benefits of we

9:00

could have very big multiplier units. uh

9:03

we could have you know high-speed memory

9:04

we could have high-speed interconnect or

9:06

later TPUs that like brought many many

9:09

chips to bear on the same problem

9:11

efficiently and u you know we've

9:14

continued to scale those up and and

9:15

improve their performance uh for over

9:18

many many generations now

9:20

>> incredible napkin math so

9:22

>> what's good

9:23

>> napkins are good

9:25

>> so actually what's a good napkin math

9:28

that everyone here who wants to be a

9:31

future founder

9:32

should run tonight to potentially build

9:35

something as consequential as the TPU.

9:39

>> Yeah, I mean, uh, it's always hard to

9:42

say. Um, [clears throat] I think, uh,

9:46

think about what problems you see in

9:49

whatever it is you're thinking about,

9:51

what what bottlenecks you see, and are

9:54

there very different ways of thinking of

9:56

the solutions to some of those problems

9:58

that would get you, you know, an order

10:00

of magnitude or two orders of magnitude

10:02

better uh, performance or capability or

10:05

whatever it is. Um, you know, because

10:08

sometimes if you just squint at a

10:10

problem and you think about not

10:13

necessarily being anchored on exactly

10:14

how that problem is solved today, but

10:16

how you would solve it from first

10:18

principles, you can come up with really

10:20

good ideas that are, you know, maybe not

10:23

what other people are thinking about.

10:25

>> That's a good tip. [clears throat]

10:27

>> No. Um, for everyone here who doesn't

10:29

know, years ago, Jeff wrote a very

10:32

famous list called the latency numbers.

10:36

every engineer should know

10:38

and these are numbers around for example

10:41

how long a cache miss takes uh disk seek

10:46

a network package traveling let's say

10:48

from California to Netherlands

10:51

um lots of numbers like this about

10:53

distributed systems and systems

10:55

engineering [clears throat]

10:56

>> and it's been sort of taped and become

10:58

the bible for a lot of distributed

10:59

systems engineers

11:00

>> okay yeah

11:01

>> now fast forward that list is up for an

11:04

update give us the AI edition for now

11:08

2026.

11:09

>> Yeah, I mean I think if you looked at

11:10

what is important in AI systems these

11:13

days, you would want to know things like

11:17

the bandwidth between you know your main

11:21

memory system on your accelerator to the

11:23

onchip memory to the um you know the

11:26

multiplier unit or whatever. You want to

11:28

know how much energy does it take to do

11:31

a single multiplier operation.

11:34

um you know uh what is the interconnect

11:36

bandwidth between chips and how much

11:39

does that uh how how many chips can you

11:42

connect with that bandwidth and then if

11:43

you go beyond that domain like what is

11:47

the fall off in in uh network bandwidth

11:50

when you need to talk to 10,000 strips

11:52

instead of instead of uh 500 or

11:55

something I think these are all really

11:57

important numbers to to learn

11:59

and and really affect how you think

12:02

about solving particular kinds the

12:03

problems.

12:04

>> H [clears throat] and one interesting

12:07

thing that I've heard you talk about is

12:08

that nowadays the unit that you measure

12:12

everything is energy.

12:15

>> Yeah.

12:15

>> You pointed out that doing a calculation

12:18

or math costs about one pico.

12:21

Uh but moving the data and doing data IO

12:24

costs thousand times that.

12:26

>> Yeah. Just bringing it in from HPM on an

12:28

accelerator into the processor so it can

12:31

actually compute on it. Yep.

12:33

That gap kind of quietly decides what

12:36

products are possible and how these

12:38

algorithms in AI are built. So what are

12:42

the kinds of problems that founders keep

12:44

calling model problems but are in fact

12:47

actually energy or data IO problems?

12:51

Yeah, I mean I think the the example you

12:54

raised of a thousandx difference in

12:56

bringing mo moving data versus actually

12:59

computing on it uh in in terms of energy

13:02

is is a pretty significant one and it

13:04

shapes a lot of aspects of what we do in

13:06

machine learning. Um because if you

13:10

didn't have that thousandx difference

13:11

then you know you wouldn't have to do

13:14

batching but you have to do batching of

13:16

you know many examples or maybe many

13:18

tokens at once in order to amortize that

13:21

data movement [clears throat] so that

13:22

you can uh you know not pay a thousandx

13:26

slowdown but pay a 1000x divided by

13:28

batch size uh energy cost. Um and you

13:32

know for for really low latency batching

13:36

is not really very good. Um so I think

13:39

these [clears throat] kinds of things

13:40

and the energy uh behind various

13:43

decisions in the computer hardware we

13:45

use really affects a lot of decisions we

13:48

make in building higher level systems.

13:51

>> A very concrete example is just how

13:53

training models is done. There's this

13:55

whole whole concept of batching the the

13:59

data sets and running epochs. That's

14:01

basically people perhaps may confuse

14:03

that as a model problem, but it's really

14:05

a systems data IO problem, right?

14:08

>> Yeah. Yeah. I mean, you have to assemble

14:10

batches to get better efficiency in your

14:13

hardware. You know, ideally you might do

14:15

batch size one training, but uh you

14:17

know, it's um not as not as good in

14:21

terms of efficiency. So people use re

14:24

pretty large batches these days.

14:26

>> Do you think uh it's possible for uh I

14:28

know you're you're well known for uh

14:30

taking off uh on a long week or weekend

14:33

and coming up with this brilliant

14:34

solution. Is there such things of Jeff

14:36

going and working on it for a couple

14:38

weeks and [clears throat]

14:38

getting batch size equals one training

14:41

done.

14:43

>> Yeah, I've been thinking more about

14:44

inference actually. So I think inference

14:46

is a pretty interesting problem because

14:49

you do want very low latency. You know

14:52

training you don't necessarily need

14:53

incredibly low latency. Um and I think

14:56

there's a lot of room for specializing

14:58

hardware more for inference than we are

15:00

today.

15:02

>> What are some of those interesting

15:03

things that are on inference that you're

15:05

really thinking a lot about?

15:07

>> Um I mean just trying to minimize data

15:09

movement. uh trying to think about

15:12

incredibly

15:14

uh low precision operations

15:17

uh and maybe not supporting lots and

15:19

lots of different kinds of precisions.

15:21

Uh if you feel like you have a a good

15:25

answer for what kinds of precision you

15:27

need, maybe just build that into the

15:29

hardware and and um not much else.

15:33

which I think it brings down to a core

15:36

[clears throat] analogy I heard from

15:38

famous computer scientists that really

15:40

the whole process of u AI is a big

15:44

compression problem because in order to

15:48

have the data to be f fully lossy and

15:50

compress it and then restore it you

15:52

basically need to understand it. Yeah, I

15:54

mean if you truly understand the data,

15:57

you should be able to compress it really

15:58

well

15:59

>> and now transformer architecture is

16:02

basically one of the ways that has

16:04

turned out to work really well.

16:07

>> Yeah. Yeah, I would say

16:08

>> working pretty well so far.

16:09

>> Good work by my colleagues. [laughter]

16:12

>> Yes. Now let's zoom out a bit. Um AI

16:16

progress used to mean just better

16:18

models. You could had more data trainer

16:20

models with bigger parameters. But

16:23

increasingly in the last years or so,

16:26

it's everything around the model. Not

16:27

just the model size and number of

16:29

parameters or more data. It's everything

16:31

around things like retrieval tools,

16:34

memory, agent tools, and it might kind

16:37

of get consolidated into what people

16:39

call uh context engineering, right?

16:42

[clears throat]

16:42

>> Yeah. I mean I think uh the model is

16:45

really only one piece of what you're

16:48

trying to do which is build an overall

16:49

system that can solve really interesting

16:52

problems and that involves you know a

16:55

model that knows how to use various

16:57

tools. It maybe knows how to retrieve

16:59

relevant information, maybe has a, you

17:02

know, a history of other uh information

17:05

that it has retrieved for past problems

17:08

and it can put information into the

17:11

context of the of the model. And the

17:14

nice thing about that is that

17:15

information is really clear to the

17:17

model, unlike the training data the

17:19

model was trained on where it's all kind

17:21

of like trillions of tokens stirred

17:24

together into a soup of of hundreds of

17:26

billions or trillions of parameters, but

17:28

it's all less clear than the actual

17:32

context uh that the model sees directly

17:34

for this particular problem or uses use

17:36

case. And then I think being able to

17:39

understand what tools are available,

17:43

which ones are going to help me solve

17:44

the help the model solve this next you

17:47

know phase of the problem, how to

17:48

decompose a problem into a sequence of

17:50

of tool calls. Maybe trying multiple

17:52

approaches to solve the problem and

17:55

seeing which ones work and being able to

17:56

evaluate that. you know this is the

17:58

whole um you know orchestration of

18:02

complex agent and multi-agent systems

18:04

that I think is going to be more and

18:06

more important and uh super exciting

18:08

times I would say

18:10

>> and I think the fun thing about this

18:11

particular problem domain set is

18:13

actually something that everyone in this

18:15

room can actually do because before to

18:18

train a model you needed incredible

18:19

amount of resources incredible amount of

18:21

access of to GPUs and data but for

18:24

context engineering everyone here could

18:27

do you have you just need the API to

18:29

something like Gemini and then work on

18:31

your own setup for your own retrieval

18:34

your own tool calls and etc etc. So how

18:38

does what are some tips for everyone

18:40

here? How does everyone get better at

18:42

and become exceptional at context

18:45

engineering? Yeah, I mean I think uh

18:49

[clears throat]

18:50

a really good way to do it is to use

18:52

these models and and sort of harnesses

18:55

and tools and so on to try to solve

18:57

problems and then some sometimes you can

19:00

actually see where the models are

19:01

failing. And often you can actually make

19:04

the model work better and succeed at

19:07

that kind of problem by not just

19:10

adjusting the model parameters which is

19:11

hard to do from the outside but from you

19:14

know creating better guidelines for the

19:16

model you know writing skills for the

19:19

model to know how to use different tools

19:21

that would be incredibly useful for

19:22

solving this particular class of

19:24

problem. And I think as you do that, you

19:28

end up on this kind of improving

19:30

self-improving of the setup that you're

19:33

trying to use to to solve things. Uh,

19:36

and you know that that's a really good

19:37

way to get better at understanding what

19:41

what additional information the model

19:43

would want in order to become more

19:45

capable.

19:46

>> Can you give an example of uh some

19:48

context engineering you personally have

19:50

done? um I don't know skills you wrote

19:52

tools that really made a huge different

19:54

in your in your workflow. Yeah, I mean I

19:57

guess uh Sanjay and I were working a few

19:59

weeks ago and we you know we often do

20:03

some amount of like uh performance

20:05

improvement for very low-level libraries

20:08

and we have a microbenchmark library

20:10

we've written at Google where you can

20:12

write microbenchmarks of how how long

20:15

different kinds of operations take or

20:17

how long does it take to populate this

20:18

data structure whatever and sometimes

20:20

those data structures are used on

20:22

millions of processes across Google. So,

20:25

it's actually pretty important to make

20:26

sure they're high performance. And so,

20:28

you can write microbenchmarks. Um, but

20:31

then without an agent-based system, what

20:33

you usually do is you measure what the

20:35

current performance is on some

20:37

benchmarks you care about. You make some

20:39

modifications to improve the performance

20:42

you hope. Then you rerun the the

20:45

benchmarks, see where things improved.

20:48

Um, you run a maybe a broader set of

20:49

benchmarks, measure the cache footprint

20:52

of things. And so we wrote a skill that

20:56

basically taught the model how to do

20:57

most of those things in in var in

21:00

various sequences so that it could

21:01

actually you know do self-improving uh

21:04

benchmark measurement benchmark improve

21:06

you know code changes measure the

21:08

performance improvement and then iterate

21:10

on that and that that seemed to work uh

21:12

pretty well for some kinds of problems.

21:14

And it really just is us giving the

21:18

approach we would use as people to the

21:21

model in a form that it could use.

21:23

>> Wow, that seems very impressive. So

21:25

you're saying you have this skill that

21:27

if someone got access to it, it could do

21:30

perform optimizations like Jeff Dean.

21:33

Seems like the world would love this and

21:36

is worth infinite amount of money to

21:38

someone have access to this.

21:39

>> Oh. Uh we actually published a document

21:41

maybe a few months ago called

21:43

performance hints that Sanjay and I

21:45

wrote that's like a 30-page document

21:47

about you know various kinds of

21:49

performance tricks and some people have

21:51

taken that and then given it in

21:53

summarized form to various models and

21:55

seen that they that model can now get

21:57

you know better at uh per reasoning

21:59

about performance issues in code.

22:01

>> So you heard it all here. You could

22:02

actually get your own optimize your own

22:05

code like Jeff Dean if you take this

22:07

this paper that you published when

22:08

performance hints. Yep.

22:10

>> It's all free available, so you should

22:11

all try it.

22:12

>> Very cool.

22:13

>> Yeah.

22:14

>> Now, you're talking about agents. Um,

22:16

everyone here is probably building one

22:17

or built one at some point. And I'm sure

22:20

everyone has seen your agent go off the

22:23

rail at perhaps like step 30 or 40. Like

22:26

agents are great for like up to step, I

22:28

don't know, 10 or something and then

22:29

gets shaky at step 50. What do you think

22:32

is the constraint today? Is it like

22:34

context

22:36

evaluators or just errors that compound

22:38

because it's basically a openloop

22:40

system?

22:41

>> Yeah, I mean obviously we want agents to

22:44

be able to run for very long periods of

22:46

time because that's how they're going to

22:47

solve more and more complicated

22:49

problems. Um but as you as you observe

22:52

today, you know, they sometimes stop

22:55

working after, you know, 10 10

22:59

interactions with the tools and so on.

23:01

Um, and sometimes that's because the

23:04

model is trying to do something it

23:05

doesn't have a lot of experience doing.

23:07

So it's been trained on a whole set of

23:09

things and as soon as you get a little

23:11

bit off the distribution of things it

23:13

knows how to do then like most machine

23:15

learning models it will you know its

23:18

performance will suddenly will start to

23:20

degrade and the farther you get off the

23:22

comfort zone of what it knows how to do

23:25

the the more likely it is to to not work

23:27

as well. Um so there's a bunch of things

23:30

you can do. So one is you know give the

23:33

model skills and hints that kind of tend

23:35

to keep it in in the uh sort of more

23:38

brightly lit path of things it does know

23:40

how to do. Um, I think you know having

23:43

multi- aent systems where you have

23:45

multiple agents trying different

23:47

approaches and you can evaluate you have

23:49

maybe another model or another agent

23:51

that's evaluating which ones of those

23:53

seem promising is another way to kind of

23:58

in some sense search the path of pos

24:01

search the space of possible solutions

24:03

and stick to the ones that seem most

24:06

promising and discard the ones that that

24:08

didn't seem to work or maybe that went

24:10

off the rails. or whatever. Um, and

24:13

that's a very very useful general

24:15

technique is you know inference time

24:18

compute to perform search over plausible

24:20

ways of solving the problem that can get

24:23

much much higher performance or much

24:25

more reliability in longunning agent

24:27

flows.

24:30

>> How are some ways you implemented this

24:32

particular workflow for your agents

24:34

internally?

24:35

Yeah, I mean we have uh you know

24:37

harnesses and then we have a whole set

24:39

of skills uh particularly in the

24:42

internal Google development environment.

24:43

We have skills so that the agents can

24:45

know how to use lots of our internal

24:48

tooling for coding or for code reviews

24:50

or for you know measuring performance or

24:54

you know fetching log files. And um

24:57

those are just skills that you can add

24:59

to make the base model more capable even

25:01

though it hasn't necessarily been

25:03

trained on exactly the way that you know

25:06

Google internal uh engineers would fetch

25:09

log files from our you know proprietary

25:12

system with the right kind of skill uh

25:15

definition you can actually get it to

25:16

work

25:17

>> uh and that that improves the usefulness

25:19

of the agents. Now let's talk about uh

25:23

where startups can can win. This section

25:27

is one that I personally care a lot

25:29

about because also everyone here in this

25:31

room needs to decide what to build in

25:33

the future of your future founder. So

25:37

the thing about Google is you co-design

25:39

everything on the system from the

25:41

processors to the products.

25:44

um which are the layers that someone

25:47

like Google would keep building and

25:49

compounding being better and and where

25:52

does a two three person team can still

25:55

win?

25:57

>> Yeah, I mean I think obviously Google

26:00

and and our Gemini models and and our

26:02

hardware infrastructure are really

26:04

trying to build very general models that

26:06

can do almost anything. But in in a lot

26:11

of cases that means that we don't have a

26:14

lot of attention on particular domains

26:16

where perhaps a really well-designed

26:19

surface that and maybe a model and set

26:22

of skills or maybe a specialized model

26:25

that uh isn't in sort of a general mix

26:28

of of things that our models do well can

26:31

actually have a significant advantage

26:33

because you can build something

26:34

delightful and you know really high

26:37

accuracy. really high quality for a

26:40

domain that you are really passionate

26:41

about. And I think that's that's where

26:44

you know the two or three people in a

26:46

room uh building that that they're

26:49

really excited about can have an

26:50

advantage. Um but I I would also caution

26:54

that the general models are definitely

26:57

getting better at a broader and broader

26:59

range of things. So you have to figure

27:00

out, you know, is that thing you're

27:03

working on, is that going to be a

27:04

durable thing or do you think the models

27:07

uh at the forefront are going to get

27:09

better at that in the next six months or

27:11

12 months or is it something they're not

27:13

going to be able to do for a couple

27:15

years or three years? And you know, you

27:17

you want to weigh that as you're as

27:19

you're deciding what to work on.

27:22

>> So let's uh dive deeper into this. So

27:24

the general models of course you're

27:26

going to keep working on and keep making

27:27

them all better.

27:29

And how should the audience reason about

27:31

what are those areas that uh it doesn't

27:35

I mean h how should founder think about

27:38

things to pick on and work on.

27:40

>> Yeah. I mean I mean the most important

27:42

thing is to pick something you're super

27:44

excited about and want to build and you

27:46

think would be useful in the world,

27:48

right? So if you do that um that that's

27:51

you're already way ahead uh than if you

27:54

wake up and you're like oh I don't

27:55

really want to do this or whatever or

27:57

you're going to build something that is

27:59

actually not that useful to to the world

28:01

or to to many people. Um so I think

28:05

that's the number one selection criteria

28:07

I try to apply for what problem should I

28:10

work on next. Um, second, I think you

28:14

want to look at what the current more

28:16

general models can do in that problem

28:19

domain, right? You can you can test them

28:22

with like, are they able to do this

28:23

thing very well? And if they're

28:26

completely failing, that's probably a

28:28

good sign. If they're kind of able to do

28:30

some of it but not very well, that's

28:33

maybe not a great sign because that's a

28:35

probably a a sign that the capability is

28:38

starting to be present in those models

28:41

and with more training data or larger

28:43

scale models or or whatever it's likely

28:45

to get better. So um you know look for

28:48

something where the model succeeds 0% or

28:51

1% of the time not not 20%.

28:54

>> How do you find those? I mean are those

28:56

things effectively

28:58

uh out of distribution from the training

29:00

set and what exactly is the problem

29:02

shape that fits that?

29:04

>> Yeah, I mean I think uh sometimes it's

29:08

uh a product that you build that might

29:11

have access to particular kind of data

29:13

that the underlying model might not the

29:15

a general model. So it might be you're

29:18

building something to help users

29:19

organize all their own personal

29:21

information and the model won't

29:22

necessarily have access to that. And so

29:24

there you can have a big advantage

29:26

because all of a sudden your model has

29:28

visibility or your product has

29:31

visibility into important data. Um it

29:35

could be some incredibly hard problem

29:38

where if you get the right training data

29:40

and you can train a more specific model

29:42

than a general purpose one, you can

29:44

actually do that in a very affordable

29:46

way. you maybe it doesn't take that much

29:48

compute to train a a niche model for

29:50

this particular problem, but you can get

29:52

something that's highly accurate. That

29:54

can sometimes be a a really good uh

29:57

building block for for solving a

30:00

important problem that is maybe not

30:03

handled very well by the general model.

30:04

>> I think that's interesting. I think

30:06

there are basically two paths. The first

30:07

path is uh a little bit funny is uh you

30:10

guys are organizing the world's

30:12

information.

30:12

>> Yeah,

30:13

>> that's probably kind of well covered.

30:15

>> Yeah. but organizing your personal

30:17

information that's open which is funny.

30:21

>> Yeah.

30:21

>> And then the second path um you talked

30:24

about more specialized models in certain

30:27

domains. Can you tell us more about what

30:29

are some of these domains?

30:30

>> Yeah. I mean I think like if you look at

30:33

uh my colleagues work on say alpha fold

30:35

that was a very specific model for uh

30:38

protein folding and it was highly

30:40

successful

30:41

um and was able to really handle that

30:44

domain quite well so that all of a

30:46

sudden you now have this amazing tool

30:48

and model that can give you answers to

30:50

questions about proteins and their

30:52

structure um really effectively um but

30:55

it's not a general model it's a very

30:58

specific one and there are other I

31:01

domains where that kind of approach can

31:03

work really well. Uh maybe in material

31:05

science or chip design or things like

31:07

that that uh will enable you to leverage

31:11

the capabilities of a very accurate but

31:14

but niche model uh to do things that are

31:18

hard today.

31:19

>> That's a good example. So if some of you

31:21

find a problem that's similar shape like

31:23

alpha fold could be a good problem to

31:24

work on. Now let's assume you found a

31:26

problem to work on. We're going to talk

31:28

a bit about how do you become a AI

31:30

native founder? How do you really become

31:32

good at it? Uh you in the past said that

31:35

managing a fleet of agents, it's like 50

31:37

or 100 agents is all about writing

31:40

really good crisp design docs or specs.

31:45

>> And how do people get good at that? What

31:47

what do those look like? Yeah, I mean I

31:50

think uh

31:54

it's

31:55

you you'll have a lot more success when

31:58

working with your virtual agents if you

32:00

can clearly specify what it is you want.

32:03

And the clearer you are on what it is

32:04

you want, the more the agent will have

32:07

sort of guidelines and sort of rules of,

32:10

you know, an outline of what it is

32:12

trying to accomplish. Um whereas if you

32:15

don't specify very much stuff, the agent

32:17

has to sort of infer what it is you

32:19

meant. And in many cases, it might infer

32:22

things that are different than what you

32:23

imagined. So we've always told computer

32:26

scientists from the very beginning that

32:27

really it's really important to specify

32:29

what it is, what's the software that

32:31

you're writing is trying to accomplish

32:34

before then going and writing it. And so

32:36

now we actually have agent-based systems

32:38

that can do the writing, but the

32:40

importance of specifying what what it is

32:42

you want has actually gone up because

32:45

before you'd be handing it off to a very

32:48

intelligent human who maybe has context

32:50

or can ask you follow-up questions. Um,

32:53

and agents can sometimes do that, but I

32:54

I think clear specifications is is a

32:58

really good idea. Um, and to give you an

33:00

example of a a a use of a coding agent

33:04

that works extremely well is you can ask

33:07

today's models to translate software

33:10

from one computer language to another

33:12

very effectively because in that case

33:15

you actually have a incredibly detailed

33:18

specification. You have the whole

33:19

software that says what the system is

33:21

supposed to do. And so if you have a

33:23

Python implementation of something and

33:25

you want a Go implementation of it, you

33:28

know, that is something that the models

33:29

seem incredibly capable at doing these

33:31

days because you can it can sort of take

33:34

all the tests that are in Python, make

33:36

sure they pass in the Go version,

33:38

translate the tests to Go, um you know,

33:41

compare uh behavioral differences

33:43

between the implementations until there

33:45

aren't any um and be you know, highly

33:48

effective because that spec is so clear.

33:51

Hm. Now let's assume now every founder

33:53

gets good at running hundreds of agents

33:55

at the same time and all the code is

33:57

written for them by the agents. What

33:59

becomes the scarce skill?

34:03

>> Yeah, I mean I think it's really having

34:07

incredibly good taste in what you ask

34:09

your agents to work on, right? That is

34:11

the the crux of you know from my

34:15

background uh a research problem. You

34:17

know, a researcher can have all the

34:19

tools and all the techniques, but often

34:22

most of the battle is what problem are

34:25

you gonna spend your time on? And if you

34:27

pick the problem well and you succeed in

34:30

in in solving it, that's way better than

34:33

if you, you know, uh, delightfully

34:36

execute a research investigation into a

34:39

rather boring problem. And so that high

34:42

level wisdom of what to work on, I think

34:44

is incredibly important. And I think

34:46

models are not necessarily going to be

34:47

that good at it. So you're going to have

34:50

people steering

34:54

uh a lot of AI assisted computation in

34:57

order to accomplish great things and

34:59

more quickly. Um but that essence of of

35:02

what it is you want your models to do is

35:05

the the the key thing you should focus

35:07

on.

35:08

>> So let's talk a bit more about taste

35:09

because it gets talked a lot about right

35:12

now in this current era with agent

35:14

coding. How do you exactly build taste

35:17

and do that? I mean, yeah, that sounds

35:20

so esoteric. How do you make it

35:22

concrete?

35:23

>> Yeah, I mean it it is a difficult thing.

35:26

It's not like there's a measurable

35:28

objective of of taste in a lot of cases.

35:31

Um, I think some of it is from

35:33

experience. You know, working on a lot

35:36

of different problems in the past kind

35:38

of teaches you about what kinds of

35:40

problems might be interesting in the

35:42

future or what kinds of things might be

35:45

just barely possible by cobbling

35:48

together these previous approaches and

35:51

then some open problems you might have

35:53

to work on in order to get to something

35:55

kind of magical or or you know, highly

35:57

useful. Um,

36:00

another way you can get more experience

36:02

for yourself is to just write down a

36:06

bunch of things you think might be

36:07

important in the next 12 months. And

36:10

maybe you pick one of them to work on,

36:12

but go back and evaluate in 12 months of

36:15

these other things, which ones actually

36:18

seemed important or which ones did other

36:20

people in the world go out and and

36:22

create and which ones did they did not

36:24

seem to do yet. um that can give you a

36:27

lot more samples for your own sort of

36:29

taste creation uh capability. Um and and

36:33

that's an important skill to have.

36:35

>> I think a third way we were talking

36:37

earlier was doing very crazy thought

36:41

experiments.

36:42

>> Oh yeah, that's another good way. I mean

36:44

I think uh sometimes it's good to

36:49

not take as a given things that most

36:52

people seem to take as a as a given. Um

36:55

so I was doing a crazy thought

36:56

experiment with some colleagues the

36:58

other day about you know

37:02

for 60 years the whole silicon

37:06

uh chip design industry uh design and

37:09

fabrication industry have

37:12

you know done tremendous work to make

37:15

smaller and smaller scale transistors

37:17

that are uh very low error rate right

37:22

like because what what the assumption

37:24

that we want is that every chip we

37:25

manufacture of the same design should be

37:28

identical to every other chip.

37:30

>> You don't want any bits to flip.

37:31

Everything

37:31

>> no bits should flip. There's all kinds

37:33

of things you there's all kinds of error

37:35

margins built into you know memories

37:38

have ECC memory these days. um you know

37:41

at the at the macro scale we don't make

37:44

that assumption when we're building

37:46

large scale distributed systems right we

37:49

we build reliable large scale

37:51

distributed file systems out of

37:53

unreliable parts right like individual

37:56

discs can fail but your data should be

37:58

safe and so we have mechanisms at a

38:00

higher level to enable us to have um you

38:04

know three copies of the data on three

38:06

different machines and three different

38:07

racks so that if any rack switch or

38:09

individual ual machine or disk fails,

38:11

you still have your data. We have read

38:13

Solomon encoding techniques. Um, but we

38:17

don't seem to do this at a really

38:19

extreme level in the uh sort of

38:23

transistor level scale of of the

38:26

technology we're working on. So what

38:28

would h basically a interesting thought

38:30

experiment is what would happen if you

38:32

tried to build a system out of

38:34

transistors that might have you know 20

38:37

errors per day.

38:38

>> Oh my god. rather than one every million

38:40

years, right? That would be a very

38:43

different design point and might be

38:45

might enable you to do really

38:46

interesting things in the fabrication

38:48

side of things. You have very different

38:50

kind of design methodologies because if

38:52

you want to get a signal from here to

38:54

there, you and you have these super

38:56

unreliable transistors. You might have

38:58

very different ways of signaling. You

39:00

might send it along multiple redundant

39:02

paths uh in order to make sure that it

39:04

gets along one of them. Um, and I think

39:07

that would be a pretty interesting set

39:08

of thought experiments. I'm not saying

39:10

we should go do this, but you know,

39:12

that's the kind of thing where you do

39:13

want to, you know, occasionally question

39:17

assumptions. Now, oftentimes these

39:20

thought experiments don't work out

39:22

because there are very good reasons

39:24

that, you know, for the last 50 years,

39:26

we've done this thing this way and not

39:28

that way. But it it's good to kind of

39:30

revisit those every so often.

39:32

>> That is so wild. Well, I mean, it's

39:33

starting to rhyme a lot with

39:35

neuromorphic computing or the human

39:37

brain and and how nature works.

39:39

>> I mean, exactly like signals in our

39:41

brain are not especially reliable from

39:43

getting one place to another. And so, I

39:46

think in brains when there are really

39:48

important things you need to get from

39:50

one place to another, there are multiple

39:51

pathways that that enable you to sort of

39:54

do that.

39:56

>> What is uh in I mean, you have such an

39:58

impressive career. What is one of these

40:01

crazy assumptions that you threw out of

40:02

the window that actually built a

40:04

consequential system in the past?

40:09

>> Yeah, I mean I guess uh

40:10

>> that worked out actually.

40:13

>> Yeah, I mean I think uh well TPUs is a

40:15

good example like being able to

40:17

specialize hardware for a very niche

40:20

>> problem domain before that problem

40:22

domain seemed as important as it is

40:24

today uh is one thought experiment. Um

40:28

you know I think the

40:30

the origin of map produce is another

40:33

good example.

40:34

So we had worked the you know my Sanjay

40:37

and myself and a number of other

40:39

colleagues had worked on various

40:41

iterations of the crawling and indexing

40:43

system at Google and you know we'd sort

40:46

of written lots of hand parallelized

40:49

code with lots of checkpointing to make

40:51

sure it would be robust and reliable if

40:53

it was running on a 100 computers or a

40:55

thousand computers and some of those

40:57

died. Um, but that code tended to be

41:02

intermixed with the actually relatively

41:04

simple thing you often were trying to do

41:06

like I just want to like look at all the

41:09

contents of all the web pages and then

41:11

compute on the side a mapping from URL

41:14

to you know what language is this page

41:16

in the text of this page. Um, and it

41:20

would get obscured by all this kind of

41:22

other code for parallelization and

41:24

reliability. And so we sort of

41:27

remembered our training in functional

41:29

languages and realized we could squint

41:31

at those problems and developed this map

41:33

produce abstraction that you could have

41:36

above this implementation and then below

41:39

the implementation you could put all the

41:42

checkpointing and reliability mechanisms

41:45

into that lower level library that

41:47

everything could then build on. And so

41:49

that became a hugely successful way of

41:52

of dealing with very large scale

41:53

computations at Google in a robust and

41:55

reliable way. From that thought

41:58

experiment of like well if we squint at

42:00

it could we find lots of problems that

42:01

fit into this abstraction.

42:03

>> That's impressive. So this thought

42:06

experiment led you to create map reduce.

42:08

>> Yeah. Awesome.

42:09

>> Now let's go back to you talked a bit

42:11

about um about your interest right now

42:15

working on a lot of customized hardware.

42:17

So right now alpha chip

42:19

>> lays out chips. Now you also got alpha

42:21

evolve that proposes solutions,

42:24

>> evaluates them and keeps all the ones

42:26

that work. Seems like you're starting to

42:27

build all these system that can compound

42:30

and build AI that builds AI.

42:32

>> Yeah. I mean I think more generally

42:34

there's a there's this sort of

42:39

the foundation of the scientific method

42:41

of you propose an experiment you

42:44

implement what you need to run the

42:46

experiment and you evaluate the

42:48

experiment and then you get results from

42:50

that and I think there are more and more

42:52

problems that are now possible to

42:55

implement where that whole loop of

42:58

running you know not just a few

43:00

experiments but running many many

43:01

experiments because you're able to

43:03

automate that loop and make the latency

43:05

of that loop extremely low is going to

43:08

be really really important. It's going

43:09

to enable us to tackle you know lots of

43:12

different problem domains in science and

43:14

engineering and machine learning uh

43:17

model design itself and also in

43:20

engineering tasks like designing chips.

43:23

And so if you can actually do those

43:25

things in an automated way and have some

43:28

orchestration framework that can take

43:31

very high level objectives and break

43:33

them down into subpros and each of those

43:35

subpros can be one of these automated

43:38

loop that is exploring the best way to

43:40

solve that sub problem and then a

43:43

orchestration framework that can put

43:46

together subpros solutions into a you

43:50

know the overall solution for the higher

43:52

level problem that's going to be really

43:54

impactful and it's really really

43:56

important and I think it'll enable us to

43:58

do you know accelerate machine learning

44:01

progress it'll enable us to accelerate

44:03

science and enable us to accelerate

44:05

engineering and I think that's that's

44:07

going to be amazing

44:09

>> that sounds awesome I mean it sounds

44:10

like a lot of fields basically where you

44:12

can have very good evaluators and maybe

44:15

adjacent to basically things that can be

44:17

formally verified right those are ripe

44:20

for AI systems that can self-improve

44:22

Yeah, I think in a lot of cases

44:25

sometimes your evaluators need to be

44:27

made much faster. Mhm.

44:28

>> So as an example, my colleagues did some

44:32

work maybe a decade ago on um some uh

44:36

problems in quantum chemistry where

44:38

you're trying to understand the

44:39

properties of a particular molecule and

44:41

you can you know generate some molecule

44:44

configuration and then you want to

44:45

understand what properties it has. And

44:48

so you can run a very computationally

44:50

intensive density functional theory

44:52

simulator which is something that might

44:54

take like a a night of computation to

44:57

tell you the answer for one thing. Um

45:00

but what my colleagues did was

45:04

take a bunch of output from those

45:07

simulation runs the input molecule

45:09

configurations and the outputs of the

45:11

the expensive simulator and then use it

45:14

to train a neural approximation to the

45:16

simulator. So this is now a validation

45:19

device, but instead of it taking a

45:22

night, they made something that was

45:25

300,000 times faster.

45:26

>> Wow.

45:27

>> And nearly as accurate as running the

45:28

full scale simulator. So now that

45:32

completely changes how you would do

45:33

science, right? Because now you have 10

45:35

million things to screen. you know, you

45:38

could do that while you go to lunch

45:40

rather than it being a six-month

45:42

endeavor where you could try to scrape

45:44

together enough compute to to run all

45:46

these simulations. And I think there's a

45:50

lot of room in a lot of domains for much

45:53

faster validation models, possibly

45:55

learned valu validation models that can

45:59

uh you know get you a a approximation to

46:02

the true answer much much more rapidly.

46:04

And that changes how those experimental

46:06

loops can be thought of and how quickly

46:08

you can go around those loops.

46:10

>> What are some of the

46:13

spaces and problems that you're super

46:14

excited that this super sped up

46:17

scientific method is going to solve or

46:20

achieve? What particular problems or

46:22

spaces?

46:23

>> Yeah, I mean I think uh

46:28

well clearly machine learning itself is

46:30

one, right? So can we have a model that

46:32

is able to recursively self-improve

46:35

itself by running lots of experiments

46:37

and you know if you think about how

46:39

models are improved today in large

46:41

research teams you know what usually

46:44

happens is people think of some ideas

46:46

they run a bunch of smallcale

46:48

experiments they see if those small

46:50

scale experiments worked out well if so

46:53

they take the most promising ones of

46:54

those they try them at larger scale and

46:56

that gets then evaluated and then the

47:00

results get integr ated together into

47:02

you know a new recipe for your model. Um

47:06

but I think there's no uh you know real

47:09

impediment to making that be a much more

47:11

automated loop where the model itself

47:15

decides it's going to explore or maybe

47:17

with a nudge from some people uh at the

47:20

various highest level like oh why don't

47:22

you try some new ideas around model

47:24

architectures that incorporate this and

47:27

then it will go run lots of experiments

47:30

uh see which ones work and then those

47:31

will get incorporated at a much more

47:33

rapid rate and uh you know effectively

47:36

you want to optimize you know your

47:39

discoveries per unit of compute input.

47:44

>> Very cool.

47:44

>> Yeah.

47:45

>> Now going back to the room as all of you

47:49

will become at some point founders or

47:51

start your careers you will probably

47:54

collect lots of rejections. That will

47:56

happen. Uh it has happened to you too

47:59

Jeff. I mean there's a story that in

48:02

2014

48:04

you with Jeff Hinton and Oral Fin wrote

48:09

a paper on distillation

48:12

>> which has to do with taking a big

48:15

teacher model to train a much smaller

48:18

and more efficient model that's a lot

48:21

cheaper to compute less model parameters

48:24

and it has become a trick that everyone

48:27

is using right now in industry. Yeah.

48:30

>> And

48:32

the thing is this paper got rejected at

48:34

Europe.

48:36

>> Yeah. I mean Yeah. I mean I think I

48:39

don't fault the program committee

48:40

because you know a lot of times a paper

48:44

gets three reviews and someone will look

48:46

at one of the reviewers will look at it

48:48

and in this case they said oh it's

48:50

unlikely to have significant impact.

48:52

>> Unlikely to have significant impact. But

48:53

you know I think you know when we wrote

48:55

the paper we actually saw this was a

48:57

super important problem because we knew

49:00

making cheaper highly capable models

49:03

from larger scale models was something

49:04

we desperately wanted to do because we

49:07

wanted to serve models to more and more

49:09

people in many different domains like

49:11

speech or vision. Um but you know

49:13

sometimes the reviewer maybe didn't have

49:15

that that experience because maybe

49:16

they're not thinking about you know

49:19

largecale AI services and are thinking

49:22

about you know is this a fundamental

49:24

advance um so so you know it gets

49:27

rejected every so often that's fine we

49:29

put it on archive people read it people

49:31

use it it's all good uh and you know we

49:34

do use it in making our flash models for

49:37

example from our larger scale pro model

49:40

that's partly why our flash models for

49:41

example in Gemini are so capable uh

49:45

relative to their size and and speed.

49:47

>> They're some of the best in the

49:48

benchmark for their model size class.

49:50

Yeah. Just impressive. And I think part

49:52

of the lesson is that even if you get

49:54

rejected, keep going.

49:56

>> Yeah. That's that's the lesson I would

49:58

distill from that.

50:00

[laughter]

50:01

>> Um no, I think the fun thing is that you

50:04

basically join when you when you join

50:05

Google as a 20 person startup back in

50:08

1999.

50:09

Now, if you were to take the young Jeff

50:12

Dean from way back then to teleransport

50:16

him to now today.

50:18

>> Yeah.

50:18

>> In this era with your skills.

50:20

>> I'm feeling so vigorous and and young

50:23

now.

50:24

>> Um what would you do? Do you join a

50:26

frontier lab, start a company? I don't

50:29

know what what would you do? The c the

50:34

Jeff theme today 25-year-old Jeff theme.

50:36

>> Yeah. I mean,

50:39

it's always hard to say and it's a very

50:41

personal choice of what it is you want

50:42

to spend your time on. Um, to me, some

50:46

of the most important questions are,

50:49

are you going to work on something you

50:52

really care about, will you're working

50:55

on that? And if you're able to make

50:57

progress on it with a bunch of

50:59

colleagues you like working with uh if

51:02

you're able to make collectively solve

51:04

it or make progress on it, will that

51:07

make a difference in the world in some

51:08

positive way, right? Like will you

51:10

suddenly be able to do something and

51:12

offer that service to you know

51:15

partically

51:19

help biochemists or something or maybe

51:21

it's a broader thing. It'll help

51:22

programmers or it will help all

51:24

consumers. uh on the internet or or

51:27

other things. Um what you you know what

51:32

you should strive to do is to have

51:33

impact in the world that is positive and

51:36

to work with people you enjoy working

51:38

with and to you know uh work hard and

51:42

and do your best. Um so in terms of say

51:46

the particular trade-off you offered

51:48

joining a frontier lab versus say

51:50

starting a company with just one or two

51:52

or three of you you and your close

51:54

friends. Um, I think those are different

51:57

experiences, right? In a in a large

51:59

established organization, you have some

52:02

structure. You have lots and lots of

52:04

amazing colleagues who know lots of

52:06

things you don't. Um, you have lots of

52:10

interesting problems that uh you can

52:12

work on and h you already have a

52:15

platform for impact by your work, you

52:18

know, influencing lots and lots of

52:20

people in the world already. Um and then

52:22

as a very small startup,

52:26

you know, you have to have something

52:28

you're passionate about and there's a

52:31

lot of risk in taking on, you know,

52:34

working on that particular problem in a

52:36

way that uh you're going to succeed and

52:38

you're going to grow a, you know, an

52:40

endeavor in order to do that. But that

52:42

can also be incredibly rewarding, I

52:44

would imagine. So I I think um you know

52:48

it's really up to personal taste but but

52:50

at the very least regardless of what

52:53

path you take ask yourself if I work on

52:56

this problem and the best possible

52:58

outcome happens you know will the world

53:01

be a lot better in some way or will the

53:03

world go eh that's kind of cool but

53:05

whatever.

53:06

>> Uh that's not the kind of thing you

53:08

should spend your time on.

53:10

Now let's talk a bit about more about

53:12

that second path of working with people

53:15

that you really like in a small team.

53:18

You've been able to be an incredible

53:20

mentor and manager to many many

53:22

engineers and you've been able to build

53:25

huge systems and what are some some of

53:29

the lessons for everyone here on how to

53:32

get the most and how to work with smart

53:34

people or find smart people?

53:36

Yeah, I mean,

53:39

you always want to find people who have

53:42

really good skills in some some area

53:45

that's needed in, you know, a team

53:47

you're trying to form, whether that's

53:50

inside a company or uh starting a

53:52

company. Um, but you also want to find

53:55

people that are people you delight being

53:59

around, right? because you're going to

54:00

spend a lot of time around people

54:02

working on really hard problems and you

54:05

want people who are low ego that are

54:08

team players that you know have

54:10

complimentary skills to your own perhaps

54:13

um I always find working in a small team

54:16

where people know things that I don't

54:18

know and where maybe I have some skills

54:20

that other people don't have as much of

54:22

you know is super fun because you're

54:24

collectively building something or

54:26

working on something that none of you

54:28

could maybe do individually. ually, but

54:30

in the process of working on that, you

54:34

actually gain a lot of new knowledge and

54:35

new skills uh for yourself and so do

54:38

they. And you you kind of want to view

54:41

your engineering or research career as

54:44

you have an amazing tool belt of

54:46

techniques. And you always want to be

54:48

adding new tools to that tool belt

54:50

because you never know when you might

54:53

come across a problem where you need

54:55

these four specialized tools rather than

54:57

these three. And adding more tools makes

55:00

it more likely that the problems you you

55:02

encounter in the future will be solvable

55:04

by you.

55:07

>> Now, one last thing. I'm pretty sure

55:09

someone in this room or multiple people

55:13

will eventually build something as

55:15

consequential as you've done with map

55:18

reduce, TPU,

55:21

distillation, etc., etc. What problem do

55:24

you hope they would be working on? Oh

55:28

yeah. I mean I I think there's a lot of

55:30

interesting problems in the world and

55:32

I'll just rattle off a few. This is not

55:34

exhaustive because the world is a very

55:36

big place and full of problems. You know

55:39

I'm particularly excited about new

55:41

approaches to hardware. You know we that

55:43

thought experiment there was kind of you

55:45

know a

55:47

you know a indication of that or much

55:50

more efficient inference hardware. You

55:52

know, I think there are radically

55:54

different kinds of algorithms for

55:56

machine learning that might be much much

55:58

more data efficient than the approaches

56:00

we're using today. If you think about

56:02

our large scale models today, they

56:04

probably see a thousand times as much

56:05

data as a human does by the age of 18.

56:09

Yet, the human by the age of 18 is

56:11

better in a lot of things and, you know,

56:13

on par uh with those frontier models

56:16

that have seen way more data. So could

56:17

you come up with much more data

56:20

efficient systems that can learn

56:22

continuously learn from their own

56:24

actions? Uh continual learning is a

56:26

really interesting thing. I think multi-

56:28

aent interactions is an interesting

56:30

thing. Um you know I think you know

56:35

creating ways of having better discourse

56:37

among people in the world uh could be

56:39

interesting. Are there ways to have much

56:41

more civil conversations and you know

56:44

helping people meet other people are all

56:46

over the world that they should know

56:48

based on their interests. You know these

56:50

are kind of interesting things. I think

56:52

there there's lots of cool things in the

56:54

world and we should all go and strive to

56:56

make even cooler things occur.

56:59

>> That sounds wonderful. Thank you so much

57:01

Jeff Dane. That's all we have today.

57:03

>> Appreciate it.

57:05

>> Thank you all.

Interactive Summary

This conversation features Jeff Dean, a veteran engineer and key architect behind major systems at Google, discussing the state of AI, engineering principles, and his philosophy on building impactful technology. He touches on the evolution of AI agents, the importance of inference hardware, the concept of 'napkin math,' and the future potential of automated scientific discovery through AI.

Suggested questions

5 ready-made prompts