HomeVideos

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Now Playing

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Transcript

1678 segments

0:00

People think I'm have a radical point of

0:02

view sometimes. They say they start

0:04

questions saying how what I'm thinking

0:06

is so different from everyone else. But

0:07

I don't see it that way at all. I see it

0:10

as like I'm thinking the ordinary way.

0:11

It's just everyone else that's thinking

0:13

a bit weird. And

0:15

>> [laughter]

0:17

>> And I mean that like, you know, it's

0:18

just the recent times people are

0:20

thinking weird. Before there was all

0:22

this AI craziness,

0:24

uh

0:25

you talk about you wouldn't have to say

0:26

continual learning cuz

0:28

it wouldn't make any sense to talk about

0:29

learning that wasn't continual. All

0:31

learning is continual. We always act and

0:33

we learn. That's just the normal way of

0:35

thinking. I'm not weird.

0:37

The field is weird. The field they need

0:39

to call it continual learning. It's just

0:41

learning.

0:50

>> [music]

0:59

>> We are honored to have the great Rich

1:01

Sutton with us here today. Rich, you

1:03

invented reinforcement learning. You

1:05

wrote the seminal textbook.

1:08

You're the

1:09

key students in the field, uh folks like

1:11

Dave Silver. You wrote the essay The

1:13

Bitter Lesson that I believe is the

1:15

Bible of the field. And and you have

1:17

just been one of the greats in

1:18

propelling the field forward. So thank

1:19

you for taking the time to join us

1:21

today. Um Rich is joined by Quorum

1:23

Javed, his co-founder uh and former

1:26

students from the University of Alberta.

1:28

Um the two of you have set off to found

1:30

Oak Lab. I'm very excited to talk to you

1:32

about that today. So for today's

1:34

session, we're going to start talking

1:36

about The Bitter Lesson, the state of

1:37

the world as we know it today,

1:39

whether LLMs will get us there or not.

1:41

And then we're going to we're going to

1:42

transition to start talking about your

1:44

your research agenda and your plan for

1:46

Oak. Um Rich, maybe take us back. I was

1:49

going to start with The Bitter Lesson,

1:50

but I actually want to start earlier

1:52

than that. Decades ago, you decided to

1:56

dedicate your career to reinforcement

1:58

learning, to deep reinforcement learning

1:59

in particular, and you established the

2:01

University of Alberta as a bastion of

2:04

that

2:05

back when I think the field was very

2:07

much in its infancy. What gave you the

2:08

conviction to do that?

2:10

>> What else you going to do?

2:12

>> [laughter]

2:12

>> We were trying to figure out the mind

2:14

and learning is a central part of the

2:16

mind.

2:17

And having a goal is a central part of

2:19

the mind. Central part of intelligence.

2:22

Yeah, so I was just doubling down on

2:24

what I was always thinking.

2:26

>> Did people think you were crazy at the

2:27

time?

2:29

>> Um

2:31

It was a winter. It was an AI winter.

2:33

>> Uh what year was this?

2:35

>> It was in 2003.

2:37

>> Okay.

2:38

>> And

2:39

it's kind of crazy actually the truth

2:42

cuz I was like really sick.

2:43

I was dying I was actually dying of

2:45

cancer in 2003. And but I I wasn't quite

2:50

dead, you know, I've been trying for a

2:52

number of years. And I wasn't dead after

2:55

another remission. And so

2:58

so I said, well, I'm not dying I haven't

3:01

succeeded in dying. So I might as well,

3:04

you know, it's going on long enough

3:05

might as well just try to get another

3:06

job. And so I so I went to Alberta and

3:10

and and started teaching there.

3:12

And then in the end I didn't die. It's

3:13

kind of amazing it's cuz

3:15

you know, it's it's like that. I'm

3:16

joking about it now but it was quite

3:18

serious. And um

3:20

it's an even more important question.

3:22

Why why did I continue

3:26

to work on this research stuff when I

3:27

was, you know, I only had a few months.

3:30

I would always

3:32

keep reminded what

3:33

I think it's Benjamin Franklin is

3:35

supposed to have said that, you know, if

3:36

you ever wonder why someone is doing

3:38

something

3:39

it's almost always one of two things.

3:41

It's either habit

3:43

or vanity. Okay? So I think I think it's

3:46

probably true. Maybe it was my habit to

3:48

just kept doing what I'd always been was

3:50

doing or maybe it was vanity.

3:52

I I know. I think it was more like habit

3:53

cuz I was I was dying.

3:55

>> Wow.

3:56

>> Um

3:56

>> Wow. Divine intervention.

3:57

>> Yeah, it's always been easy for me to

3:59

keep be very determined.

4:02

Um and

4:04

I'm I'm I'm going to go even longer on

4:05

this answer.

4:06

>> Please go.

4:08

>> People think I am have a radical point

4:10

of view sometimes. They they they start

4:12

questions saying how how what I'm

4:14

thinking is so different from everyone

4:16

else. But I don't see it that way at

4:18

all. I see it as like I'm thinking the

4:20

ordinary way. It's everyone else that's

4:22

thinking a bit weird.

4:24

>> [laughter]

4:26

>> And I mean that like, you know, it's

4:28

just the recent times people are

4:30

thinking weird. If you go look back what

4:31

what what what people thought about the

4:33

mind for, you know, even just a decade,

4:36

you'll find the kind of thoughts that,

4:38

you know, learning is important. You've

4:40

got to have a goal. Um

4:43

and you know, perception is important.

4:45

We have a We are We are low-level We are

4:47

low-level beings. We are generating

4:49

actions and perceiving data at a fast

4:52

speed and yet we have to think at higher

4:54

levels. And you know, go back a few

4:57

before there was all this AI craziness,

5:00

uh

5:01

you talk about you wouldn't have to say

5:02

continual learning cuz

5:04

it wouldn't make any sense to talk about

5:05

learning that wasn't continual. All

5:07

learning is continual.

5:08

You know, you

5:10

It's not a special phase. We always act

5:13

and we learn. That's just the normal way

5:14

of thinking. I'm not weird.

5:17

The field is weird. The field they need

5:19

to call it continual learning. It's just

5:21

learning.

5:22

>> I'm not weird, everybody else is. That's

5:23

a

5:24

good good motto to live by.

5:26

>> We're going to have to send out a an ex

5:28

post about that.

5:29

>> [laughter]

5:30

>> We're very happy that you you you lived

5:32

on with the field is happy that you

5:34

lived on and thank you for pushing the

5:36

frontier of AI.

5:36

>> happy.

5:37

>> [laughter]

5:38

>> Thank you for pushing the frontier of

5:39

AI. I'm sure really happy and thank you

5:41

for all of that and push you've like

5:43

been able to sort of educate a lot of

5:45

students who pushed the frontier as

5:46

well.

5:47

How did you pick them? How did you

5:49

over the last 20, 30 years?

5:51

>> Oh, well, you are giving me

5:52

opportunities to be uh

5:54

to be humble. Like I like to be humble

5:56

and point out how all these great

5:58

decisions are are just happen.

6:01

And it's that that's the way I feel

6:03

about students. I don't feel that I

6:05

choose them very well. I've just

6:07

Sometimes I'm lucky, sometimes I'm

6:08

unlucky.

6:10

I don't feel I'm particularly good at

6:11

picking my students.

6:14

I'm looking at Karam. I think sometimes

6:16

you end up with the really great ones.

6:18

David picked David Silver picked me.

6:20

>> Yeah.

6:20

>> How is it that

6:21

that I got you, Karam?

6:23

>> Yeah, that was also so I

6:25

finished my master's not with you

6:27

uh and I was planning to join industry.

6:30

And then we were collaborating on a

6:33

project which also just had started

6:35

organically. Like there was something I

6:37

worked on that Rich was in a meeting,

6:39

then they mentioned that I worked on it.

6:40

So I got pulled into it. We started

6:42

collaborating. It went really well. Like

6:44

I felt so happy with that collaboration.

6:47

Rich also felt really good about it. And

6:49

then 6 months down the road we had made

6:51

some progress and it just made sense to

6:53

convert that into a thesis proposal. So

6:56

at no point did I apply, at no point did

6:58

I ask should you be my PhD advisor. We

7:01

worked together, then we decided this

7:03

would be a pretty good thesis.

7:05

And then then after that I applied for

7:07

the PhD.

7:08

>> Life works in unexpected ways.

7:10

Uh take us to 2019. You wrote the bitter

7:13

lesson which has become the the mother

7:16

in tome.

7:17

2019 was a funny time to be writing that

7:19

piece because ImageNet was 2009. AlphaGo

7:23

was 2015. What caused you in 2019 to

7:26

reflect and and to write that? Because

7:28

it was before the current kind of

7:30

scaling paradigm around large language

7:32

models had taken off, but it was after

7:34

deep learning had really proven itself.

7:36

>> Well, it was a long time coming. You

7:38

know, as the

7:39

bitter lesson expresses, it's something

7:41

that you can

7:42

for a long time, for many decades.

7:45

And it's definitely at least as much due

7:48

to the round of

7:50

symbolic AI, which I lived through.

7:53

It's all about not getting distracted by

7:56

trying to put in your human knowledge

7:58

and just paying attention to what the

7:59

problem needs and how you can scale with

8:02

computation.

8:03

I know I I I wrote versions of it at

8:06

least a year before and I I gave talks.

8:10

I gave a talk a year before.

8:12

And um

8:13

it wasn't a particular response to the

8:16

moment. It was a particular response

8:19

to my my long experience, different

8:22

people trying to think in different ways

8:25

about how you can make smart systems.

8:28

>> Mhm.

8:29

What is the essence of the bitter

8:31

lesson?

8:32

>> The essence of the bitter lesson.

8:33

>> You know, and maybe the phrase that I

8:35

hear used the most in my meetings these

8:37

days is is bitter lesson pill, is it not

8:39

bitter lesson pill?

8:41

I would imagine given the popularity of

8:42

the phrase it's probably been tortured

8:44

and misused in different ways that you

8:46

didn't originally intend it. So, what do

8:48

you think what is the essence of it and

8:49

where do you think people go wrong in

8:51

their in their attempt to understand it?

8:53

>> Yeah, you're you're making me think

8:55

about X now and my I recently made a

8:58

post where I tried to do the the bitter

9:00

lesson in in 26 words.

9:03

It [laughter] goes something like

9:05

don't be distracted by human knowledge

9:09

as AI traditionally has been many times.

9:13

Instead focus on

9:15

learning methods that will scale with

9:18

computation

9:20

like search and like learning. So, it's

9:23

really all about uh focusing on

9:27

algorithms and improvements. It's not

9:30

it's not saying you don't need

9:32

fancy algorithms. You need fancy

9:34

algorithms, but you want fancy

9:36

algorithms that will scale with scale

9:38

with computation.

9:40

>> Rather than scaling with data.

9:41

>> Rather than scaling with human input.

9:45

Yeah, and then the the question if I can

9:47

anticipate it

9:49

um

9:51

Yeah, what about large language models?

9:52

>> Yeah.

9:53

>> Are they

9:54

>> consistent or inconsistent with your

9:56

>> Yeah.

9:57

>> And and I thought about this and I think

10:00

there's an ex- there's another exposed

10:01

model, but the conclusion is that it's

10:04

both a a positive example and a negative

10:06

example of the big lesson. First uh

10:10

large language models enabled uh

10:13

enormous scaling with computation. And

10:16

you could just

10:17

drink in the internet and scale so much.

10:20

So it was a it was a way of getting uh

10:22

much more capable system just by

10:25

methods that scale. Then after that, as

10:28

you go on further um it eventually gets

10:31

limited by by that information. The the

10:34

internet is finite and it's hard to get

10:38

more examples.

10:39

And uh the world is big and the world is

10:43

massively bigger than everything we

10:44

stored

10:45

on the internet. And so in the end, it

10:48

seems like it could be

10:50

uh I guess that would be a

10:52

positive example of when, you know, we

10:54

relied too much on human knowledge and

10:56

it eventually

10:58

holds us back.

10:59

>> Mhm.

11:00

Can I just push on this a little bit?

11:02

>> Yeah.

11:03

>> It seems like a lot of what the

11:04

foundation model labs are working on

11:06

right now is synthetic data generation

11:09

in order to kind of get us beyond

11:12

the fossil fuel that is the existing

11:14

human internet.

11:15

Um is synthetic data generation kind of

11:18

as part of this LLM scaling paradigm, is

11:20

that

11:21

a general method that leverages

11:23

computation?

11:24

>> No, that's that's just a big mistake.

11:27

>> Why?

11:27

>> [laughter]

11:28

>> Well, it's such a big it's such a maybe

11:30

it's the next the next big lesson.

11:33

Um

11:34

it's been floating around

11:37

uh Alberta for 5 or 10 years.

11:40

>> Okay.

11:40

>> And uh we call it the big world

11:42

perspective or big world hypothesis.

11:44

Khuram, who eventually wrote it up as a

11:46

paper. There's a little paper called the

11:49

big world hypothesis.

11:50

>> So, the big world is

11:52

that the world is infinitely big. There

11:55

are infinitely many things to learn.

11:57

And you can have people generating the

12:00

synthetic data sets, but there will

12:02

always be more things to learn. And

12:04

because of that, if you could just

12:06

learn from experience, if you could

12:08

remove the humans from the loop, then

12:10

you would have systems that can do

12:12

everything. Because, you know, the world

12:14

is big. There are many tasks that we

12:15

want them to do.

12:17

And they would be able to do anything by

12:18

learning from their experience. Going

12:20

back to the synthetic data question,

12:21

too. Uh

12:23

who decides what's a good synthetic data

12:25

and what's a bad synthetic data? Because

12:27

I can write a program that can output a

12:28

lot of synthetic data, which would hurt

12:30

programs. Right now, I would say humans

12:32

decide. And that's the bottleneck where

12:35

okay, you can have humans deciding how

12:37

to generate these data sets, but you

12:39

need human experts who know what's a

12:40

good data set and what's a bad data set

12:42

for that approach to scale. So, it is

12:44

bottlenecked by humans.

12:45

>> Doesn't my loss curve decide like how

12:47

much better did I get with this data set

12:49

versus that better data set?

12:50

>> Right. But, if all the engineers open AI

12:53

and tropic or all the big new labs and

12:55

engineers went on vacation,

12:57

who would generate the synthetic data?

12:58

That's the question. It doesn't doesn't

13:01

come from agent's experience. It's not

13:03

something that the agent is generating

13:04

itself.

13:05

Some human has to decide what is the

13:08

right synthetic data to generate. And

13:09

that requires human expertise. So, for

13:11

example,

13:12

if you want a system to do something

13:15

very challenging from a physics point of

13:16

view. Maybe you want a drone that flies

13:18

with echolocation, like a bat, for

13:20

example. Um what's the right synthetic

13:22

data for that? I think you would need to

13:24

hire domain experts to go figure out

13:28

what is the right data and generate it

13:29

and then maybe you would be able to

13:31

learn from that. But

13:33

the domain expert has to exist first.

13:35

So, we are bottlenecked by human

13:36

expertise at that point.

13:37

>> But you can have infinite synthetic

13:39

worlds. The existing world is finite.

13:42

>> But let's go back to the echolocation

13:44

thing, right? That's what I want. I want

13:46

a drone that can uh

13:48

localize itself and move with

13:50

echolocation.

13:51

That's my goal.

13:53

Um the robot that that's a robot that's

13:55

generating its own experience. So, it

13:56

could totally learn from its own

13:58

experience,

13:59

but it wouldn't be able to It doesn't

14:00

matter how much synthetic data you

14:02

generate. Doesn't matter if you generate

14:04

synthetic data that captures 50

14:06

different universes, it will not be not

14:08

allow you to do that task without humans

14:11

figuring it out first.

14:12

>> I first just say it's it's the synthetic

14:14

data is wrong.

14:16

I mean, it won't be correct. It'll be a

14:18

synthetic world. It won't be the real

14:20

world.

14:22

And it will matter.

14:23

The world is incredibly complex. If you

14:26

write a little program, cuz this is

14:28

going to be a little program that will

14:29

generate the synthetic data,

14:31

it'll be a very It'll be a small world.

14:35

>> Mhm.

14:36

>> So,

14:37

for example, what's important to me is

14:39

what's going on in your mind right now.

14:42

Okay? And why You're saying, "Why don't

14:43

I get some synthetic data to tell me

14:46

what's going on in other people's

14:47

minds?" No, there's no way we can have

14:50

synthetic data for other people's minds.

14:53

And other people's minds matter to us.

14:55

You know, like

14:57

I talked to you guys about investing

14:58

today, so I care what you going on in

15:00

your minds.

15:01

And how how can I get synthetic data on

15:03

such a thing?

15:05

Really, you can't even get synthetic

15:06

data on anything. You can't get

15:07

synthetic data on on the how the drone

15:09

is going to interact with with its

15:11

environment in the physical world and

15:14

the the

15:15

the

15:16

the friction and where in the motors of

15:19

this of this robot. The world is

15:21

infinitely complex, and any simulation

15:24

of it is like microscopic. The big world

15:27

hypothesis, let's say what it is, is

15:30

that the world is massively more complex

15:33

than your mind, than any agents any

15:36

agent.

15:37

And this is obvious because the world

15:39

contains many other agents. So, because

15:42

the world is is massively complex, as

15:46

you could no way you can do anything

15:48

like

15:49

anything that might claim to be optimal

15:52

or perfect. You're going to be

15:54

imperfect, and you have to have

15:55

approximations, and those approximations

15:57

will will be severe. And so, because of

16:01

that, that is the ultimately the reason

16:03

why we have to continue learning, if you

16:05

want to think of it as a reason.

16:08

We have to continue learning because

16:09

we'll encounter some particular part of

16:11

this immense world, and we'll have to

16:13

learn an approximation that's tuned to

16:15

the part of the world we're in, not to

16:16

the

16:17

all the other parts that we're not in.

16:19

>> Yeah. I'm I'm going to push on this one

16:21

more time. Um, and sorry, I'm being

16:23

argumentative for the sake of being

16:24

argumentative, [clears throat] but I'm

16:25

trying to understand. My understanding

16:27

is that the newest cohort of

16:28

self-driving car companies, many of them

16:31

were primarily trained in sim,

16:33

and then they, you know, have to do some

16:36

some sort of

16:37

post-training, I guess, to to make sure

16:39

they work in the real world, but but

16:40

that it's been a very effective

16:42

pipeline.

16:42

>> Yeah. So, I think the important question

16:45

to ask here is, how many engineers were

16:47

involved in building that simulation,

16:49

and are we ready to say that the only

16:52

problem worth solving are those where we

16:54

can hire a large team of engineers to

16:56

first make a simulation. And I'm sure

16:59

they had to do multiple iterations where

17:00

they made the simulation, they learned

17:02

in it, they realized there was a

17:03

sim-to-real gap that was not acceptable,

17:05

then they fixed it. So, there is this

17:06

human in the loop fixing the simulation.

17:10

Like, they're getting feedback from the

17:11

real world, humans, and then they're

17:12

fixing the simulation. Why can't we just

17:15

remove the human and let the agent do it

17:17

itself.

17:18

>> And then when it actually drives, again,

17:20

something unexpected will happen.

17:22

>> And that's when you really want to learn

17:23

from experience.

17:24

>> So, your point is there's just so much

17:26

more data that's going to come from

17:27

experience than there possibly can be

17:30

from

17:31

humans

17:32

curating and creating data.

17:34

>> Yeah, and think there is a there is

17:36

obviously value in learning from

17:38

simulation. And and there is a way of

17:40

doing it. The agent can learn a model

17:41

from its own experience. And when the

17:43

agent learns it, it's very it's much

17:45

better because if the model is

17:46

incorrect, it can fix it by continuous

17:48

learning.

17:49

If the humans are making a simulator,

17:51

then the model only gets updated when

17:52

the humans figure out that something is

17:54

wrong. So,

17:55

yes, planning is important. The agents

17:57

should learn from uh simulators, but

17:59

simulators they make themselves.

18:01

>> Okay, I want to move to another part of

18:02

the bitter lesson, removing human

18:04

knowledge. From your essay, quote,

18:07

"Seeking the improvement that makes a

18:08

difference in the shorter term,

18:10

researchers seek to leverage their human

18:11

knowledge of the domain. But the only

18:13

thing that matters in the long run is

18:15

the leveraging of computation."

18:17

And if I, you know, your former student

18:19

Dave, uh with AlphaGo and AlphaZero, for

18:23

me that was a an example of a triumph of

18:25

removing human priors.

18:27

Did that result surprise you? I guess

18:29

why or why not?

18:31

>> Of course, it made me very happy. It

18:32

made me, you know, feel vindicated. Um

18:36

you know, it could have gone either way.

18:38

It wasn't that

18:40

uh cuz cuz prior knowledge can help. You

18:43

know, there's nothing wrong with prior

18:45

knowledge. You know, and and I say this

18:47

right at the very beginning of the

18:48

right at the very beginning of the

18:50

bitter lesson, I say there's no reason

18:52

why there has to be a conflict between

18:54

prior knowledge and then learning

18:55

knowledge. You know, you can put some

18:57

prior in there and then start learning.

18:59

There's no reason in principle why these

19:01

have to be opposed. In fact, they're all

19:03

about knowledge. You know, life is

19:05

gaining knowledge and having knowledge.

19:08

And

19:09

why are these how somehow, you know,

19:10

nature and nurture became enemies? but

19:12

really you know, prior learning is what

19:15

you already and then and you then you

19:16

learn more and it's they should be

19:18

friends.

19:20

But, as I say at the beginning of the

19:21

bitter lesson, in practice they have

19:24

been enemies. In practice, people who

19:27

who had a an affection for existing

19:30

human knowledge ended up, you know,

19:32

wanting that to win and so they wanted

19:34

to

19:35

to minimize or or dismiss learning.

19:38

And so now, I'm sure your sense of me is

19:40

that I'm someone who who loves learning

19:43

and wants to dismiss prior knowledge. Um

19:47

but you know, I'm really

19:50

someone who's interested in the mind.

19:51

The mind is you have prior knowledge and

19:53

then you get more and then once you've

19:55

gotten more, then that becomes your

19:56

prior knowledge as you get more and more

19:58

and more.

19:59

And this these two things work together.

20:02

I end up appearing to be

20:04

someone who's who's interested in

20:06

learning

20:08

primarily because all the rest of the

20:10

world is is is is talking about all you

20:13

need is enough knowledge. You don't need

20:14

to learn.

20:15

You know, large language models are

20:17

we're going to put all this knowledge in

20:18

the into the system and the large

20:20

language model will not learn when it

20:22

runs. You know, you know, it's talking

20:24

to people, it's interacting. It is

20:25

absolutely the weights never change.

20:28

So,

20:29

you know,

20:30

I am not the weird one. It's you guys

20:32

that are the weird one that think that

20:34

that's possible that you could possibly,

20:36

you know, they claim they can like a PhD

20:38

level experience and expertise out of

20:40

something that doesn't learn at all

20:42

anymore.

20:44

You know, so

20:45

you know, I'm not the weird one.

20:47

>> [laughter]

20:47

>> So, your recommendation is let it let

20:49

the algorithms run for much much longer

20:51

period of time before feeding

20:52

>> Continually learn.

20:54

>> before you feed it data.

20:56

with prior data.

20:58

So, drip or drip the prior data along

21:01

the way.

21:02

prior knowledge.

21:04

>> So, both are important.

21:06

Um but in the long run, you've got to

21:08

gain and structure the gaining of new

21:10

knowledge.

21:12

That's that's what all that matters in

21:14

the long run. And as you are doing this,

21:17

yeah, there'll be some that you had

21:20

got previously.

21:22

Um,

21:23

like how it would work if we, you know,

21:25

look into the future when we have

21:26

intelligent robots.

21:28

We will we will will we

21:30

uh, have them all learn from scratch?

21:34

Or will we like copy them and ask them

21:36

to keep learning from wherever they are?

21:39

I mean, we they'll be digital and it'll

21:40

be easy to copy them. And so, instead of

21:42

having like this huge thing where we're

21:45

spending zillions of dollars to retrain

21:48

them from the internet, we'll just copy

21:50

the agent and keep learning

21:53

from there. And and so, in some sense,

21:55

the prior knowledge will be

21:58

should be dismissed cuz you're just

21:59

going to copy it from the previous

22:01

robot.

22:02

>> So, why don't you describe for us what

22:04

you think a machine or computer that

22:07

learns from experience looks like?

22:09

>> Well, it could look like a robot. It

22:11

could be It also could be live entirely

22:14

on the internet.

22:15

You could like, for example, routing a

22:17

package through the internet and do that

22:19

in a way that's sensitive to experience

22:21

and and becomes better over time.

22:23

Or you can interact via the user

22:25

interface that's interacting with

22:26

people, like on your phone or on your

22:28

computer, and uh, it becomes better over

22:31

time. Yeah, like an intelligent

22:32

assistant, you know, has to become

22:35

better over time. It has to know what

22:36

you want.

22:37

>> Would your contention be that the

22:39

current paradigm of,

22:42

you know, the popular A L M based

22:44

assistants, would your contention be

22:46

that these are not experiential learners

22:48

or continual learners? And if so, what

22:50

is the fundamental gap?

22:52

>> Are you serious?

22:54

>> [laughter]

22:55

>> I mean, obviously they

22:56

>> They they have memo- they learn memories

22:58

about me.

22:59

They're

23:01

you know, they're they're they're doing

23:02

some in-context learning.

23:03

>> Their weights never change.

23:05

>> And and by the way, is a small number of

23:07

the weights changing sufficient or do

23:08

you need all the weights to be changing?

23:10

>> Well, so all think of all the

23:12

structuring and generation of new

23:14

concepts that went into creating the

23:17

large language models. All that is the

23:19

weight learning.

23:21

The

23:22

and you you want to continue be able to

23:24

continue doing that. You don't want that

23:26

to happen just once.

23:27

>> Is another way of saying it is we do too

23:29

much pre-training and post-training

23:31

before we launch the the models. They

23:33

they don't learn after that.

23:35

>> The only point that the big disagreement

23:36

is we don't let them learn after that.

23:38

>> Yeah, we don't let them learn after

23:39

that.

23:40

>> as much pre-training as we want, that's

23:41

okay. Post-training is fine, but then

23:43

when I'm when I'm using the model I

23:46

cannot it's it's it stops learning.

23:48

Uh you can give it more context. You can

23:49

change the state of the model by giving

23:51

it more context. And so it it already

23:53

learned that if the state is different,

23:55

if the state says something new, then it

23:57

will use that to make the next

23:59

prediction. But the model is not

24:00

learning.

24:01

>> Cursor's tab other complete model. It

24:03

is, you know, it does get updated based

24:06

on

24:07

>> Those models those weights change.

24:09

>> weights change. Those are two examples

24:10

like Cursor's tab and I think the

24:11

composer they were also updating. Those

24:13

are two examples of continual learning.

24:15

>> Okay.

24:16

>> it's can be much better. So the way they

24:18

do it as far as I understand is a lot of

24:20

people are using tab, they collect all

24:22

this data so coming from millions of

24:24

users or thousands of users and then

24:26

they do one update of the policy from

24:28

this batch data. Um so now

24:31

this could work, but imagine I want to

24:32

teach this model something specific. I

24:34

don't want to fight with 100,000 other

24:36

peoples about what they want to treat

24:38

teach their models. I want to teach my

24:40

model something very specific and I want

24:41

to do it

24:43

to my version of the model. I like I

24:44

don't care about the shared knowledge

24:46

that the model has coming from other

24:47

people.

24:48

And so it's a very inefficient way of

24:50

doing it.

24:50

>> Mhm.

24:51

It seems like the way that this is

24:53

currently done is that there's

24:54

fundamental skills maybe that are

24:56

learned

24:58

in the weights that are common to

24:59

everybody.

25:00

And then there's personalization that

25:01

happens in the form of context, right?

25:04

>> Yeah.

25:05

>> Is that not the right mental model for

25:07

how learning should work? Like should

25:08

should all the context live in the

25:10

weights themselves?

25:11

>> So, context can be in the state, too.

25:13

Could be both.

25:15

But, you still need to be able to update

25:17

the weights. So,

25:18

if I give you an example, some really

25:21

good use studies are for with human

25:23

disabilities. When human go through

25:24

something that changes their mind or or

25:26

some sensors, you can see them adapt.

25:28

So, for example, we have proprioception,

25:30

we have internal sensors that tell us

25:32

where the how the body is positioned,

25:33

and we use this for walking.

25:35

There are cases where people lose this

25:37

ability completely, and then they can't

25:38

walk at all because that is literally

25:41

the foundation of their walking

25:43

policies. It is ingrained in the brain.

25:45

But then over the course of 2 3 years,

25:47

they can learn to walk again by looking

25:50

at their feet. So, visual feedback

25:52

through that. So, brain is

25:54

insanely plastic in the sense that it it

25:56

can learn a lot of things. Something

25:58

that has been true for 20 years, when it

26:00

stops being true, it can go and update

26:02

that and get rid of that. And that is

26:04

the capability I think that's extremely

26:06

useful we would want in our systems.

26:09

>> Hm. What is there for us to learn from

26:11

how human babies or animals learn? And

26:14

how much inspiration do you take from

26:16

that?

26:16

>> Well, we take uh a lot of inspiration.

26:20

We don't take it as a requirement that

26:22

the AI has to

26:24

behave like

26:25

uh the natural system, like babies or

26:28

people or animals.

26:30

Um but it's it's a source of

26:31

inspiration. Inspiration, but not

26:33

constraint from animal learning.

26:35

>> Consistent with the better lesson?

26:36

>> Yeah.

26:37

>> [laughter]

26:37

>> Yeah.

26:39

>> Where do you think we should most seek

26:42

to draw inspiration from the way that

26:43

biological learnings works that is not

26:46

present in today's systems?

26:48

>> I feel like I'm just giving opinions

26:49

now, but they're just obvious opinions.

26:52

So, so I think it's apparent that

26:55

no animal learns by supervised learning.

27:00

Because

27:02

we don't get examples of how our muscles

27:04

should twitch.

27:06

And that's our output.

27:07

>> But all of school is supervised

27:08

learning.

27:09

>> I I know. Absolutely not.

27:12

Uh, but even if it was, school is like a

27:15

tiny fraction of what we learn.

27:18

Like we learn to see, we we learn to

27:20

walk.

27:22

And we learn, um,

27:24

how the world works.

27:26

But even Yeah, and even in school, you

27:27

know, no one tells us how we should

27:29

twitch our muscles.

27:31

>> The knowledge skills I acquire

27:33

were were from supervised learning in

27:34

school.

27:35

>> don't want to say that that that

27:37

learning from

27:38

from others, transmission from others,

27:40

is not important. It's like extremely

27:42

important. And language is extremely

27:43

important.

27:45

Um, but what are we

27:47

what are we missing?

27:48

You know,

27:49

there is there is no

27:51

supervised learning. There's no targets

27:54

that are given to us.

27:55

You know, you you hear the right answer

27:57

is, you know,

27:59

where where is

28:00

what's the capital of France? And we

28:02

know the answer is Paris.

28:05

Okay, but no one tells me how I should

28:07

pronounce Paris.

28:09

You say the answer is Paris, and I

28:11

listen to you, and I hear your words,

28:13

and, you know, I will make some other

28:16

uh, muscle motions to produce the answer

28:19

Paris.

28:20

It's not literally supervised learning.

28:23

Um,

28:24

anyway, yeah. So, I think it's really

28:27

true. I mean,

28:29

well, anyway, the first thing is a

28:30

school is

28:31

is irrelevant. Like, you know, squirrels

28:34

don't go to school and and and learn

28:36

[laughter] that.

28:37

>> They might.

28:37

>> Animals don't learn that way. It's And

28:40

school is a very special thing that that

28:42

even we didn't have up until, you know,

28:44

I don't know, a few hundred years ago.

28:46

It's not part of in not part of the

28:48

essence of intelligence? And it's a

28:50

distraction to think of that as your

28:52

primary example of learning is this

28:54

thing which we didn't do

28:58

as animals.

28:59

>> I wish you had been around to tell my

29:01

parents that before I was made to have

29:03

good go to school and

29:06

deal with all the structure.

29:08

>> The thing is like squirrels are

29:09

wonderful at jumping off trees, but

29:10

squirrels can't prove math theorems. And

29:12

if I want to learn how to prove a math

29:13

theorem, I go to school.

29:14

>> Yeah.

29:15

Uh

29:16

they also don't have

29:18

uh

29:19

DVDs and

29:20

>> [laughter]

29:21

>> and iPods. You know, there are a lot of

29:23

things

29:23

they can do things that we can't do.

29:25

Um but

29:27

>> [sighs]

29:28

>> math theorems

29:30

uh

29:31

Yeah, and they don't play chess.

29:33

You know, it's sort of like more of X

29:34

paradox. They're uh

29:35

they're these advanced things that we

29:36

think of as really intelligent. But uh

29:40

they're sort of easy for computers to do

29:42

as opposed to all these regular things

29:44

that are hard. Like moving and seeing

29:47

with attention and everything.

29:50

Um

29:50

I think supervised learning is a is a

29:52

good thing.

29:53

You know, just mentions I like to think

29:55

look for obvious things. No one tells us

29:58

how to twitch our muscles by giving us

30:00

examples cuz they couldn't possibly cuz

30:02

we have had to twitch our muscles. We've

30:04

had to figure that out.

30:06

>> Yeah.

30:06

>> And their answer would be wrong, right?

30:08

So, if I moved my

30:10

mouth and my tongue and my vocal cords

30:12

exactly the same way that Rich does to

30:14

pronounce Paris, I'm sure a very

30:16

different sound would come out. So, in

30:18

some sense

30:19

Rich or no one knows the right way of

30:21

producing a sound with my body. Only I

30:23

know that.

30:24

>> Yeah.

30:25

It seems to me that many of the most, I

30:27

guess the most raw like sensory motor

30:29

capabilities, especially related to

30:31

movement in the physical world.

30:34

I agree with you that that seems

30:36

something that is inherently learns from

30:37

experience.

30:39

It seems to me though that there are

30:40

higher levels of abstraction that bring

30:42

us closer to, you know, what makes

30:44

humans great.

30:46

And much of that

30:47

doesn't live in this low level of

30:49

sensory motor learning. Does your world

30:51

model, I guess, span

30:54

sensory motor learning all the way up?

30:56

>> Yeah, that's the ambition, absolutely.

30:59

And squirrels, by the way, can do some

31:01

enormously abstract things.

31:03

>> What's the coolest thing a squirrel can

31:04

do?

31:05

>> Well, it can always get into your bird

31:07

feeder,

31:08

>> [laughter]

31:08

>> no matter what obstacles you put in the

31:09

way, you know, it can find new ways to

31:12

jump and climb and

31:14

>> Okay.

31:14

>> and do lots of things.

31:16

>> Calculate trajectories pretty well.

31:18

Animals are pretty good at understanding

31:20

the physical world without the mental

31:22

calculations that we think we are doing

31:24

when we

31:26

think about launching ourselves

31:28

into space.

31:29

>> Breaking a fall, they can do it in real

31:30

time in the right way to prevent

31:32

injuries.

31:33

>> Okay, fair enough.

31:34

>> I think it's just a question of degree

31:36

between and I like to think that

31:38

animals, other animals, are are very

31:41

close to humans. I think it's hubristic

31:43

to try to emphasize what we do

31:45

differently, you know, how we're

31:47

different from animals.

31:48

It's better to see the commonalities.

31:51

And I I think we are just a question of

31:53

degree. It's degree and of course

31:55

society and culture give us big

31:57

advantages. Language give us big

31:59

advantages.

32:00

>> Can I just push on some of this?

32:02

>> Yeah, good.

32:02

>> Because I want to back up Sonia. So, I

32:05

believe

32:07

animals and children learn from

32:09

experience and do incredible things

32:12

learning from experience.

32:13

And

32:15

when my son was two or three or four,

32:17

I'm like, "Wow, this is really

32:18

interesting that they're my son can

32:20

learn these things without nobody

32:23

really teaching him how to do these

32:24

things."

32:25

But at the same time, what Sonia's

32:27

saying is like what makes human uniquely

32:29

human, to be able to

32:31

go to outer space, build a rocket.

32:34

Those Those not things that are learned

32:37

100% from experience because before you

32:40

launch the rocket,

32:41

you actually have to abstract thinking

32:44

through it in a way that is not learned

32:47

from {quote} {unquote} experience.

32:49

Because you don't know if it's going to

32:51

work or not. You have to imagine it. How

32:53

do you we teach

32:55

a machine to imagine things that were

32:57

not available before? That's probably

32:59

the thing that we're trying to like push

33:01

on

33:02

because that we're not quite

33:03

understanding that.

33:04

>> actually going to agree with you there.

33:06

You have to be able to plan. You have to

33:08

be able to imagine.

33:09

>> Yeah.

33:10

>> Would you say that humans 1,000 years

33:12

ago, before they had done all most of

33:14

the thing that we're talking about, were

33:15

they as intelligent?

33:18

Um if for example someone from that era

33:20

was exposed to this new culture, would

33:23

they be able to get the same skills and

33:25

and start doing useful things?

33:27

>> Even over the last 10,000 years, I don't

33:28

think the human brain has evolved that

33:30

much. Because

33:31

>> Fundamentally the same machine.

33:33

>> Fundamentally the same machine,

33:35

but we've built up 10,000 years of

33:36

knowledge.

33:37

>> Yes.

33:37

>> And I get to learn 10,000 years of

33:39

knowledge by going to school through

33:41

supervised learning.

33:42

>> Right.

33:42

>> And I get all that much, much faster

33:45

than trying to learn through experience.

33:47

>> Right.

33:48

>> So I think you're like totally right. So

33:49

we we would want our systems to learn

33:52

from experience and part of their

33:54

experience would be getting exposed to

33:55

our culture and then learning from about

33:57

our culture. They should learn from

33:59

that. That's all good. But let's talk

34:01

about when someone goes and does a

34:03

paradigm shifting thing. So everyone

34:06

gives the example of Einstein, but I

34:07

think there are many examples. Learning

34:09

is that too, like looking at

34:11

learning thing versus

34:12

programming thing.

34:14

So when these paradigm shifts happen, I

34:15

would say it's a human who has

34:17

accumulated all this knowledge and then

34:19

from their experience they're building

34:20

new abstractions and they're planning

34:22

with them and then they're discovering

34:23

new knowledge. And that skill of of

34:26

coming up with new abstractions and then

34:27

learning what models and planning with

34:29

them, that problem is that skill is

34:31

totally missing in our current systems.

34:33

And you can expose this at the edge of

34:36

human knowledge, but you can also study

34:37

this problem at the sensory motor stream

34:40

level.

34:41

>> So, we're not arguing with the

34:44

with the principle. We need to form

34:47

abstractions so we can reason at a high

34:48

level.

34:49

You guys are coming close to doing that

34:52

thing that I said we should never do,

34:53

which is argue is

34:55

prior knowledge important or gaining

34:57

knowledge important. You know, that's

34:58

what you guys just said. You said It's a

35:00

you're you're still going to have to

35:01

learn things. And you're saying, "Oh, I

35:03

can get things from my culture and from

35:06

prior knowledge."

35:08

But these should not fight for each

35:09

other.

35:10

>> On the exact thing around paradigm

35:12

shifts, how do we create a machine that

35:14

understands when to shift the paradigm?

35:16

>> Yeah, I think through its experience,

35:18

right? So, it would have to

35:20

through its own experience. It can't

35:22

rely on human knowledge because we're

35:23

assuming the humans see one paradigm and

35:26

we want a different way of looking at

35:28

things.

35:28

And so, through its experience, it has

35:31

to find something that is better. Maybe

35:34

it it generalizes better and makes

35:36

better predictions. Maybe it's better in

35:38

some other ways, but it has to be

35:39

through its own experience.

35:40

>> The big challenge that we don't see in

35:42

our field, the ability we don't see in

35:44

our field yet, is the ability to learn a

35:46

model and then plan with the model.

35:49

We can do the

35:51

the math things and we can do AlphaGo

35:53

because

35:54

the games, we know the model. We know

35:57

how the moves work.

35:58

And in math, we know what the operators

36:00

are.

36:02

We you know, we know lean will take us

36:03

from one state of knowledge to the to

36:05

the state of the proof to the next

36:07

state.

36:08

But if we have to learn the models,

36:10

there are no

36:12

I'm I'm going to say it. It's probably

36:13

maybe a a weird example, but

36:16

a counterexample, but I can see that

36:17

there's no instances of learning the

36:19

model and then planning with the model

36:21

in our field.

36:23

>> At least not with

36:24

uh uh,

36:26

like self-discovered

36:27

abstractions. So,

36:29

there are people who say, "I'm just

36:31

going to learn a model of what happens

36:32

in the next second or next millisecond."

36:34

But, that's not how our models work. Our

36:36

models are more abstract. Our models

36:39

are, uh,

36:41

quite different.

36:42

>> So, one of the things I like about what

36:44

you're doing here is you're not just

36:45

sitting around pontificating or

36:47

lamenting the state of the world as it

36:48

is. You're very action-oriented. It's

36:50

why you started a company. So, let's

36:52

let's start talking about that a bit.

36:54

Uh, in 2022,

36:55

Rich, you laid out a very specific

36:57

12-point plan, the Alberta plan for AI

37:01

research.

37:02

Maybe tell us about that.

37:04

>> So, the Alberta plan came about because

37:07

we just have general ideas, but we also

37:09

needed to convert them into smaller

37:11

chunks.

37:12

And so, the 12 steps

37:15

are the attempt to uh, crystallize

37:17

particular chunks.

37:18

There's a very important early step,

37:21

step two,

37:22

uh, which is

37:24

continual deep learning.

37:26

And And we think that one is like almost

37:29

the most important because it unlocks

37:31

everything else. If you could do

37:32

continual deep learning,

37:35

you could then continually update your

37:37

model of the world.

37:39

And then, if you knew how to do the

37:40

abstraction rights in like

37:42

the second half

37:43

of of the the steps are all about how to

37:45

get the abstractions right.

37:47

So, and not only my abstractions right,

37:50

what I mean by I don't mean get the

37:51

right abstractions cuz no one can say

37:53

what the right abstractions are. That

37:54

depends on the world that you're in.

37:56

Your agent would have to learn the

37:58

correct abstractions for whatever world

37:59

it's in.

38:01

And so, you know, if you maybe those are

38:04

the two key things. You have to find the

38:05

right abstractions, and then you have to

38:07

do able to continual deep learning.

38:10

>> I think that a lot of the people in the

38:12

field realize that we need models, we

38:13

need to plan with them. But, the

38:15

abstractions tell us what the model

38:17

should be conditioned on. So, what

38:19

should you What should the model

38:21

predict? What are you going to do and

38:22

then something is going to happen. And

38:24

more importantly, where would that come

38:26

from? So, I really like the example of

38:28

elite athletes. If you ask elite

38:30

athletes about how they do certain

38:32

things, they would have weird niche

38:35

terminologies for doing very specific

38:37

things. They were like, you know, I do

38:38

this thing and they would have a name

38:40

for it. If they communicate, sometimes

38:42

they don't even have a name for it if

38:43

they're just doing it alone.

38:45

So, how did they come up with those

38:46

abstractions? That's in some sense a

38:49

crucial thing that's missing that the

38:51

later half of Alberta plan answers.

38:53

>> Can we talk about the continual deep

38:54

learning part?

38:56

Is it an algorithmic gap that exists

38:59

today or is it a just a practical

39:01

deployment infrastructure data privacy

39:03

gap? Because if I wanted to do

39:05

call it naive updating of weights based

39:08

on user interaction, I can do that

39:10

today, right? And so,

39:12

what in your opinion is the biggest

39:14

thing that we're missing to kind of get

39:16

to continual deep learning?

39:18

>> Yeah, so it's absolutely an algorithmic

39:20

gap. You You can do the naive thing, but

39:22

then you'll see all sorts of problems.

39:24

So, for example, if you say um I'm going

39:26

to take one sample and then I'm going to

39:29

update my whole model with that one

39:30

sample. You will run into this problem

39:32

that now all of the previous knowledge

39:34

in the model it's impacted negatively.

39:36

And the way currently we we

39:38

get around this is exactly what cursor

39:40

does. They don't use one example. They

39:42

use a large batch coming from a lot of

39:43

users. So, in use cases where you can

39:45

have that, you can do continual

39:47

learning. But, most use cases you don't

39:49

have that. Most use cases you have a

39:51

single stream of data. And then if you

39:53

apply it to the naive thing, it just

39:55

completely destroys your prior knowledge

39:57

in a very

39:58

um

39:59

destructive way.

40:00

>> Catastrophic forgetting.

40:02

>> Yeah.

40:02

>> That is the Yeah.

40:04

>> But, it's totally curable. You have to

40:06

[laughter] have the right algorithm.

40:07

>> cure?

40:07

>> Well, yeah, exactly.

40:09

>> What is the cure?

40:09

>> Well, you know, first you need to do

40:13

what we call step size optimization.

40:16

And it means every weight in your

40:17

network has to have a separate step

40:19

size. So, some will move fast, some will

40:21

move slow. And you will we you will have

40:24

to metalearn the step sizes for each

40:27

weight. Most of your network will be

40:29

have have weights that have tiny step

40:32

sizes. So, then when you train on a new

40:34

example, they don't get destroyed.

40:37

Happens just to the right places.

40:39

And then secondly, you have to use some

40:43

form of generate and test.

40:45

Um

40:46

which is

40:47

in feature space. So, you come up with

40:49

new features or new units and and

40:52

without following gradients. Cuz

40:54

gradients are very slow process. You

40:55

only move in a direction if you know

40:57

it's the helpful one. And that's always

41:01

going to be very slow and doesn't give

41:02

you a path to grow more and more complex

41:06

and and to have sustained learning. You

41:08

need to have something that just

41:09

proposes a bunch of new units.

41:12

And

41:14

and and then goes from there. I guess

41:17

so, there is a specific thing I can say

41:19

that make it at least concrete, which is

41:21

to say we have this algorithm called

41:23

continual backprop. We used published in

41:26

in nature a couple years ago. And it it

41:29

is exactly like backprop, but every but

41:32

you also plant new seeds of units that

41:34

are newly initialized with random

41:36

weights.

41:37

Backprop only has random weights at the

41:39

beginning of time. And then as you go

41:41

on, all that randomness all that variety

41:43

from the randomness gets used up.

41:46

And with continual backprop, we keep

41:47

injecting a bit of randomness, a bit of

41:49

generate and test, a bit of generate and

41:52

then the

41:53

the operation of backprop is the tester.

41:56

So, you need you need that.

41:58

And and if you put those together really

42:00

well, I think you'll have a new

42:01

generation of massively

42:04

superior

42:05

continual deep learning.

42:07

And that's what we hope to do in the

42:09

next couple years.

42:10

>> Wonderful.

42:11

Do you think that these algorithms

42:14

can be applied to the current state of

42:15

affairs with people scaling LLMs and

42:18

trying to get them to do continual

42:20

learning without catastrophic

42:21

forgetting?

42:22

>> Yeah, absolutely. I think it's

42:24

So, I don't think that you could take an

42:27

existing model and say I'm going to just

42:29

start updating it with these algorithms

42:31

because these algorithms meta learn how

42:33

to learn. So, really you have to say,

42:36

I'm going to learn from scratch. So,

42:37

let's say I learn a new foundation

42:38

model, but I'm going to learn with this

42:40

these new algorithms. These new

42:42

algorithms in addition to

42:44

learning the knowledge, they're also

42:46

going to learn how to learn future

42:48

things. So, they're learning two things

42:49

at the same time. And then um then I

42:53

think you would be able to learn new

42:55

things without catastrophic forgetting.

42:57

>> the most radical thing in that you're

42:59

trying to do in your company in terms of

43:01

from the current state of affairs

43:04

to try to do these two things at the

43:05

same time?

43:06

>> Most radical thing.

43:07

>> Is I think that's

43:09

>> This goes back to I'm not crazy,

43:10

everyone else is crazy.

43:11

>> [laughter]

43:12

>> Yeah.

43:13

>> That's perhaps not totally radical.

43:14

There was a point

43:16

in like 2016 to 2018 where a lot of

43:19

people were exploring these ideas quite

43:21

a bit. They were doing it

43:23

in a much more limited setting. So, they

43:25

would say, we have a distribution of

43:27

problems and then in this specific case

43:29

we'll do it whereas we want to do it

43:30

from a single stream of experience. So,

43:32

our method should be more generally

43:34

applicable. So, I think many people have

43:35

explored this,

43:37

but no one has explored this in the

43:39

general setting where the resulting

43:41

algorithm would be applicable

43:42

everywhere.

43:43

>> So, what would be the most radical thing

43:45

that your company your new company is

43:47

trying to do that other people are not

43:49

doing?

43:49

>> What's the most ambitious thing?

43:52

Remember, I don't think I'm weird, so I

43:54

don't want to say it's radical.

43:56

>> the most ambitious

43:56

>> ambitious thing, I think is to try to

44:00

have the full spectrum of knowledge

44:04

both about the tiny things and about the

44:06

big things. You know, like thinking

44:08

about how you take an airplane from one

44:10

city to another. That's a a very big

44:12

thing. You know, it's it's more it's

44:14

like your

44:15

your space flight example, but it's just

44:18

kind of more common sensical to think

44:21

about cuz we all

44:22

many of us take airplanes and any all of

44:24

us use abstractions on all kinds of our

44:26

life. And even even the squirrels use

44:28

abstractions. So, to have that spectrum

44:32

of of knowledge

44:33

uh

44:34

from the small to the big and to treat

44:36

it in a uniform way and

44:39

to be able to help have it self uh

44:43

maintaining. You know, the big question

44:45

is always you have your knowledge-based

44:47

system and what keeps the knowledge in

44:49

it correct?

44:51

Well, what keeps the knowledge correct

44:53

in a large language model is well,

44:55

people did a lot of post-training

44:58

and and they they they made it sure it

45:00

was correct. And then they freeze it

45:01

after that. So, that's

45:03

what keeps it correct. But really our

45:04

minds, we are we're always changing

45:07

things and yet something keeps it

45:09

organized and coherent and and and

45:12

settling back into a good place rather

45:13

than drifting off into crazy land.

45:16

That is an I think our our biggest

45:18

ambition to have a mind that is

45:20

self-consistent

45:22

and and can

45:23

keep training itself and making it

45:25

coherent.

45:26

>> that.

45:27

Can I ask?

45:28

It almost seems that it's it's such a

45:31

ambitious vision and the idea that all

45:34

these things can be unified into a

45:35

single mind is so ambitious.

45:37

>> It's within reach.

45:39

I think it's within reach. It's here

45:40

it's 2026 and our computers are so fast.

45:44

You know, is it is it

45:46

so ambitious that it's out of reach?

45:48

Uh or or do we have

45:50

already inklings of how all the steps

45:52

can be done? And we

45:54

I I think we have

45:56

a vision

45:58

and inklings.

45:59

Um I don't think it's I don't think it's

46:01

inappropriate.

46:02

>> Your vision involves a trillion

46:04

parameter model with 20 watts.

46:07

That seems pretty

46:09

ambitious.

46:10

>> That is ambitious. Uh, in some sense

46:12

with current technology, I would say

46:13

it's also impossible.

46:15

Like just storing a trillion parameters

46:18

in memory would probably use more than

46:20

20 watts of energy with current memory

46:22

technologies, but

46:25

we are really thinking of okay,

46:28

things are getting better, computation

46:29

is getting cheaper or it is getting more

46:30

energy efficient. So, where would be

46:32

would be in 5 to 10 years? And I think 5

46:35

to 10 years with the right algorithms

46:38

and

46:39

we can totally be in a world where this

46:42

would be possible.

46:43

>> So, 5 to 10 years is two orders of

46:46

magnitude of Moore's law.

46:49

It's a standard improvement.

46:52

If we double every 18 months, 10 years

46:55

would give you two orders of magnitude.

46:57

And so,

46:58

for Kurzweil's statement to be

47:00

plausible,

47:02

then today you should be able to do it

47:03

for

47:05

for what? 20 watts? Two orders of

47:06

magnitude?

47:08

>> 2,000

47:08

>> 2,000 watts. If you can do it with 2,000

47:10

watts today, yeah, then in 10 years

47:12

you'll be able to do it for 20 watts.

47:14

>> You think you can do it for a 2,000

47:15

watts? You have I think lots of people

47:17

at research labs that have access to way

47:20

more than that.

47:21

>> Yeah, I think we can

47:24

can be more efficient than that even now

47:26

with the right algorithm.

47:28

>> If we can be more efficient than that,

47:29

then why aren't we? It's not like people

47:32

just want to spend all their money on

47:34

spend all their money.

47:34

>> Sometimes it seems like they

47:36

>> I'll pay money.

47:38

>> Sometimes it [laughter] seems like they

47:39

want to.

47:39

>> Yeah, I I

47:40

>> Doesn't it?

47:41

I think that's how they show they're

47:43

they're real men

47:44

by using lots of energy.

47:46

>> At least when I look at different

47:47

research groups, I don't even see anyone

47:50

believing in that it's possible. And I

47:51

think if you don't believe in it, you're

47:53

just not going to work on the technical

47:55

problems and work through them.

47:56

>> Is it that it's not possible or it's

47:58

that there's so much waste in the

47:59

system? Like one which one is it? Like

48:02

is it is is there someone who knows how

48:04

to do it efficiently?

48:05

>> Yeah.

48:06

>> And then there's 10 times the number of

48:08

people in the same lab doing all these

48:10

other things.

48:11

And

48:12

so nine out of 10 people are wasting

48:16

>> In some sense I the way I think about it

48:18

is that we are stuck in a local minimum.

48:20

So, if we want to move towards these new

48:22

kind of algorithms,

48:24

it is almost impossible that things will

48:26

not get worse before they get better.

48:27

So, when we start exploring these new

48:28

directions, you're not going to get

48:30

state of the art performance from day

48:31

one, but it is because it is a different

48:33

paradigm. Um but that path leads to

48:36

similar performance at a higher energy

48:38

scale. And these big labs, they are so

48:41

locked into a product that they

48:45

like it is not possible for them to

48:47

pursue a path where things get worse

48:48

first.

48:49

>> Because their current paradigm allows

48:51

them to keep scaling and this new

48:53

paradigm they have to take a bet. And

48:55

then

48:55

>> And they have to figure out some some

48:57

technical things that are difficult that

49:01

we have thought about it for many years.

49:03

We know people who have thought about

49:04

these things for many years and

49:06

when I talk to them, it makes sense that

49:08

it's doable, but

49:09

you need to think about those challenges

49:11

for a long period of time.

49:13

>> So, if everything goes right with Oak,

49:16

what happens with the company? What do

49:18

you what kind of company are you

49:19

building?

49:20

>> Uh if everything goes right, we

49:23

uh implement the architecture, we can

49:25

have a genuine uh continual learning and

49:27

we can form abstractions so that we can

49:30

do planning and reasoning and

49:33

and we have so sort of like true

49:36

intelligence. And then, you know, it's

49:37

hard to imagine just exactly what will

49:39

happen by then.

49:40

But I think

49:41

>> Humans will become irrelevant.

49:43

>> I I don't think that's true at all. I I

49:46

I

49:46

>> We don't either.

49:47

>> I think the world becomes exciting and

49:49

even more exciting and interesting and

49:52

and for humans.

49:53

But in particular, I think there are the

49:56

You have to wonder about the large

49:58

language models.

49:59

They might be at risk.

50:02

Uh when this eventually happens, you

50:04

know, I'm sure they'll get a good run.

50:05

They've already had a good run. You

50:07

know, they've been very successful. And

50:08

let me say,

50:10

just for for clarity that uh large

50:12

language models are an amazing

50:14

scientific breakthrough, a breakthrough

50:16

in the skillful use of language by

50:18

neural networks, wholly unanticipated.

50:22

You know, it was a It was always a hold

50:23

out for uh symbolic methods in language,

50:26

and they

50:28

They have totally changed how that's

50:30

thought about now. Yeah. It's a It's a

50:32

big breakthrough.

50:33

It's so It's frustrating to me that we

50:35

have to you know, just celebrate that

50:38

we've made this great progress in the

50:39

sub subset of the problem of AI, and

50:42

enjoy that. Instead,

50:44

it has to pretend to be all of AI. All

50:46

of intelligence is not fluid, capable

50:50

use of language. There's so much more.

50:52

It's an important part. You know, it's

50:53

like

50:55

20% or a quarter of

50:57

intelligence. There's There's more.

50:59

>> Yeah.

51:00

>> We're not done.

51:01

>> Yeah.

51:02

If everything goes right, are you

51:04

imagining that there's a single minds

51:06

that can do everything from learn how to

51:08

swing from tree branches to make a

51:11

spaceship,

51:12

uh to you know, all these various things

51:15

we've talked about today. Is it a single

51:16

mind, and is it a single set of weights

51:19

that can that can do all these things,

51:21

or is it

51:22

>> It's a single design.

51:23

>> Okay.

51:24

>> And there'll be many different There'll

51:25

be many minds.

51:26

>> Mm. Okay. So, it's a single design that

51:28

reacts to different environments.

51:30

>> And it did it And different versions of

51:33

that mind would learn different things

51:34

because they have different experience.

51:35

This sort of goes back to the big world

51:38

hypothesis that there are infinitely

51:40

many things to learn, so one system

51:42

cannot learn infinitely many things. And

51:45

like I think Rich already mentioned

51:46

this, but if you have two of these

51:48

systems, if you have two of the largest

51:50

systems in the world, then it is trivial

51:52

that they cannot model each other

51:54

because they are equally complex. So,

51:55

single system would never be able to get

51:58

to a point where it can learn

51:59

everything. It would always be multiple

52:01

systems

52:02

that are learning from their own

52:02

experience.

52:03

>> Guys, are you hiring? What kind of

52:05

people are you looking for?

52:06

>> We are hiring.

52:09

The initial team, most of it we already

52:10

have in our mind. So, these are people

52:12

who have thought about these ideas in

52:14

the past in many

52:16

uh you know, for a long time.

52:17

And

52:19

we are going to take a slightly

52:20

different approach because this is a

52:22

different paradigm. It doesn't make

52:23

sense to

52:25

become large very quickly because uh in

52:27

some sense, everyone we hire has to

52:31

uh come to see what we see. And not

52:33

everyone sees that. So, we're going to

52:34

start small, slowly grow to maybe a

52:37

handful or two or three uh and then go

52:39

from there.

52:40

>> We want to be super aligned.

52:42

>> We want to be super aligned.

52:43

>> So, that we can be um

52:46

very productive working together and and

52:49

scaling the progress.

52:50

>> Absolutely.

52:51

>> Very, very cool.

52:52

>> Wonderful.

52:53

I love this conversation. Thank you for

52:55

taking the time to share what you're up

52:57

to. You are, you know, an an unusually

53:00

deep thinker about

53:02

where reinforcement learning and

53:04

algorithmic design will go.

53:06

And it was a true pleasure to get to

53:07

explore it together with you today. So,

53:09

thank you.

53:10

>> Thank you very much. Thank you. It's our

53:11

pleasure.

53:15

>> [music]

53:21

[music]

53:37

[music]

Interactive Summary

In this video, AI researcher Rich Sutton and his co-founder Quorum Javed discuss their work at Oak Lab, focusing on the concepts of reinforcement learning, 'The Bitter Lesson,' and the necessity of 'continual learning.' They argue that intelligence arises from agents interacting with and learning from their environments over time, rather than relying solely on pre-trained models or static human knowledge. The discussion highlights their 'Big World' hypothesis—the idea that the world is infinitely complex and that systems must learn from their own experiences rather than being bottlenecked by human-curated data or simulations. They outline their research agenda, which includes developing algorithms that enable continual deep learning, self-discovery of abstractions, and efficient weight updates, all with the goal of creating more robust, self-maintaining intelligent systems.

Suggested questions

4 ready-made prompts