Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again
1678 segments
People think I'm have a radical point of
view sometimes. They say they start
questions saying how what I'm thinking
is so different from everyone else. But
I don't see it that way at all. I see it
as like I'm thinking the ordinary way.
It's just everyone else that's thinking
a bit weird. And
>> [laughter]
>> And I mean that like, you know, it's
just the recent times people are
thinking weird. Before there was all
this AI craziness,
uh
you talk about you wouldn't have to say
continual learning cuz
it wouldn't make any sense to talk about
learning that wasn't continual. All
learning is continual. We always act and
we learn. That's just the normal way of
thinking. I'm not weird.
The field is weird. The field they need
to call it continual learning. It's just
learning.
>> [music]
>> We are honored to have the great Rich
Sutton with us here today. Rich, you
invented reinforcement learning. You
wrote the seminal textbook.
You're the
key students in the field, uh folks like
Dave Silver. You wrote the essay The
Bitter Lesson that I believe is the
Bible of the field. And and you have
just been one of the greats in
propelling the field forward. So thank
you for taking the time to join us
today. Um Rich is joined by Quorum
Javed, his co-founder uh and former
students from the University of Alberta.
Um the two of you have set off to found
Oak Lab. I'm very excited to talk to you
about that today. So for today's
session, we're going to start talking
about The Bitter Lesson, the state of
the world as we know it today,
whether LLMs will get us there or not.
And then we're going to we're going to
transition to start talking about your
your research agenda and your plan for
Oak. Um Rich, maybe take us back. I was
going to start with The Bitter Lesson,
but I actually want to start earlier
than that. Decades ago, you decided to
dedicate your career to reinforcement
learning, to deep reinforcement learning
in particular, and you established the
University of Alberta as a bastion of
that
back when I think the field was very
much in its infancy. What gave you the
conviction to do that?
>> What else you going to do?
>> [laughter]
>> We were trying to figure out the mind
and learning is a central part of the
mind.
And having a goal is a central part of
the mind. Central part of intelligence.
Yeah, so I was just doubling down on
what I was always thinking.
>> Did people think you were crazy at the
time?
>> Um
It was a winter. It was an AI winter.
>> Uh what year was this?
>> It was in 2003.
>> Okay.
>> And
it's kind of crazy actually the truth
cuz I was like really sick.
I was dying I was actually dying of
cancer in 2003. And but I I wasn't quite
dead, you know, I've been trying for a
number of years. And I wasn't dead after
another remission. And so
so I said, well, I'm not dying I haven't
succeeded in dying. So I might as well,
you know, it's going on long enough
might as well just try to get another
job. And so I so I went to Alberta and
and and started teaching there.
And then in the end I didn't die. It's
kind of amazing it's cuz
you know, it's it's like that. I'm
joking about it now but it was quite
serious. And um
it's an even more important question.
Why why did I continue
to work on this research stuff when I
was, you know, I only had a few months.
I would always
keep reminded what
I think it's Benjamin Franklin is
supposed to have said that, you know, if
you ever wonder why someone is doing
something
it's almost always one of two things.
It's either habit
or vanity. Okay? So I think I think it's
probably true. Maybe it was my habit to
just kept doing what I'd always been was
doing or maybe it was vanity.
I I know. I think it was more like habit
cuz I was I was dying.
>> Wow.
>> Um
>> Wow. Divine intervention.
>> Yeah, it's always been easy for me to
keep be very determined.
Um and
I'm I'm I'm going to go even longer on
this answer.
>> Please go.
>> People think I am have a radical point
of view sometimes. They they they start
questions saying how how what I'm
thinking is so different from everyone
else. But I don't see it that way at
all. I see it as like I'm thinking the
ordinary way. It's everyone else that's
thinking a bit weird.
>> [laughter]
>> And I mean that like, you know, it's
just the recent times people are
thinking weird. If you go look back what
what what what people thought about the
mind for, you know, even just a decade,
you'll find the kind of thoughts that,
you know, learning is important. You've
got to have a goal. Um
and you know, perception is important.
We have a We are We are low-level We are
low-level beings. We are generating
actions and perceiving data at a fast
speed and yet we have to think at higher
levels. And you know, go back a few
before there was all this AI craziness,
uh
you talk about you wouldn't have to say
continual learning cuz
it wouldn't make any sense to talk about
learning that wasn't continual. All
learning is continual.
You know, you
It's not a special phase. We always act
and we learn. That's just the normal way
of thinking. I'm not weird.
The field is weird. The field they need
to call it continual learning. It's just
learning.
>> I'm not weird, everybody else is. That's
a
good good motto to live by.
>> We're going to have to send out a an ex
post about that.
>> [laughter]
>> We're very happy that you you you lived
on with the field is happy that you
lived on and thank you for pushing the
frontier of AI.
>> happy.
>> [laughter]
>> Thank you for pushing the frontier of
AI. I'm sure really happy and thank you
for all of that and push you've like
been able to sort of educate a lot of
students who pushed the frontier as
well.
How did you pick them? How did you
over the last 20, 30 years?
>> Oh, well, you are giving me
opportunities to be uh
to be humble. Like I like to be humble
and point out how all these great
decisions are are just happen.
And it's that that's the way I feel
about students. I don't feel that I
choose them very well. I've just
Sometimes I'm lucky, sometimes I'm
unlucky.
I don't feel I'm particularly good at
picking my students.
I'm looking at Karam. I think sometimes
you end up with the really great ones.
David picked David Silver picked me.
>> Yeah.
>> How is it that
that I got you, Karam?
>> Yeah, that was also so I
finished my master's not with you
uh and I was planning to join industry.
And then we were collaborating on a
project which also just had started
organically. Like there was something I
worked on that Rich was in a meeting,
then they mentioned that I worked on it.
So I got pulled into it. We started
collaborating. It went really well. Like
I felt so happy with that collaboration.
Rich also felt really good about it. And
then 6 months down the road we had made
some progress and it just made sense to
convert that into a thesis proposal. So
at no point did I apply, at no point did
I ask should you be my PhD advisor. We
worked together, then we decided this
would be a pretty good thesis.
And then then after that I applied for
the PhD.
>> Life works in unexpected ways.
Uh take us to 2019. You wrote the bitter
lesson which has become the the mother
in tome.
2019 was a funny time to be writing that
piece because ImageNet was 2009. AlphaGo
was 2015. What caused you in 2019 to
reflect and and to write that? Because
it was before the current kind of
scaling paradigm around large language
models had taken off, but it was after
deep learning had really proven itself.
>> Well, it was a long time coming. You
know, as the
bitter lesson expresses, it's something
that you can
for a long time, for many decades.
And it's definitely at least as much due
to the round of
symbolic AI, which I lived through.
It's all about not getting distracted by
trying to put in your human knowledge
and just paying attention to what the
problem needs and how you can scale with
computation.
I know I I I wrote versions of it at
least a year before and I I gave talks.
I gave a talk a year before.
And um
it wasn't a particular response to the
moment. It was a particular response
to my my long experience, different
people trying to think in different ways
about how you can make smart systems.
>> Mhm.
What is the essence of the bitter
lesson?
>> The essence of the bitter lesson.
>> You know, and maybe the phrase that I
hear used the most in my meetings these
days is is bitter lesson pill, is it not
bitter lesson pill?
I would imagine given the popularity of
the phrase it's probably been tortured
and misused in different ways that you
didn't originally intend it. So, what do
you think what is the essence of it and
where do you think people go wrong in
their in their attempt to understand it?
>> Yeah, you're you're making me think
about X now and my I recently made a
post where I tried to do the the bitter
lesson in in 26 words.
It [laughter] goes something like
don't be distracted by human knowledge
as AI traditionally has been many times.
Instead focus on
learning methods that will scale with
computation
like search and like learning. So, it's
really all about uh focusing on
algorithms and improvements. It's not
it's not saying you don't need
fancy algorithms. You need fancy
algorithms, but you want fancy
algorithms that will scale with scale
with computation.
>> Rather than scaling with data.
>> Rather than scaling with human input.
Yeah, and then the the question if I can
anticipate it
um
Yeah, what about large language models?
>> Yeah.
>> Are they
>> consistent or inconsistent with your
>> Yeah.
>> And and I thought about this and I think
there's an ex- there's another exposed
model, but the conclusion is that it's
both a a positive example and a negative
example of the big lesson. First uh
large language models enabled uh
enormous scaling with computation. And
you could just
drink in the internet and scale so much.
So it was a it was a way of getting uh
much more capable system just by
methods that scale. Then after that, as
you go on further um it eventually gets
limited by by that information. The the
internet is finite and it's hard to get
more examples.
And uh the world is big and the world is
massively bigger than everything we
stored
on the internet. And so in the end, it
seems like it could be
uh I guess that would be a
positive example of when, you know, we
relied too much on human knowledge and
it eventually
holds us back.
>> Mhm.
Can I just push on this a little bit?
>> Yeah.
>> It seems like a lot of what the
foundation model labs are working on
right now is synthetic data generation
in order to kind of get us beyond
the fossil fuel that is the existing
human internet.
Um is synthetic data generation kind of
as part of this LLM scaling paradigm, is
that
a general method that leverages
computation?
>> No, that's that's just a big mistake.
>> Why?
>> [laughter]
>> Well, it's such a big it's such a maybe
it's the next the next big lesson.
Um
it's been floating around
uh Alberta for 5 or 10 years.
>> Okay.
>> And uh we call it the big world
perspective or big world hypothesis.
Khuram, who eventually wrote it up as a
paper. There's a little paper called the
big world hypothesis.
>> So, the big world is
that the world is infinitely big. There
are infinitely many things to learn.
And you can have people generating the
synthetic data sets, but there will
always be more things to learn. And
because of that, if you could just
learn from experience, if you could
remove the humans from the loop, then
you would have systems that can do
everything. Because, you know, the world
is big. There are many tasks that we
want them to do.
And they would be able to do anything by
learning from their experience. Going
back to the synthetic data question,
too. Uh
who decides what's a good synthetic data
and what's a bad synthetic data? Because
I can write a program that can output a
lot of synthetic data, which would hurt
programs. Right now, I would say humans
decide. And that's the bottleneck where
okay, you can have humans deciding how
to generate these data sets, but you
need human experts who know what's a
good data set and what's a bad data set
for that approach to scale. So, it is
bottlenecked by humans.
>> Doesn't my loss curve decide like how
much better did I get with this data set
versus that better data set?
>> Right. But, if all the engineers open AI
and tropic or all the big new labs and
engineers went on vacation,
who would generate the synthetic data?
That's the question. It doesn't doesn't
come from agent's experience. It's not
something that the agent is generating
itself.
Some human has to decide what is the
right synthetic data to generate. And
that requires human expertise. So, for
example,
if you want a system to do something
very challenging from a physics point of
view. Maybe you want a drone that flies
with echolocation, like a bat, for
example. Um what's the right synthetic
data for that? I think you would need to
hire domain experts to go figure out
what is the right data and generate it
and then maybe you would be able to
learn from that. But
the domain expert has to exist first.
So, we are bottlenecked by human
expertise at that point.
>> But you can have infinite synthetic
worlds. The existing world is finite.
>> But let's go back to the echolocation
thing, right? That's what I want. I want
a drone that can uh
localize itself and move with
echolocation.
That's my goal.
Um the robot that that's a robot that's
generating its own experience. So, it
could totally learn from its own
experience,
but it wouldn't be able to It doesn't
matter how much synthetic data you
generate. Doesn't matter if you generate
synthetic data that captures 50
different universes, it will not be not
allow you to do that task without humans
figuring it out first.
>> I first just say it's it's the synthetic
data is wrong.
I mean, it won't be correct. It'll be a
synthetic world. It won't be the real
world.
And it will matter.
The world is incredibly complex. If you
write a little program, cuz this is
going to be a little program that will
generate the synthetic data,
it'll be a very It'll be a small world.
>> Mhm.
>> So,
for example, what's important to me is
what's going on in your mind right now.
Okay? And why You're saying, "Why don't
I get some synthetic data to tell me
what's going on in other people's
minds?" No, there's no way we can have
synthetic data for other people's minds.
And other people's minds matter to us.
You know, like
I talked to you guys about investing
today, so I care what you going on in
your minds.
And how how can I get synthetic data on
such a thing?
Really, you can't even get synthetic
data on anything. You can't get
synthetic data on on the how the drone
is going to interact with with its
environment in the physical world and
the the
the
the friction and where in the motors of
this of this robot. The world is
infinitely complex, and any simulation
of it is like microscopic. The big world
hypothesis, let's say what it is, is
that the world is massively more complex
than your mind, than any agents any
agent.
And this is obvious because the world
contains many other agents. So, because
the world is is massively complex, as
you could no way you can do anything
like
anything that might claim to be optimal
or perfect. You're going to be
imperfect, and you have to have
approximations, and those approximations
will will be severe. And so, because of
that, that is the ultimately the reason
why we have to continue learning, if you
want to think of it as a reason.
We have to continue learning because
we'll encounter some particular part of
this immense world, and we'll have to
learn an approximation that's tuned to
the part of the world we're in, not to
the
all the other parts that we're not in.
>> Yeah. I'm I'm going to push on this one
more time. Um, and sorry, I'm being
argumentative for the sake of being
argumentative, [clears throat] but I'm
trying to understand. My understanding
is that the newest cohort of
self-driving car companies, many of them
were primarily trained in sim,
and then they, you know, have to do some
some sort of
post-training, I guess, to to make sure
they work in the real world, but but
that it's been a very effective
pipeline.
>> Yeah. So, I think the important question
to ask here is, how many engineers were
involved in building that simulation,
and are we ready to say that the only
problem worth solving are those where we
can hire a large team of engineers to
first make a simulation. And I'm sure
they had to do multiple iterations where
they made the simulation, they learned
in it, they realized there was a
sim-to-real gap that was not acceptable,
then they fixed it. So, there is this
human in the loop fixing the simulation.
Like, they're getting feedback from the
real world, humans, and then they're
fixing the simulation. Why can't we just
remove the human and let the agent do it
itself.
>> And then when it actually drives, again,
something unexpected will happen.
>> And that's when you really want to learn
from experience.
>> So, your point is there's just so much
more data that's going to come from
experience than there possibly can be
from
humans
curating and creating data.
>> Yeah, and think there is a there is
obviously value in learning from
simulation. And and there is a way of
doing it. The agent can learn a model
from its own experience. And when the
agent learns it, it's very it's much
better because if the model is
incorrect, it can fix it by continuous
learning.
If the humans are making a simulator,
then the model only gets updated when
the humans figure out that something is
wrong. So,
yes, planning is important. The agents
should learn from uh simulators, but
simulators they make themselves.
>> Okay, I want to move to another part of
the bitter lesson, removing human
knowledge. From your essay, quote,
"Seeking the improvement that makes a
difference in the shorter term,
researchers seek to leverage their human
knowledge of the domain. But the only
thing that matters in the long run is
the leveraging of computation."
And if I, you know, your former student
Dave, uh with AlphaGo and AlphaZero, for
me that was a an example of a triumph of
removing human priors.
Did that result surprise you? I guess
why or why not?
>> Of course, it made me very happy. It
made me, you know, feel vindicated. Um
you know, it could have gone either way.
It wasn't that
uh cuz cuz prior knowledge can help. You
know, there's nothing wrong with prior
knowledge. You know, and and I say this
right at the very beginning of the
right at the very beginning of the
bitter lesson, I say there's no reason
why there has to be a conflict between
prior knowledge and then learning
knowledge. You know, you can put some
prior in there and then start learning.
There's no reason in principle why these
have to be opposed. In fact, they're all
about knowledge. You know, life is
gaining knowledge and having knowledge.
And
why are these how somehow, you know,
nature and nurture became enemies? but
really you know, prior learning is what
you already and then and you then you
learn more and it's they should be
friends.
But, as I say at the beginning of the
bitter lesson, in practice they have
been enemies. In practice, people who
who had a an affection for existing
human knowledge ended up, you know,
wanting that to win and so they wanted
to
to minimize or or dismiss learning.
And so now, I'm sure your sense of me is
that I'm someone who who loves learning
and wants to dismiss prior knowledge. Um
but you know, I'm really
someone who's interested in the mind.
The mind is you have prior knowledge and
then you get more and then once you've
gotten more, then that becomes your
prior knowledge as you get more and more
and more.
And this these two things work together.
I end up appearing to be
someone who's who's interested in
learning
primarily because all the rest of the
world is is is is talking about all you
need is enough knowledge. You don't need
to learn.
You know, large language models are
we're going to put all this knowledge in
the into the system and the large
language model will not learn when it
runs. You know, you know, it's talking
to people, it's interacting. It is
absolutely the weights never change.
So,
you know,
I am not the weird one. It's you guys
that are the weird one that think that
that's possible that you could possibly,
you know, they claim they can like a PhD
level experience and expertise out of
something that doesn't learn at all
anymore.
You know, so
you know, I'm not the weird one.
>> [laughter]
>> So, your recommendation is let it let
the algorithms run for much much longer
period of time before feeding
>> Continually learn.
>> before you feed it data.
with prior data.
So, drip or drip the prior data along
the way.
prior knowledge.
>> So, both are important.
Um but in the long run, you've got to
gain and structure the gaining of new
knowledge.
That's that's what all that matters in
the long run. And as you are doing this,
yeah, there'll be some that you had
got previously.
Um,
like how it would work if we, you know,
look into the future when we have
intelligent robots.
We will we will will we
uh, have them all learn from scratch?
Or will we like copy them and ask them
to keep learning from wherever they are?
I mean, we they'll be digital and it'll
be easy to copy them. And so, instead of
having like this huge thing where we're
spending zillions of dollars to retrain
them from the internet, we'll just copy
the agent and keep learning
from there. And and so, in some sense,
the prior knowledge will be
should be dismissed cuz you're just
going to copy it from the previous
robot.
>> So, why don't you describe for us what
you think a machine or computer that
learns from experience looks like?
>> Well, it could look like a robot. It
could be It also could be live entirely
on the internet.
You could like, for example, routing a
package through the internet and do that
in a way that's sensitive to experience
and and becomes better over time.
Or you can interact via the user
interface that's interacting with
people, like on your phone or on your
computer, and uh, it becomes better over
time. Yeah, like an intelligent
assistant, you know, has to become
better over time. It has to know what
you want.
>> Would your contention be that the
current paradigm of,
you know, the popular A L M based
assistants, would your contention be
that these are not experiential learners
or continual learners? And if so, what
is the fundamental gap?
>> Are you serious?
>> [laughter]
>> I mean, obviously they
>> They they have memo- they learn memories
about me.
They're
you know, they're they're they're doing
some in-context learning.
>> Their weights never change.
>> And and by the way, is a small number of
the weights changing sufficient or do
you need all the weights to be changing?
>> Well, so all think of all the
structuring and generation of new
concepts that went into creating the
large language models. All that is the
weight learning.
The
and you you want to continue be able to
continue doing that. You don't want that
to happen just once.
>> Is another way of saying it is we do too
much pre-training and post-training
before we launch the the models. They
they don't learn after that.
>> The only point that the big disagreement
is we don't let them learn after that.
>> Yeah, we don't let them learn after
that.
>> as much pre-training as we want, that's
okay. Post-training is fine, but then
when I'm when I'm using the model I
cannot it's it's it stops learning.
Uh you can give it more context. You can
change the state of the model by giving
it more context. And so it it already
learned that if the state is different,
if the state says something new, then it
will use that to make the next
prediction. But the model is not
learning.
>> Cursor's tab other complete model. It
is, you know, it does get updated based
on
>> Those models those weights change.
>> weights change. Those are two examples
like Cursor's tab and I think the
composer they were also updating. Those
are two examples of continual learning.
>> Okay.
>> it's can be much better. So the way they
do it as far as I understand is a lot of
people are using tab, they collect all
this data so coming from millions of
users or thousands of users and then
they do one update of the policy from
this batch data. Um so now
this could work, but imagine I want to
teach this model something specific. I
don't want to fight with 100,000 other
peoples about what they want to treat
teach their models. I want to teach my
model something very specific and I want
to do it
to my version of the model. I like I
don't care about the shared knowledge
that the model has coming from other
people.
And so it's a very inefficient way of
doing it.
>> Mhm.
It seems like the way that this is
currently done is that there's
fundamental skills maybe that are
learned
in the weights that are common to
everybody.
And then there's personalization that
happens in the form of context, right?
>> Yeah.
>> Is that not the right mental model for
how learning should work? Like should
should all the context live in the
weights themselves?
>> So, context can be in the state, too.
Could be both.
But, you still need to be able to update
the weights. So,
if I give you an example, some really
good use studies are for with human
disabilities. When human go through
something that changes their mind or or
some sensors, you can see them adapt.
So, for example, we have proprioception,
we have internal sensors that tell us
where the how the body is positioned,
and we use this for walking.
There are cases where people lose this
ability completely, and then they can't
walk at all because that is literally
the foundation of their walking
policies. It is ingrained in the brain.
But then over the course of 2 3 years,
they can learn to walk again by looking
at their feet. So, visual feedback
through that. So, brain is
insanely plastic in the sense that it it
can learn a lot of things. Something
that has been true for 20 years, when it
stops being true, it can go and update
that and get rid of that. And that is
the capability I think that's extremely
useful we would want in our systems.
>> Hm. What is there for us to learn from
how human babies or animals learn? And
how much inspiration do you take from
that?
>> Well, we take uh a lot of inspiration.
We don't take it as a requirement that
the AI has to
behave like
uh the natural system, like babies or
people or animals.
Um but it's it's a source of
inspiration. Inspiration, but not
constraint from animal learning.
>> Consistent with the better lesson?
>> Yeah.
>> [laughter]
>> Yeah.
>> Where do you think we should most seek
to draw inspiration from the way that
biological learnings works that is not
present in today's systems?
>> I feel like I'm just giving opinions
now, but they're just obvious opinions.
So, so I think it's apparent that
no animal learns by supervised learning.
Because
we don't get examples of how our muscles
should twitch.
And that's our output.
>> But all of school is supervised
learning.
>> I I know. Absolutely not.
Uh, but even if it was, school is like a
tiny fraction of what we learn.
Like we learn to see, we we learn to
walk.
And we learn, um,
how the world works.
But even Yeah, and even in school, you
know, no one tells us how we should
twitch our muscles.
>> The knowledge skills I acquire
were were from supervised learning in
school.
>> don't want to say that that that
learning from
from others, transmission from others,
is not important. It's like extremely
important. And language is extremely
important.
Um, but what are we
what are we missing?
You know,
there is there is no
supervised learning. There's no targets
that are given to us.
You know, you you hear the right answer
is, you know,
where where is
what's the capital of France? And we
know the answer is Paris.
Okay, but no one tells me how I should
pronounce Paris.
You say the answer is Paris, and I
listen to you, and I hear your words,
and, you know, I will make some other
uh, muscle motions to produce the answer
Paris.
It's not literally supervised learning.
Um,
anyway, yeah. So, I think it's really
true. I mean,
well, anyway, the first thing is a
school is
is irrelevant. Like, you know, squirrels
don't go to school and and and learn
[laughter] that.
>> They might.
>> Animals don't learn that way. It's And
school is a very special thing that that
even we didn't have up until, you know,
I don't know, a few hundred years ago.
It's not part of in not part of the
essence of intelligence? And it's a
distraction to think of that as your
primary example of learning is this
thing which we didn't do
as animals.
>> I wish you had been around to tell my
parents that before I was made to have
good go to school and
deal with all the structure.
>> The thing is like squirrels are
wonderful at jumping off trees, but
squirrels can't prove math theorems. And
if I want to learn how to prove a math
theorem, I go to school.
>> Yeah.
Uh
they also don't have
uh
DVDs and
>> [laughter]
>> and iPods. You know, there are a lot of
things
they can do things that we can't do.
Um but
>> [sighs]
>> math theorems
uh
Yeah, and they don't play chess.
You know, it's sort of like more of X
paradox. They're uh
they're these advanced things that we
think of as really intelligent. But uh
they're sort of easy for computers to do
as opposed to all these regular things
that are hard. Like moving and seeing
with attention and everything.
Um
I think supervised learning is a is a
good thing.
You know, just mentions I like to think
look for obvious things. No one tells us
how to twitch our muscles by giving us
examples cuz they couldn't possibly cuz
we have had to twitch our muscles. We've
had to figure that out.
>> Yeah.
>> And their answer would be wrong, right?
So, if I moved my
mouth and my tongue and my vocal cords
exactly the same way that Rich does to
pronounce Paris, I'm sure a very
different sound would come out. So, in
some sense
Rich or no one knows the right way of
producing a sound with my body. Only I
know that.
>> Yeah.
It seems to me that many of the most, I
guess the most raw like sensory motor
capabilities, especially related to
movement in the physical world.
I agree with you that that seems
something that is inherently learns from
experience.
It seems to me though that there are
higher levels of abstraction that bring
us closer to, you know, what makes
humans great.
And much of that
doesn't live in this low level of
sensory motor learning. Does your world
model, I guess, span
sensory motor learning all the way up?
>> Yeah, that's the ambition, absolutely.
And squirrels, by the way, can do some
enormously abstract things.
>> What's the coolest thing a squirrel can
do?
>> Well, it can always get into your bird
feeder,
>> [laughter]
>> no matter what obstacles you put in the
way, you know, it can find new ways to
jump and climb and
>> Okay.
>> and do lots of things.
>> Calculate trajectories pretty well.
Animals are pretty good at understanding
the physical world without the mental
calculations that we think we are doing
when we
think about launching ourselves
into space.
>> Breaking a fall, they can do it in real
time in the right way to prevent
injuries.
>> Okay, fair enough.
>> I think it's just a question of degree
between and I like to think that
animals, other animals, are are very
close to humans. I think it's hubristic
to try to emphasize what we do
differently, you know, how we're
different from animals.
It's better to see the commonalities.
And I I think we are just a question of
degree. It's degree and of course
society and culture give us big
advantages. Language give us big
advantages.
>> Can I just push on some of this?
>> Yeah, good.
>> Because I want to back up Sonia. So, I
believe
animals and children learn from
experience and do incredible things
learning from experience.
And
when my son was two or three or four,
I'm like, "Wow, this is really
interesting that they're my son can
learn these things without nobody
really teaching him how to do these
things."
But at the same time, what Sonia's
saying is like what makes human uniquely
human, to be able to
go to outer space, build a rocket.
Those Those not things that are learned
100% from experience because before you
launch the rocket,
you actually have to abstract thinking
through it in a way that is not learned
from {quote} {unquote} experience.
Because you don't know if it's going to
work or not. You have to imagine it. How
do you we teach
a machine to imagine things that were
not available before? That's probably
the thing that we're trying to like push
on
because that we're not quite
understanding that.
>> actually going to agree with you there.
You have to be able to plan. You have to
be able to imagine.
>> Yeah.
>> Would you say that humans 1,000 years
ago, before they had done all most of
the thing that we're talking about, were
they as intelligent?
Um if for example someone from that era
was exposed to this new culture, would
they be able to get the same skills and
and start doing useful things?
>> Even over the last 10,000 years, I don't
think the human brain has evolved that
much. Because
>> Fundamentally the same machine.
>> Fundamentally the same machine,
but we've built up 10,000 years of
knowledge.
>> Yes.
>> And I get to learn 10,000 years of
knowledge by going to school through
supervised learning.
>> Right.
>> And I get all that much, much faster
than trying to learn through experience.
>> Right.
>> So I think you're like totally right. So
we we would want our systems to learn
from experience and part of their
experience would be getting exposed to
our culture and then learning from about
our culture. They should learn from
that. That's all good. But let's talk
about when someone goes and does a
paradigm shifting thing. So everyone
gives the example of Einstein, but I
think there are many examples. Learning
is that too, like looking at
learning thing versus
programming thing.
So when these paradigm shifts happen, I
would say it's a human who has
accumulated all this knowledge and then
from their experience they're building
new abstractions and they're planning
with them and then they're discovering
new knowledge. And that skill of of
coming up with new abstractions and then
learning what models and planning with
them, that problem is that skill is
totally missing in our current systems.
And you can expose this at the edge of
human knowledge, but you can also study
this problem at the sensory motor stream
level.
>> So, we're not arguing with the
with the principle. We need to form
abstractions so we can reason at a high
level.
You guys are coming close to doing that
thing that I said we should never do,
which is argue is
prior knowledge important or gaining
knowledge important. You know, that's
what you guys just said. You said It's a
you're you're still going to have to
learn things. And you're saying, "Oh, I
can get things from my culture and from
prior knowledge."
But these should not fight for each
other.
>> On the exact thing around paradigm
shifts, how do we create a machine that
understands when to shift the paradigm?
>> Yeah, I think through its experience,
right? So, it would have to
through its own experience. It can't
rely on human knowledge because we're
assuming the humans see one paradigm and
we want a different way of looking at
things.
And so, through its experience, it has
to find something that is better. Maybe
it it generalizes better and makes
better predictions. Maybe it's better in
some other ways, but it has to be
through its own experience.
>> The big challenge that we don't see in
our field, the ability we don't see in
our field yet, is the ability to learn a
model and then plan with the model.
We can do the
the math things and we can do AlphaGo
because
the games, we know the model. We know
how the moves work.
And in math, we know what the operators
are.
We you know, we know lean will take us
from one state of knowledge to the to
the state of the proof to the next
state.
But if we have to learn the models,
there are no
I'm I'm going to say it. It's probably
maybe a a weird example, but
a counterexample, but I can see that
there's no instances of learning the
model and then planning with the model
in our field.
>> At least not with
uh uh,
like self-discovered
abstractions. So,
there are people who say, "I'm just
going to learn a model of what happens
in the next second or next millisecond."
But, that's not how our models work. Our
models are more abstract. Our models
are, uh,
quite different.
>> So, one of the things I like about what
you're doing here is you're not just
sitting around pontificating or
lamenting the state of the world as it
is. You're very action-oriented. It's
why you started a company. So, let's
let's start talking about that a bit.
Uh, in 2022,
Rich, you laid out a very specific
12-point plan, the Alberta plan for AI
research.
Maybe tell us about that.
>> So, the Alberta plan came about because
we just have general ideas, but we also
needed to convert them into smaller
chunks.
And so, the 12 steps
are the attempt to uh, crystallize
particular chunks.
There's a very important early step,
step two,
uh, which is
continual deep learning.
And And we think that one is like almost
the most important because it unlocks
everything else. If you could do
continual deep learning,
you could then continually update your
model of the world.
And then, if you knew how to do the
abstraction rights in like
the second half
of of the the steps are all about how to
get the abstractions right.
So, and not only my abstractions right,
what I mean by I don't mean get the
right abstractions cuz no one can say
what the right abstractions are. That
depends on the world that you're in.
Your agent would have to learn the
correct abstractions for whatever world
it's in.
And so, you know, if you maybe those are
the two key things. You have to find the
right abstractions, and then you have to
do able to continual deep learning.
>> I think that a lot of the people in the
field realize that we need models, we
need to plan with them. But, the
abstractions tell us what the model
should be conditioned on. So, what
should you What should the model
predict? What are you going to do and
then something is going to happen. And
more importantly, where would that come
from? So, I really like the example of
elite athletes. If you ask elite
athletes about how they do certain
things, they would have weird niche
terminologies for doing very specific
things. They were like, you know, I do
this thing and they would have a name
for it. If they communicate, sometimes
they don't even have a name for it if
they're just doing it alone.
So, how did they come up with those
abstractions? That's in some sense a
crucial thing that's missing that the
later half of Alberta plan answers.
>> Can we talk about the continual deep
learning part?
Is it an algorithmic gap that exists
today or is it a just a practical
deployment infrastructure data privacy
gap? Because if I wanted to do
call it naive updating of weights based
on user interaction, I can do that
today, right? And so,
what in your opinion is the biggest
thing that we're missing to kind of get
to continual deep learning?
>> Yeah, so it's absolutely an algorithmic
gap. You You can do the naive thing, but
then you'll see all sorts of problems.
So, for example, if you say um I'm going
to take one sample and then I'm going to
update my whole model with that one
sample. You will run into this problem
that now all of the previous knowledge
in the model it's impacted negatively.
And the way currently we we
get around this is exactly what cursor
does. They don't use one example. They
use a large batch coming from a lot of
users. So, in use cases where you can
have that, you can do continual
learning. But, most use cases you don't
have that. Most use cases you have a
single stream of data. And then if you
apply it to the naive thing, it just
completely destroys your prior knowledge
in a very
um
destructive way.
>> Catastrophic forgetting.
>> Yeah.
>> That is the Yeah.
>> But, it's totally curable. You have to
[laughter] have the right algorithm.
>> cure?
>> Well, yeah, exactly.
>> What is the cure?
>> Well, you know, first you need to do
what we call step size optimization.
And it means every weight in your
network has to have a separate step
size. So, some will move fast, some will
move slow. And you will we you will have
to metalearn the step sizes for each
weight. Most of your network will be
have have weights that have tiny step
sizes. So, then when you train on a new
example, they don't get destroyed.
Happens just to the right places.
And then secondly, you have to use some
form of generate and test.
Um
which is
in feature space. So, you come up with
new features or new units and and
without following gradients. Cuz
gradients are very slow process. You
only move in a direction if you know
it's the helpful one. And that's always
going to be very slow and doesn't give
you a path to grow more and more complex
and and to have sustained learning. You
need to have something that just
proposes a bunch of new units.
And
and and then goes from there. I guess
so, there is a specific thing I can say
that make it at least concrete, which is
to say we have this algorithm called
continual backprop. We used published in
in nature a couple years ago. And it it
is exactly like backprop, but every but
you also plant new seeds of units that
are newly initialized with random
weights.
Backprop only has random weights at the
beginning of time. And then as you go
on, all that randomness all that variety
from the randomness gets used up.
And with continual backprop, we keep
injecting a bit of randomness, a bit of
generate and test, a bit of generate and
then the
the operation of backprop is the tester.
So, you need you need that.
And and if you put those together really
well, I think you'll have a new
generation of massively
superior
continual deep learning.
And that's what we hope to do in the
next couple years.
>> Wonderful.
Do you think that these algorithms
can be applied to the current state of
affairs with people scaling LLMs and
trying to get them to do continual
learning without catastrophic
forgetting?
>> Yeah, absolutely. I think it's
So, I don't think that you could take an
existing model and say I'm going to just
start updating it with these algorithms
because these algorithms meta learn how
to learn. So, really you have to say,
I'm going to learn from scratch. So,
let's say I learn a new foundation
model, but I'm going to learn with this
these new algorithms. These new
algorithms in addition to
learning the knowledge, they're also
going to learn how to learn future
things. So, they're learning two things
at the same time. And then um then I
think you would be able to learn new
things without catastrophic forgetting.
>> the most radical thing in that you're
trying to do in your company in terms of
from the current state of affairs
to try to do these two things at the
same time?
>> Most radical thing.
>> Is I think that's
>> This goes back to I'm not crazy,
everyone else is crazy.
>> [laughter]
>> Yeah.
>> That's perhaps not totally radical.
There was a point
in like 2016 to 2018 where a lot of
people were exploring these ideas quite
a bit. They were doing it
in a much more limited setting. So, they
would say, we have a distribution of
problems and then in this specific case
we'll do it whereas we want to do it
from a single stream of experience. So,
our method should be more generally
applicable. So, I think many people have
explored this,
but no one has explored this in the
general setting where the resulting
algorithm would be applicable
everywhere.
>> So, what would be the most radical thing
that your company your new company is
trying to do that other people are not
doing?
>> What's the most ambitious thing?
Remember, I don't think I'm weird, so I
don't want to say it's radical.
>> the most ambitious
>> ambitious thing, I think is to try to
have the full spectrum of knowledge
both about the tiny things and about the
big things. You know, like thinking
about how you take an airplane from one
city to another. That's a a very big
thing. You know, it's it's more it's
like your
your space flight example, but it's just
kind of more common sensical to think
about cuz we all
many of us take airplanes and any all of
us use abstractions on all kinds of our
life. And even even the squirrels use
abstractions. So, to have that spectrum
of of knowledge
uh
from the small to the big and to treat
it in a uniform way and
to be able to help have it self uh
maintaining. You know, the big question
is always you have your knowledge-based
system and what keeps the knowledge in
it correct?
Well, what keeps the knowledge correct
in a large language model is well,
people did a lot of post-training
and and they they they made it sure it
was correct. And then they freeze it
after that. So, that's
what keeps it correct. But really our
minds, we are we're always changing
things and yet something keeps it
organized and coherent and and and
settling back into a good place rather
than drifting off into crazy land.
That is an I think our our biggest
ambition to have a mind that is
self-consistent
and and can
keep training itself and making it
coherent.
>> that.
Can I ask?
It almost seems that it's it's such a
ambitious vision and the idea that all
these things can be unified into a
single mind is so ambitious.
>> It's within reach.
I think it's within reach. It's here
it's 2026 and our computers are so fast.
You know, is it is it
so ambitious that it's out of reach?
Uh or or do we have
already inklings of how all the steps
can be done? And we
I I think we have
a vision
and inklings.
Um I don't think it's I don't think it's
inappropriate.
>> Your vision involves a trillion
parameter model with 20 watts.
That seems pretty
ambitious.
>> That is ambitious. Uh, in some sense
with current technology, I would say
it's also impossible.
Like just storing a trillion parameters
in memory would probably use more than
20 watts of energy with current memory
technologies, but
we are really thinking of okay,
things are getting better, computation
is getting cheaper or it is getting more
energy efficient. So, where would be
would be in 5 to 10 years? And I think 5
to 10 years with the right algorithms
and
we can totally be in a world where this
would be possible.
>> So, 5 to 10 years is two orders of
magnitude of Moore's law.
It's a standard improvement.
If we double every 18 months, 10 years
would give you two orders of magnitude.
And so,
for Kurzweil's statement to be
plausible,
then today you should be able to do it
for
for what? 20 watts? Two orders of
magnitude?
>> 2,000
>> 2,000 watts. If you can do it with 2,000
watts today, yeah, then in 10 years
you'll be able to do it for 20 watts.
>> You think you can do it for a 2,000
watts? You have I think lots of people
at research labs that have access to way
more than that.
>> Yeah, I think we can
can be more efficient than that even now
with the right algorithm.
>> If we can be more efficient than that,
then why aren't we? It's not like people
just want to spend all their money on
spend all their money.
>> Sometimes it seems like they
>> I'll pay money.
>> Sometimes it [laughter] seems like they
want to.
>> Yeah, I I
>> Doesn't it?
I think that's how they show they're
they're real men
by using lots of energy.
>> At least when I look at different
research groups, I don't even see anyone
believing in that it's possible. And I
think if you don't believe in it, you're
just not going to work on the technical
problems and work through them.
>> Is it that it's not possible or it's
that there's so much waste in the
system? Like one which one is it? Like
is it is is there someone who knows how
to do it efficiently?
>> Yeah.
>> And then there's 10 times the number of
people in the same lab doing all these
other things.
And
so nine out of 10 people are wasting
>> In some sense I the way I think about it
is that we are stuck in a local minimum.
So, if we want to move towards these new
kind of algorithms,
it is almost impossible that things will
not get worse before they get better.
So, when we start exploring these new
directions, you're not going to get
state of the art performance from day
one, but it is because it is a different
paradigm. Um but that path leads to
similar performance at a higher energy
scale. And these big labs, they are so
locked into a product that they
like it is not possible for them to
pursue a path where things get worse
first.
>> Because their current paradigm allows
them to keep scaling and this new
paradigm they have to take a bet. And
then
>> And they have to figure out some some
technical things that are difficult that
we have thought about it for many years.
We know people who have thought about
these things for many years and
when I talk to them, it makes sense that
it's doable, but
you need to think about those challenges
for a long period of time.
>> So, if everything goes right with Oak,
what happens with the company? What do
you what kind of company are you
building?
>> Uh if everything goes right, we
uh implement the architecture, we can
have a genuine uh continual learning and
we can form abstractions so that we can
do planning and reasoning and
and we have so sort of like true
intelligence. And then, you know, it's
hard to imagine just exactly what will
happen by then.
But I think
>> Humans will become irrelevant.
>> I I don't think that's true at all. I I
I
>> We don't either.
>> I think the world becomes exciting and
even more exciting and interesting and
and for humans.
But in particular, I think there are the
You have to wonder about the large
language models.
They might be at risk.
Uh when this eventually happens, you
know, I'm sure they'll get a good run.
They've already had a good run. You
know, they've been very successful. And
let me say,
just for for clarity that uh large
language models are an amazing
scientific breakthrough, a breakthrough
in the skillful use of language by
neural networks, wholly unanticipated.
You know, it was a It was always a hold
out for uh symbolic methods in language,
and they
They have totally changed how that's
thought about now. Yeah. It's a It's a
big breakthrough.
It's so It's frustrating to me that we
have to you know, just celebrate that
we've made this great progress in the
sub subset of the problem of AI, and
enjoy that. Instead,
it has to pretend to be all of AI. All
of intelligence is not fluid, capable
use of language. There's so much more.
It's an important part. You know, it's
like
20% or a quarter of
intelligence. There's There's more.
>> Yeah.
>> We're not done.
>> Yeah.
If everything goes right, are you
imagining that there's a single minds
that can do everything from learn how to
swing from tree branches to make a
spaceship,
uh to you know, all these various things
we've talked about today. Is it a single
mind, and is it a single set of weights
that can that can do all these things,
or is it
>> It's a single design.
>> Okay.
>> And there'll be many different There'll
be many minds.
>> Mm. Okay. So, it's a single design that
reacts to different environments.
>> And it did it And different versions of
that mind would learn different things
because they have different experience.
This sort of goes back to the big world
hypothesis that there are infinitely
many things to learn, so one system
cannot learn infinitely many things. And
like I think Rich already mentioned
this, but if you have two of these
systems, if you have two of the largest
systems in the world, then it is trivial
that they cannot model each other
because they are equally complex. So,
single system would never be able to get
to a point where it can learn
everything. It would always be multiple
systems
that are learning from their own
experience.
>> Guys, are you hiring? What kind of
people are you looking for?
>> We are hiring.
The initial team, most of it we already
have in our mind. So, these are people
who have thought about these ideas in
the past in many
uh you know, for a long time.
And
we are going to take a slightly
different approach because this is a
different paradigm. It doesn't make
sense to
become large very quickly because uh in
some sense, everyone we hire has to
uh come to see what we see. And not
everyone sees that. So, we're going to
start small, slowly grow to maybe a
handful or two or three uh and then go
from there.
>> We want to be super aligned.
>> We want to be super aligned.
>> So, that we can be um
very productive working together and and
scaling the progress.
>> Absolutely.
>> Very, very cool.
>> Wonderful.
I love this conversation. Thank you for
taking the time to share what you're up
to. You are, you know, an an unusually
deep thinker about
where reinforcement learning and
algorithmic design will go.
And it was a true pleasure to get to
explore it together with you today. So,
thank you.
>> Thank you very much. Thank you. It's our
pleasure.
>> [music]
[music]
[music]
Ask follow-up questions or revisit key timestamps.
In this video, AI researcher Rich Sutton and his co-founder Quorum Javed discuss their work at Oak Lab, focusing on the concepts of reinforcement learning, 'The Bitter Lesson,' and the necessity of 'continual learning.' They argue that intelligence arises from agents interacting with and learning from their environments over time, rather than relying solely on pre-trained models or static human knowledge. The discussion highlights their 'Big World' hypothesis—the idea that the world is infinitely complex and that systems must learn from their own experiences rather than being bottlenecked by human-curated data or simulations. They outline their research agenda, which includes developing algorithms that enable continual deep learning, self-discovery of abstractions, and efficient weight updates, all with the goal of creating more robust, self-maintaining intelligent systems.
Videos recently processed by our community