HomeVideos

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Now Playing

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Transcript

1384 segments

0:00

If I play football, for example, it

0:02

looks very very closely to reinforcement

0:04

learning. I get all a lot of times and

0:06

every time I adjusted a little bit and I

0:08

see if it roughly matches what I what I

0:10

wanted and there are there are there are

0:11

some self-reinforcement happening. When

0:13

I learn mathematics, it's a very

0:15

different type of thing. It's it's like

0:17

reading

0:18

about hard concepts and thinking about

0:20

them very deeply inside my head until

0:23

things click and I until until I have

0:24

them connected. And both of those in

0:26

some way are learning from experience.

0:29

They are just very different. We

0:31

probably are spending the most compute

0:32

than ever on learning from experience,

0:34

but there reinforcement learning is is

0:36

not the end of learning from experience

0:37

and there will be better approaches that

0:39

researchers will be coming up in the

0:41

coming [music] years on how to how to

0:42

use that data.

0:55

>> [music]

1:00

>> Jerry Rohan, thank you so much for

1:01

joining us today. The two of you are the

1:03

founders of Core Animation, one of the

1:06

hottest new labs in San Francisco right

1:07

now. And before starting Core Animation,

1:10

you led some of the most important

1:12

research projects of the AI era. Jerry,

1:15

you were VP at OpenAI, where you worked

1:18

among amongst other things on running

1:20

the strawberry and reasoning teams. And

1:23

Rohan, you were one of two of the four

1:26

pre-training leads at Gemini. And before

1:28

that, led a lot of the fundamental AI

1:30

research at Google Brain and were the

1:32

fix-it guy across Google and then at

1:34

Anthropic. And so, between the two of

1:36

you, you've seen more than a fair your

1:38

fair share of

1:40

what the world looks like in terms of

1:41

doing frontier research. And so, I'm

1:43

very very excited to dig in. Um, let's

1:45

start with you, Jerry. You tweeted a

1:47

very spicy take recently. The first step

1:50

to replacing transformers is

1:52

appreciating deeply how far they were

1:54

able to carry us. Is that a eulogy for

1:56

the transformer? What does that mean?

1:57

>> Thank you very much for inviting us

1:59

here, Sonia. I feel like a lot of my

2:02

interviews these days is explaining my

2:04

tweets. And what did I

2:06

>> [laughter]

2:06

>> What did I mean?

2:08

But appreciating Transformer means like

2:11

understanding what it does well. So, if

2:13

you're not solving the problems that it

2:15

is solving well, you have to focus on

2:16

its on its weaknesses. You have to You

2:18

have to understand good parts and bad

2:20

parts. And it's very easy in a lot of

2:22

the work what people are doing in

2:24

architectures is trying to make

2:26

Transformers cheaper. They are trying to

2:28

make Transformer more efficient. I I

2:30

very rarely see people thinking about

2:32

how do we make Transformers more

2:34

powerful, trying to do more express. But

2:36

But like seeing someone's weak parts and

2:38

seeing someone's strong parts are almost

2:40

almost almost the same thing. It's just

2:41

just understanding the shape of

2:43

Transformer a little bit more.

2:44

>> What I think right now we are in this

2:47

stage. We got to really really good at

2:49

training really really big models. We

2:51

mastered two algorithms. We mastered

2:54

pre-training at a large scale. And we

2:56

mastered reinforcement learning at a

2:57

large scale.

2:59

And I'm asking myself a lot what is next

3:02

in machine learning. I think at this

3:04

moment what the bottleneck is to better

3:07

models and to smarter systems is the

3:09

architecture itself. It is this moment

3:11

to revisit uh the train we've been

3:14

riding for the last 6 years of trying to

3:17

add more and more parameters to

3:19

essentially two of the same

3:21

operations, which is MoE and and

3:24

attention. And when I am when I'm

3:26

thinking about it, like what's what like

3:29

where we are today and what we are

3:31

doing, I am thinking a lot about about

3:34

like what Codex and what Cloud Code are

3:36

doing for us. And I'm really really

3:38

appreciative of those systems and of the

3:40

coding and of the workflow automation

3:42

and of the systems and of the products

3:45

that we have today that we essentially

3:46

have built over those those 6 years of

3:49

scaling. And I think I think this is

3:50

this is the first step of thinking.

3:52

Like, what is the

3:53

If we want to work on the replacement,

3:55

we need to like see where we are, what

3:57

problems we have solved to like start

3:59

seeing what the next stage is, what

4:01

problems we haven't solved yet, what

4:03

what kind of are we missing. And this is

4:05

this is kind of whenever whenever I use

4:07

Codex and I am successful at a task, I

4:11

also start thinking, why did I why

4:12

didn't I try to push that thing harder?

4:15

Whenever whenever I come to work, there

4:17

are a lot of things I do with Codex, but

4:19

I still come to work. I still ask it to

4:21

do certain things for me. And I'm always

4:24

asking myself, why am why am I even

4:25

needed there? Why is uh core automation

4:29

is name and its concept is we want to we

4:31

want to be automating tasks. And why why

4:34

why why are those things not yet

4:35

automated? Why why why is not Codex

4:37

doing everything for me? And this is

4:40

this is the question of like the

4:41

research, where we want to go. And with

4:44

that research, I'm trying to think, what

4:45

kind of models, what kind of what kind

4:47

of systems do we need, what kind of

4:49

qualities do we need what that that that

4:51

we don't we don't have today. And that's

4:52

that's what I'm thinking a lot these

4:54

days.

4:54

>> And you have the starting premise of the

4:57

architecture is the issue, which I think

4:59

is a contrarian point of view. So,

5:00

what's what led you to that point of

5:02

view? What did you see that made you

5:04

think the architecture was the issue?

5:05

>> It's fundamentally what is the issue.

5:08

It's it's it's like comes back from that

5:10

from from the previous implication. What

5:11

I I think is the issue is that the

5:14

models are being trained in the lab and

5:17

are being deployed in the in the real

5:19

world. That is that that is the

5:20

fundamental fundamental tension that is

5:23

that is there.

5:25

And mm

5:27

like

5:28

a bit of a bit of my disappointment came

5:30

comes from my my my personal story.

5:32

Whenever we were starting uh the

5:34

research and progress on scaling up

5:36

reinforcement learning at OpenAI, I

5:38

basically believed that scaling up

5:41

reinforcement learning is a necessary

5:43

stepping stone on a path to AGI

5:46

uh since since since I started like uh

5:48

working working at Open AI. And I was

5:50

always reinforcement learning

5:52

maximalist. I was always believe this is

5:54

what we need to focus on. This is what

5:55

we need to do. I've seen LLMs being

5:57

scaled up to

5:59

to to higher and higher levels through

6:01

GPT-3 to GPT-4. And we're still doing

6:04

very little RL. And I had this this

6:06

internal belief that the moment we start

6:08

scaling up RL we'll we'll we'll solve

6:10

everything. We'll we'll we'll be able to

6:12

solve solve all the problems. And we

6:15

eventually started

6:17

scaling up RL. I was I was just in

6:19

there. I was I was in the center of it.

6:20

I was thinking here we are. If you ask

6:22

Jerry in 2024, when do we get AGI? I

6:26

would say 2025 will be will be that

6:28

year. This is

6:29

this is where we solve everything. And I

6:31

saw us training model after model. This

6:35

model was getting better and better. All

6:36

the benchmark scores were going up. And

6:40

did we also solve all the real-world

6:44

tasks at that moment? Unfortunately,

6:46

unfortunately not. We we we we still

6:48

have work. And I realized there was this

6:51

bit of

6:53

distinction as all the benchmarks that

6:56

we are evaluating our models, they were

6:58

essentially the same thing as we were

6:59

training the models on. Like all the

7:01

evals and training side tasks are the

7:03

same sides of the coin. But the

7:05

real-world distribution and real-world

7:07

task is much messier, much murkier, much

7:09

more much more different. Our our

7:10

training data didn't really replicate

7:13

the real-world use cases. And despite us

7:16

basically maximizing all the task, if

7:18

you see ask anyone training models,

7:21

"Hey, what is one of your main issues?"

7:23

"I don't have hard enough tasks. I don't

7:25

I don't have what to what to train our

7:26

model on." Yet we are still not covering

7:29

the the entirety of the of the

7:30

real-world distribution. From that, my

7:33

conclusion is we need to have models

7:35

that learn at test time. We need to have

7:37

models that learn with users on their

7:39

data, on their real-world task, on the

7:41

real world distribution. And

7:44

there when you ask why why why why why

7:46

why don't we have that today? Why why

7:48

why why are transformers not not

7:49

learning any anywhere? And there are

7:52

essentially two types of learning that

7:55

we could be doing at test time. We could

7:57

be doing in-context learning essentially

8:00

of transformers, which is

8:02

it it it doesn't have

8:04

fundamental problems of catastrophic

8:07

forgetting. It doesn't have that issue.

8:08

It is pretty data efficient. So that

8:10

that is great, but it's not very

8:11

scalable. We only can have so much of

8:14

it. It is limited and it has some more

8:17

even more of mechanical limitations of

8:21

what actually are you doing when you

8:23

when you build context? But maybe maybe

8:25

we can we can come back to it later. But

8:28

we have we have in-context learning,

8:29

which is very limited and very very

8:31

small amount of data. Whenever I'm using

8:32

Codex roughly after around 20 minutes of

8:35

usage, I need to I need to compact it

8:37

and move and move it afterwards, which

8:39

is not that much not that much data. If

8:41

we have all all we can learn is for for

8:43

20 minutes, it's not that much. And the

8:44

second thing is

8:46

fine-tuning. We could try to

8:48

continuously fine-tune our models, but

8:50

then those have the issues of

8:51

catastrophic forgetting. We have issues

8:53

of very low data efficiency. And neither

8:57

those are very solvable. Neither of

8:59

those are very easy to

9:01

uh to find ways. People have been

9:03

trying. If they was were easy to solve,

9:05

someone already solved it. So

9:08

my personal belief is we need to find an

9:12

algorithm that we can we can meta learn,

9:14

that we can express on the architectural

9:16

layer, that can represent how does how

9:19

does learning look like. How does

9:21

learning look like that can work on much

9:22

much longer horizons.

9:24

>> Mhm. Do you expect the architecture will

9:26

look transformer-like? Cuz my my my

9:29

understanding from the from the chief

9:31

seats is that you know, OpenAI have been

9:33

trying to scale up reinforcement

9:34

learning for a long time, and it wasn't

9:36

until the transformer came about that it

9:38

seemed like there was an even kind of

9:41

scalable prior on the world upon which

9:43

to even scale RL. And so, how do you

9:46

even go about trying to think about

9:48

scaling up this this new resume?

9:50

>> Yeah, it's a it's a it's a great

9:51

question. I think those two things

9:53

happened at the same time, but

9:57

if anything that that happened there was

9:58

was mostly about about economics,

10:00

because technically you can scale up

10:02

LSTMs. Just no one no no one really

10:05

really dared to go in that direction.

10:07

And they did scale much much more

10:11

poorly. They are they are they are

10:12

scaling in a scaling loss paper there is

10:15

presented a comparison of LSTMs and

10:17

Transformers. And fundamentally scaling

10:19

law of Transformers was better. There is

10:21

a world where we never invented

10:23

Transformers and we would be scaling

10:24

LSTMs and we would be having some

10:26

models. But because they would be much

10:28

more expensive to train and much less

10:32

impressive as a product, we would have

10:34

we would have just this worse experience

10:36

and maybe no one would be able to

10:37

convince people to spend as many dollars

10:39

training that those gigantic LSTMs

10:41

because we wouldn't get a market return.

10:43

The the the majestic thing about

10:45

Transformer, which goes back to like why

10:47

why why why do we have to appreciate

10:48

Transformers so deeply, is that

10:50

Transformers are economically valuable.

10:53

The training them the cost of training

10:55

them is lower than the revenue that they

10:57

that they generate, which is which is

10:59

magic of machine learning and it's not

11:01

guaranteed by itself. But for for LSTMs

11:03

it probably wouldn't be that way, which

11:05

which which may have happened. But you

11:06

can in many ways you can scale most of

11:09

the architectures.

11:11

I think I think a lot of reasons why

11:12

people didn't scale things before was

11:14

because researchers before OpenAI had a

11:17

lot of reluctance to scaling. It was

11:20

often seen as unscientific and

11:23

research in algorithm was was providing

11:25

how do we become more and more

11:27

efficient? How do we for the same

11:28

compute budget get better better

11:29

results? And it was it was a bit of a

11:31

bet by OpenAI at that moment. Try to

11:35

say, "Hey, we don't care about better

11:37

and better algorithms. We care about

11:38

smaller and more scalable algorithms."

11:41

And how do we how do we pour more and

11:42

more compute and get better and get

11:44

better results, which OpenAI was

11:45

criticized repeatedly by many people in

11:48

the community for for a long time. But

11:49

thanks to that, we have we have the

11:51

models that we that we have today.

11:53

Uh and I think there are tons of

11:54

architectures that can be scaled up. And

11:58

I am part of

12:00

the core automations mission, and our

12:02

our belief is that a lot of

12:04

architectural research happened at too

12:06

small scale for

12:07

too long time. A lot of people are

12:10

trying to say, "Hey, let's let's try to

12:12

first try our architecture on a on a on

12:15

a small data set on a small in a small

12:17

computer regime, and then and then and

12:19

then see where where where where where

12:20

where it scales only after after you

12:22

prove itself." But for example, when you

12:25

do work on reinforcement learning, you

12:27

know that to get to any interesting

12:29

results, you only you need certain level

12:30

of compute to even see the capabilities

12:33

in a model. Reinforcement learning needs

12:35

a baseline of of of ability to only to

12:37

only start working. So so where I am

12:39

coming from, probably there are many

12:40

architectures that need a baseline of

12:43

compute to even start doing anything

12:45

anything interesting and anything

12:46

useful.

12:47

>> Can I ask you then maybe a touchy

12:48

question?

12:49

>> Please do.

12:49

>> If you need a baseline of compute, that

12:51

sounds like a job that would be well

12:53

served inside of a big research lab. Why

12:55

start a company to go do this?

12:57

>> [laughter]

12:58

>> It's a it's a it's a great question, and

13:00

it's I think in in in many ways it's

13:03

likely a timing thing, timing issue.

13:06

Market is right now in a in a very

13:08

specific place where

13:11

the biggest and the most successful

13:13

labs, by coincidence or by fate, are

13:16

probably in the most competitive market

13:20

fight ever right now, which makes them

13:23

not very keen on trying different paths,

13:27

trying alternatives. If Transformer is

13:29

profitable and if you can spend more

13:31

efforts and more resources scaling

13:33

Transformer to win in the next quarter,

13:36

it's very hard to put at least a lot of

13:38

attention and a lot of energy to work on

13:40

something that will that will maybe

13:42

better or maybe or maybe will will

13:45

redefine the field in a year or two. So

13:47

so I think the big biggest labs and I

13:50

talked to basically all of them don't

13:53

don't have that much interest in trying

13:55

the alternatives to to Transformer. And

13:58

the labs that are not the biggest are

13:59

doing whatever they can to do what they

14:02

what the what the what the most

14:03

successful labs are doing. And everyone

14:04

is trying to trying the same the same

14:06

coding agent. If you look at the

14:08

last week's releases, everyone everyone

14:10

is is trying to release a coding agent

14:12

right now. And and I think

14:15

we need different paths and different

14:18

and different approaches here. So so

14:19

that's what that's the niche in the

14:21

ecosystem we are trying to trying to

14:23

fill in.

14:24

>> Mhm. And you were you were at

14:25

you were at Brain when the Transformer

14:26

was invented. Do you Do you agree with

14:28

Jerry's eulogy for the Transformer?

14:30

>> Yes. Um in some sense like um once the

14:34

first when the Transformers Ashish Noam

14:36

and others came up with it I I had like

14:38

work I worked on my work on online

14:41

distillation around the same time. We

14:42

presented it at the same internal

14:44

research conference. It wasn't a big

14:46

deal internally. There was only few

14:47

people who actually got it.

14:49

>> Mhm.

14:49

>> A lot of few people were like, oh, it's

14:51

it's like, yeah, it's another work. And

14:54

people were finding

14:55

ways to And it was also very focused on

14:58

at least the original work was very

15:00

focused on a real problem, which was

15:01

translation. So they solved like they

15:03

beat LSTM on translation. And it took

15:05

opening up I mean, internally at Google

15:08

there was definitely like Noam and a few

15:09

others were definitely interested in

15:11

scaling language models. I think it is

15:13

until GPT-2 and GPT-3 that we saw the

15:16

benefit of Transformers working quite

15:18

well. At least the way I think about

15:20

architecture is how do we spend

15:21

computation and transformer is one way

15:24

very efficient way to spend computation.

15:27

But now that I look at the industry,

15:28

it's it's a lot of our computation is

15:31

inference time and spending it on

15:32

tokens. Let me ask the question like if

15:35

I want to optimize for a better

15:36

architecture, I want to look at both

15:38

pre-training and RL together and I would

15:40

like to find architectures that spend

15:42

computation much better than current

15:45

chain of thought token generation. Uh in

15:48

at a like to give a much better

15:51

overview, it's I think of like

15:52

pre-training is built the transformer

15:54

with certain context length and RL comes

15:57

in and it's like, well, that's not

15:58

sufficient. I need more computation. Let

16:00

me do it via adding one token at a time.

16:03

This is quite inefficient from uh like

16:06

inference perspective. You're doing one

16:07

token at a time, so most of the

16:09

solutions have been finding to do better

16:12

ways of speculative decoding. So it's

16:14

like a band-aid to a problem that we've

16:15

picked something that's can only

16:17

generate one token at a time. So

16:19

autoregressive decoding. Uh there is

16:21

problems with the transformer in terms

16:23

of how do we spend the computation? Uh

16:26

for all the longest time, I think uh

16:28

most of the world was training very

16:30

large dense models and it took uh like

16:33

the industry like roughly two to three

16:35

years to get to refine the architecture

16:38

to what we now take for granted was not

16:40

obvious to a lot of people. Sparsity and

16:42

the mixtures of experts and getting good

16:45

training efficiencies with them, right?

16:47

So then you can ask like what's wrong

16:49

with the transformer? Well, it's it the

16:51

computational depth is poor. You how do

16:54

we increase computational depth? And

16:55

just posing that question opens up like

16:58

20 new directions on how we can modify

17:00

the mechanism to incorporate it. So um

17:04

I see like to do work like this, it

17:07

takes time and usually like fundamental

17:10

research in the past have taken like

17:13

five, six years to land into industry

17:16

and it's largely from organizational

17:19

knowing that it is important. This is

17:21

the bet. Like just like Jerry had the

17:23

inner belief that RL is needed.

17:26

Absolutely do not have that belief at

17:27

Google. I was a pre-training maximalist.

17:30

Pre-train a bigger smaller

17:31

>> are a good fit.

17:32

>> Right? Like

17:33

all right. And then so that inner belief

17:35

and second is you need your architecture

17:38

to run efficiently on hardware. A

17:40

theoretically optimal architecture is

17:41

not useful to anyone. It is something

17:44

when it comes into practice. So you need

17:46

the research inception to getting it

17:49

productionized and getting kernels and

17:51

everything written. The end-to-end loop.

17:53

And there's only few places right now

17:55

which have integrated teams doing that

17:58

and I think we have built a team in a

18:01

way that puts every like the experts

18:03

together not in different silos that

18:06

like we are accelerating on having

18:08

everybody look at the problem

18:09

holistically from end to end. So that's

18:11

like where I'm quite bullish. That's why

18:13

I'm here. Um the current mechanisms are

18:17

quite poor and it will if if you leave

18:20

it to the world, I am afraid that it'll

18:22

take us like a much longer time horizon

18:24

before we replace the transformer and I

18:27

think a lot of folks are already

18:28

complaining on

18:29

a lot on token costs. And that seems

18:32

like as someone

18:33

>> complaining.

18:34

>> Uh

18:35

I mean

18:36

in terms of like yeah, exactly.

18:39

I come from the Google mindset where we

18:40

had to like serve billions of people. So

18:42

like finding more efficient

18:43

architectures that fit have a like a

18:45

deadline on latency and the number of

18:47

tokens that you can serve. So like when

18:50

I look at that like the amount of like

18:51

the world that can use like frontier

18:53

tech is very little and we someone or

18:57

some group has to

18:58

like accelerate and make this better and

19:01

we are taking that shot at doing that.

19:03

The current technology just scaled up is

19:05

still only relevant to like a subset of

19:08

humans.

19:09

>> Yeah.

19:09

>> And this is the bet we're making.

19:11

>> So one of the things I heard you say was

19:13

the problem with transformers is the

19:15

computational depth is poor.

19:17

>> Yeah.

19:17

>> If that's the crux of the issue, tell us

19:19

what does that mean?

19:21

Why is that the case? How do you fix it?

19:23

>> I can

19:25

give you like one insight like most

19:27

transformers that we train are quite

19:29

shallow. It's at most like 100 layers

19:32

deep.

19:33

Depth is like it's called deep learning

19:35

cuz you want a deeper representations.

19:38

There has been experiments on going with

19:40

no one has actually shown

19:43

us learning extremely deep

19:44

representations. Chain of thought

19:46

reasoning and RL to do chain of thought

19:48

on by model itself is one way to

19:51

increase computational depth because

19:52

every token you add you add like one

19:54

more pathway. And there has been

19:57

So, then you can get out of like this

19:59

bottleneck that the pre-trained

20:00

architecture has set you up on. You can

20:03

only do layer number of layers times

20:06

sequence length. Now, you can increase

20:08

the sequence length and you get much

20:10

stronger results. You can do inference

20:12

time scaling. Now, the issue with

20:14

inference time scaling is that models

20:16

now have to produce more tokens to get

20:19

better results and that's very one token

20:21

at a time. And from this you can see

20:24

like you can directly address many of

20:26

these things and this is like a subset

20:28

of work that we're looking at, right?

20:30

Making this much more efficient.

20:32

>> What's your forecast for the

20:33

transformer-based architecture?

20:35

If it's not the end state, how far can

20:37

it get us? When do we start to see it

20:40

topping out?

20:41

>> I I I

20:41

think it all it all comes back to what

20:44

we are what we are training transformers

20:46

for and what we can what we can do with

20:48

them. We're doing pre-training which is

20:51

very good at distilling all the

20:52

knowledge from the internet into

20:54

transformers and then we can RL them

20:56

through which is we basically can bake

20:58

all the workflows that we want into a

21:00

transformer. So, what what what where

21:03

transformer has like capped out is we

21:05

have all the knowledge of of of humanity

21:07

in the model together with the

21:09

relationships and how do they how do

21:11

they work together, how we how they can

21:13

be combined. And basically any task that

21:16

we have training data for, we can we can

21:19

put into the model and we this can be

21:21

gigantic model trained with a lot of

21:23

compute on all the on all the all the

21:24

data in the world. And then if we ever

21:27

stop training that model, what would

21:29

what would happen? Like the the question

21:31

worth asking often and thinking about

21:32

transformer, what would happen if OpenAI

21:35

and Anthropic stopped training new

21:36

models and we got we got the transformer

21:38

we have today and say this is this is

21:40

it. This is this is this is the best the

21:41

best model we have. Uh months pass,

21:45

uh year year year years pass and the

21:47

model is getting less and less useful.

21:50

Maybe the lab really like

21:52

recorded of every human on earth what

21:54

they were doing and what their what

21:55

their tasks were and their environments

21:57

and put them put them in the model, put

21:59

them in the in the in the learning

22:00

environment. But then that that then

22:02

then then what happens if the if

22:04

anything of that stuff changes? If there

22:07

are new events in the world, if those

22:09

new events have new relationships

22:11

between them,

22:12

if there are there are new types of

22:14

tasks, if there are new code bases, new

22:16

tools to use, transformers are getting a

22:19

lot of their usefulness and all value

22:21

through the things that are that are

22:23

valuable have to be present in in in in

22:25

in in training. And when when when when

22:27

when they are not, they they they they

22:29

suffer. There's there's some ability to

22:31

adapt, but it's not very not very big

22:33

and not very not very flexible. So in my

22:35

mind this is this is kind of the level

22:37

where the where the transformers uh top,

22:40

which in many ways what I think is a is

22:42

a tool to use for us. If we if if we

22:46

kind of know if there ever there's human

22:47

who knows the limitations of a

22:49

transformer, they can they can schedule

22:51

that model, they can they can write a

22:52

prompt of what is what is it the task

22:54

that you want. And by doing the training

22:56

we are doing, you can you can like get

22:59

very successful in that and any task the

23:01

model fails, you get added to training

23:02

data and the model and the model can can

23:04

succeed. But that that loop is has to go

23:07

for the lab training training the model

23:09

for you. And if the model that that

23:11

fundamentally needs to be trained in the

23:12

lab, like how much do you think of it

23:14

that this is this is the goal or or or

23:17

you would want to be able to update the

23:19

model somehow not having to to go back

23:21

there.

23:22

>> Have you read the Rich Sutton and David

23:24

Silver have this paper the the age of

23:26

experience? Have you read it? I'm

23:27

curious how much you agree or if you

23:29

have any different opinions where your

23:31

opinions diverge.

23:32

>> Reinforcement learning is not a

23:33

particularly new

23:35

approach particularly particularly new

23:38

thing to do. So,

23:40

uh and some some way age of experience I

23:42

think I think always has been there and

23:45

people have been criticizing a bit

23:49

pre-training because pre-training work

23:50

clearly is this is the other way of of

23:53

looking at the models which is like we

23:54

have static data that data is mostly

23:56

generated by others. Although I I have

23:59

this personal view that pre-training

24:00

today is largely distilling other models

24:03

other models into the new model because

24:04

most of the tokens in the internet are

24:06

are coming from from AI. But there

24:09

there's clearly pre-training which is

24:10

which is behavioral cloning which is

24:12

mimicry which is which is compression of

24:15

of of of of internet data. But

24:17

reinforcement learning is not something

24:18

that that people haven't been thinking

24:20

and people haven't been haven't been

24:22

doing. Reinforcement learning was used

24:24

to solve backgammon back in the day.

24:25

They used to solve go, Starcraft, Dota

24:28

and to solving programming right now and

24:32

every time it comes down to model

24:35

writing its own experience and learning

24:37

learning from that experience. But this

24:39

is very clear and what I think is

24:41

interesting and what I think is still is

24:43

still perplexing to people that

24:45

reinforcement learning is not really the

24:47

only way to learn from experience and

24:50

there there will be

24:53

more and there will be there will be a

24:54

little bit more of

24:56

I think

24:58

you can call it algorithmic, but

24:59

essentially innovation of how we learn

25:01

from experience. Just because just

25:04

because reinforcement learning is only

25:06

one way to do it. It's a mathematical

25:08

formulation. And

25:11

especially right now how we are using

25:13

it. It really likes those parallel

25:14

rollouts for variance reduction and for

25:17

for comparing how the model does in in

25:19

in parallel versions of the world, which

25:20

is not how not how we do not how we

25:22

learn from experience. We learn from our

25:24

experience much more efficiently and

25:25

much more

25:27

Um we use we use those in many in many

25:30

many ways. At some moment I've been

25:31

trying to explain to people that what

25:33

what brain does how how we learn. So uh

25:35

there's one learning algorithm in a

25:37

brain. I think I think there are there

25:39

are multiple actually and they and they

25:41

they work together. But if I play

25:44

football for example, it looks very very

25:47

closely to reinforcement learning. I

25:48

kick ball a lot of times and every time

25:50

I adjusted a little bit and I see if it

25:52

roughly matches what I what I wanted and

25:54

I learn or or some self-reinforcement

25:56

had money. When I when I when I learn

25:58

mathematics, it's very different type of

26:01

thing. It's it's like reading about hard

26:04

concepts and thinking about them very

26:06

deeply inside my head until until things

26:09

click and I until until I have them

26:10

connected. And both of those in some way

26:13

are learning from experience. They are

26:15

just very different. So so summarizing

26:17

my thinking of the learning from

26:19

experience is that we've been doing it

26:20

for a while. We probably are spending

26:22

the most compute than ever on learning

26:25

from experience, but the reinforcement

26:26

learning is is not the end of learning

26:28

from experience and there will be

26:30

there will there will be better

26:31

approaches that will

26:33

researchers will be coming up in the

26:35

coming years on how to how to use that

26:36

data in in a richer new chart settings.

26:39

>> Thank you. Same.

26:41

>> I'm Rohan. I'm curious since a lot of

26:42

your work has been around optimization

26:45

and and efficiency, how do we get to a

26:48

orders of magnitude

26:50

more efficient, I guess more compute

26:52

efficient and more data efficient

26:54

learning algorithm.

26:55

>> You know, I would start with

26:56

measurement. I think pre-training as we

26:58

define it right now is about

27:00

compression. We look at perplexity and

27:02

then measure

27:03

how do we decrease the perplexity and

27:05

then we find that scaling and increasing

27:08

parameter count and putting more compute

27:10

is the way and every time we increase

27:13

compute in log scale, we get this

27:16

epsilon more improvement in these

27:18

metrics.

27:19

This is I think this is fine for

27:21

building the prior.

27:23

But I think this is the wrong way to

27:24

look at the problem. We should be

27:26

looking at the end-to-end. What are we

27:28

training these models for? Look at the

27:30

outcome. Like for example, I train this

27:32

model and give it to Jerry. Jerry will

27:34

do RL and destroy all the perplexity

27:37

metrics that I have created, right? Like

27:39

so then it's sort of like it was the

27:41

best way we had so far to attempt to

27:46

solve the problem and I think the labs

27:47

and everyone else have done a great job

27:49

in producing intelligence that's super

27:52

valuable and makes my work so much fun.

27:55

But it was the bootstrap process to get

27:57

there. We have to combine pre-training

27:59

and RL together and that's like

28:02

where one order of magnitude improvement

28:04

would come from and that's like a

28:06

training procedure. You can say it's a

28:08

learning algorithm. Um in terms of

28:11

optimization,

28:12

my story is I started optimization at

28:15

Google for logistic regression back in

28:18

2016. Got nerds like me and I worked on

28:21

some solvers for what we used to call

28:23

Sybil, which was the large-scale linear

28:25

solver that was used at Google before

28:27

neural network took off and then

28:30

replaced this. Then I asked myself like

28:33

what do I want to work on with neural

28:34

network and it was quite clear like I

28:36

want to understand the training

28:37

algorithm and make it better. And then

28:39

someone Vinit Gupta just showed up one

28:41

day at my desk. It's like I heard you're

28:43

really good at writing optimization

28:45

methods. We have this idea that you

28:47

know, uh like we worked out on a

28:49

whiteboard

28:50

uh and what turned out to be the Shampoo

28:52

algorithm. Can you help us implement

28:54

scale, make sure it works at large scale

28:56

for neural network training? So, as

28:59

working on this, I I thank my manager,

29:01

Yanghui Wu, who has like supported it

29:03

throughout then to like till my end of

29:05

my tenure 2024. But largely like the

29:08

community and most of the other people

29:09

were not as excited by this idea. And

29:12

for me, this is the most exciting thing

29:14

cuz I was like I'm putting in

29:15

computation and making training better.

29:17

This is the thing I have to figure out.

29:19

I will spend as much time I would would

29:22

take to do it. And then people were um

29:25

making this assumption, "Oh, why like

29:27

what's the upper bound? You could still

29:29

use Adam. That's fine. Like why are we

29:31

You could spend all the time on

29:33

everything else, not optimization." But

29:35

in some sense, optimization is like like

29:37

you have a model, you're optimizing it,

29:40

you want to optimize it better. Now,

29:41

I'll connect it to some of the stuff

29:42

that we talked about architecture. What

29:45

has happened is that a lot of the work

29:48

that we've done in architecture is to

29:50

make these networks train. And in some

29:53

sense, it's like two sides of the

29:55

coin and optimization and architecture

29:57

go together. You could have a stronger

29:59

optimizer train a much more harder

30:02

harder to optimize model and get better

30:04

performance. Or you can use a weaker

30:06

optimizer on easier to optimize models

30:09

and get decent performance. So, there

30:11

are these tradeoffs that appear all

30:13

over. And for me, it's a lot of time

30:15

working on it. I think we used it for

30:17

Gemini 1.5 Flash. And then like the

30:19

community started getting like more

30:21

interested in it. There was the Soap

30:22

paper published. We have like an entire

30:25

literature of like shampoo, soap, all

30:27

like bath time the things that you would

30:30

use. And then it was quite clear like so

30:32

that was like maybe a 2x improvement

30:34

over what was happening. But even then,

30:36

if you look at Shampoo, it's quite weak

30:38

in what it's doing. It's not using all

30:40

the information that's available to you

30:43

when you train. And as you use more and

30:45

more information as part of training,

30:47

you can get better improvement. And in

30:51

some sense, like your optimization

30:52

algorithm defines what architectures you

30:55

will discover. Like I have colleagues,

30:57

it's not very popular in the literature.

31:00

It's only like maybe sub like four

31:02

people in the world care about it. Kind

31:04

of ideas that are extremely interesting

31:06

like residual connections have been

31:08

extremely useful for training neural

31:10

networks. There are folks who have like

31:12

gotten rid of them and learned deeper

31:14

representations. But they needed a

31:16

better optimization method. So like for

31:18

me, optimization methods and the

31:20

question you asked is just like how do

31:22

we get there? It's It's combined with

31:24

architecture and thinking about the

31:26

problem end to end is where a lot of the

31:28

computational efficiency is in it.

31:29

>> Yeah.

31:30

>> And I also see RL as spending a lot of

31:32

compute not as efficiently. And so like

31:35

if you could spend it because you don't

31:37

get much feedback and you're

31:39

spending a lot more compute because you

31:41

have to decode all this long chain of

31:42

thought to get this one bit of

31:44

information into the network. Seems

31:46

quite inefficient and an easy target to

31:48

get orders of magnitude on top of. I can

31:51

go on talking about optimization all

31:53

day, but I

31:54

>> it.

31:55

>> Rohan really has

31:56

>> Do you think we'll ever approach or

31:57

surpass biological learning efficiency?

32:01

>> I do not think so because I think

32:04

we would need to change

32:05

I maybe like that was a strong

32:07

statement. At least with the hardware we

32:08

have, it seems pretty unlikely. Um it

32:12

our biological like we have something as

32:15

Jeff Hinton says, model computation. So

32:18

we build our own circuit as we grow up

32:20

and we build our own learning algorithm

32:21

with the hardware. And then we die and

32:23

then we're gone. Uh neural networks are

32:26

quite different. Our hard The hardware

32:27

stays, the neural network stays, but

32:29

it's learning very inefficiently and you

32:31

need a lot more of them and a lot of

32:33

parallelism to get small amounts of

32:35

information through. So until I think we

32:38

design hardware to be much more like how

32:41

humans operate, maybe more analog,

32:44

figure out how to deal with analog

32:45

circuits, and to figure out how to do

32:47

with error correction, figure out how to

32:49

get information through it will be much

32:51

harder.

32:52

Uh we

32:53

I think we're safe.

32:55

>> Safe. That's an interesting way to put

32:56

it.

32:57

Um the idea that pre-training and RL

33:00

should be optimized end-to-end

33:03

seems like such a, you know, clear

33:05

maybe obvious statement. Do do you think

33:08

Do you think the labs realize this? And

33:09

then is it just hard for them to, you

33:12

know, get rid of org charts and process

33:14

to be able to make that come together?

33:16

Or what stops the labs from from being

33:18

able to, you know, unify the the two?

33:20

>> Uh I don't think it's that obvious cuz

33:23

uh it's a completely, again, different

33:25

optimization problem. You have a prior,

33:28

you're doing rollouts, you have higher

33:30

variance, you're

33:32

and then

33:33

pre-training is much large batch, like

33:36

more parallelism or compute for the unit

33:38

of time that you can spend. So, it is

33:41

not an obvious thing for folks to

33:42

combine these two training procedures

33:44

until

33:45

uh you think a bit more like, "Why is it

33:48

that the naive combination doesn't

33:49

work?" So, that's one.

33:52

Uh the second one is if I pull like some

33:54

of the best researchers in these labs,

33:56

they would say, "Oh, this makes sense.

33:57

We should probably explore it, but it

33:59

would be probably not in the top bucket

34:02

because they have to train a model for

34:03

the next cycle before It's like as as

34:06

Jerry said, like there are companies now

34:09

competing for release cycles because

34:11

tokens are not sticky. So, it's much

34:14

harder to do long sight, like even

34:17

long-term research of 6 months in many

34:20

of these labs in the environment they

34:22

are in.

34:24

>> So, it seems like one of the core

34:25

premises for Coral Lemonade is, you

34:27

know, you're starting you're starting a

34:28

lab at a time when you know, Sam's been

34:30

talking about the AI scientists. I think

34:32

Dario's been talking about the AI

34:33

scientist. It seems like your job as

34:35

researchers has actually fundamentally

34:36

changed.

34:38

Um and you're get you get to start the

34:39

company native to that era. And as a

34:41

result, you're maybe able to run a lot

34:43

more experiments than otherwise might be

34:46

possible. How automatable do you think

34:49

the research job is?

34:51

Um and how are you guys approaching

34:54

building your lab to be as I believe

34:55

your mission one of your missions is to

34:57

be the most autonomous lab there is?

34:59

>> Most most automated lab in the world and

35:03

to to start with I think that automation

35:05

the version of automation

35:07

by core automation is about giving each

35:10

human maximum level of agency in some

35:13

way. It is we we're not trying to really

35:17

get humans out of the loop which is like

35:18

one version to automate, but it is about

35:21

give humans ability to do the most with

35:23

with their their amount of time.

35:25

Whenever whenever you are walking you

35:27

can get some distance. Whenever you get

35:28

a bike you can go go work a larger

35:30

distance. Whenever you are a car you can

35:32

you can go much much much much larger.

35:34

When whenever whenever

35:37

humans started farming they had to farm

35:38

by hand and will work on a small plot of

35:41

land. When you have a machine you work

35:42

on a much much larger plot of land.

35:45

Uh personally I am I'm both really great

35:48

fan of the of the current coding agents

35:50

and and and very happy it's in some way

35:52

it is what I've been working for many

35:54

years both doing coding research and

35:57

working on various versions of of AI

35:59

scientist inside inside of OpenAI and uh

36:03

in the end I realized starting a company

36:06

to realize that vision is is is is one

36:08

of the best one of the best ways to

36:10

realize it because the way you can do

36:12

research today is very very different

36:15

because a single researcher can do much

36:17

more. In the end the speed of iteration

36:20

the speed of research the speed of how

36:21

quickly you can you can move through

36:23

ideas and how quickly you can get data

36:25

on your ideas is something something

36:27

very very different and you can try to

36:31

move the old structures around it and

36:35

the teams workflows

36:38

how data is gathered or you can try to

36:40

build like like like you said you can

36:42

you can try to build build natively for

36:44

it for processes that that maximally

36:47

empower each researcher and allow them

36:49

to just to just iterate on their on

36:50

their idea much much quicker. We are we

36:53

are here and we are trying to to to

36:55

rebuild deep learning stack and try to

36:57

think how we can do almost each

37:01

operation differently. What are what are

37:02

various options and if we can execute at

37:06

least even one of those experiments a

37:08

day, that's already a pretty good

37:10

iteration speed versus anything that was

37:12

that was done before and there isn't

37:14

really like any any fundamental like

37:16

laws of physics reason why not and maybe

37:18

maybe one day we get to to 10 of those a

37:20

day. Maybe one day we get we get 200

37:22

those a day and fundamentally for that

37:25

like search process optimization

37:27

process, we should be able to just just

37:29

find things that work in a better deep

37:31

learning setting and what what we are

37:33

what we are trying to do is like we've

37:35

been like almost all of us we are we are

37:38

we are a team that is very agent built

37:40

and automation built. We are we are we

37:42

are we are we are trying to do an

37:43

experiment like how far how far we can

37:44

we can push those thing and how far how

37:46

how much an organization that tries to

37:49

do as much as we can with a small team,

37:51

how far how far we can get with that.

37:53

>> When will we know that we've reached

37:55

AGI?

37:56

>> At some moment I used to say it's it's

37:58

very much in everyone's heart, whatever

38:00

whatever they consider AGI. Open AI says

38:02

the system that can

38:04

outperform all humans in economically

38:06

valuable work, but it goes to my my my

38:08

my previous statement. What if what if

38:10

Open AI stops training models? Would

38:12

that would that would that would that

38:13

still keep working and I still would

38:15

would that still keep the automation

38:17

level the same the same or would that

38:19

would that drive?

38:20

And for me AGI is a model that can

38:24

improve itself without human in the loop

38:28

in any in any in any way. That's That's

38:30

I think I think the moment where we can

38:32

meaningfully talk about about AGI

38:34

because that is in some way it is a sub

38:36

definition of the previous one because

38:38

improving air models is actually a job

38:40

that humans can do. It is economically

38:41

valuable work. And it definitely

38:44

definitely is the case, but removing

38:46

humans from from from loops with models

38:49

has been actually notoriously

38:51

notoriously difficult so far. We haven't

38:53

come anywhere close to it. It's very

38:55

very hard for me to find any task where

38:56

all of them were able to like get get

38:58

humans out in the loop. We are We are

39:00

the the the the human LLM hybrid is is

39:03

really really successful right now, but

39:05

LLMs without humans not so much, not at

39:07

all. And what I have seen there in 2024

39:11

and in early 2025 is that the current

39:14

like path doesn't doesn't get us there.

39:17

And I think we need we need some pretty

39:19

pretty serious research on my this

39:21

company or some other to to try to

39:23

unlock how do we how do we make our

39:25

models learn and and adapt at a test on

39:30

a deeper level than we've been so far.

39:32

>> Mhm.

39:32

>> I feel like we've alluded to this

39:33

throughout the conversation, but it's

39:35

kind of been one of these, you know,

39:36

five blinded blindfolded people trying

39:38

to find the elephant things. Uh, what is

39:41

the grand master plan for for Core AI

39:43

that you're willing to share?

39:45

>> I can share our 6-month road map. In

39:47

some sense, like building architectures

39:49

as I said, like wouldn't it's not just

39:51

how good the architectures does it run

39:53

well and can we get users including

39:55

ourselves as part of the lab to use it?

39:58

Right? So, then

40:00

that's directly like we can do many

40:03

things now, but the thing that is going

40:04

to be difficult and thing that we want

40:06

to automate away is kernel generation.

40:09

So, we would we have a set of hardware

40:11

GPUs, black holes that we have to train

40:13

and run inference on. Uh, we will build

40:16

the best model that we can to

40:19

basically reduce the time from having a

40:22

very cool idea that can make

40:26

these bottlenecks go away in the

40:27

architecture to having them run at the

40:30

highest TFLOPS on GPUs. And in some

40:33

sense like current coding agents plus

40:35

humans can go a long way, but like an

40:38

example of this is our urgent QR kernel

40:41

competition that we hosted with GPU

40:44

mode. It's for running this fairly old

40:48

linear algebra operation QR. Um, it's

40:51

used for optimization like shampoo line

40:53

of work uses it. Many other places uses

40:55

it. And you want to run this efficiently

40:57

on a B200 node. Uh, if you use cool

41:01

solver for the shapes that we care

41:02

about, you get some efficiency. And then

41:05

a human plus some search loop can get

41:08

you like something like 7x.

41:10

But it requires the highest taste human

41:14

like there's maybe three people on the

41:15

world to and spend about $100,000 on

41:18

these coding agents over span of 4 weeks

41:21

to get to a solution that's 60x faster.

41:25

So these models today are no way close

41:27

to getting that 60x faster kernel. And

41:29

there's there is a real bottleneck now.

41:32

That was a single problem. It has like

41:35

perhaps three different operators. Work

41:38

on this panel, do this matrix multiply,

41:40

fold it back in, and do this repeatedly.

41:43

That's what the square factorization of

41:44

a matrix would look like. And if you

41:46

give this problem to entropic models of

41:48

an AI model Gemini, it just wouldn't

41:51

solve it. It just is not our models are

41:53

not even close to solving this problem.

41:55

So for us it's like something that we've

41:57

talked about, something that we're

41:58

getting close to as sort of like

42:01

getting to that point because that's our

42:03

inner loop to having more efficient

42:05

architectures.

42:06

>> Why kernels? Is it just cuz like

42:09

maximize intelligence per flop of

42:11

compute you need to generate your own

42:12

>> In some sense, um

42:14

I've had like three projects.

42:17

Two of them have

42:19

kind of landed in the industry. So,

42:21

first is

42:22

secondary methods. Kernels was a

42:24

bottleneck cuz you have to run it well.

42:26

If you were at a place like Google, you

42:28

cannot spend 10x uh amount of compute

42:30

and get a 2x win.

42:31

So, I could only spend maybe a budget of

42:33

20% and get the 2x win. Everyone's

42:35

happy.

42:36

I don't like great. So, like I think

42:38

that's the market, right? You spend uh

42:40

less than

42:41

>> Yeah.

42:41

>> you uh get. So, kernels ended up being a

42:45

bottleneck there cuz most of the

42:47

operations were novel that we haven't

42:50

gotten a lot of people to look at it.

42:52

There's only like two humans at Google

42:54

who could write it, Rasmus and Peter

42:56

Hawkins. Cuz it was deep XLA LLO code

42:59

that you have to write to make this work

43:01

and then that took them 2 years to do.

43:04

The other idea I had with one of my

43:06

coworkers at that time was like

43:08

replacing some of the parameters in a

43:10

transformer with extra memory and we

43:13

called it N-gram and N-gram memory.

43:15

We worked on it in 2020. It's a we had

43:18

versions of it internally deployed, not

43:21

the big version, the smaller version.

43:23

But, there I needed like something that

43:26

can accelerate sparse uh gathers and

43:28

scatters as part of training.

43:29

>> Yeah.

43:30

>> Uh it required hardware change and a

43:31

hardware making use of the hardware. It

43:34

never arrived. I had uh conferences set

43:37

up with the TPU team, us and a bunch of

43:39

others. We were talking about it and

43:41

during COVID like oh, we're going to

43:42

have it happen and it never arrived. I

43:45

also was using uh TPUs at Anthropic.

43:48

While I was leaving, just barely started

43:50

the surface of being able to um do it.

43:53

But, at the same time, 6 months before

43:55

that, Deep Seek wrote uh their N-gram,

43:57

which is a improved version of adding

43:59

more memory. Short scaling loss that

44:02

yeah, you don't need MOEs. You could

44:04

actually replace it with these

44:06

>> Mhm.

44:06

>> N-gram embeddings. for me that was like

44:08

ah

44:09

>> Yeah.

44:10

>> I It was like a five-year thing and I

44:13

was very happy for them.

44:15

>> isn't even possible if you're not

44:17

writing kernels.

44:18

>> Kernels and you need to be assisted in

44:20

writing kernels or solve that kernel

44:23

to have the highest performance. Like

44:25

and the

44:26

the roof line is pretty high. So it's

44:29

like the QR. If I use CuSolver's QR, I

44:32

get some performance. If I we use our

44:34

the competition winners QR, you get 60x

44:36

faster.

44:37

And that is a completely different

44:39

playing field. Now it opens up an entire

44:41

new set of algorithms you can apply and

44:43

in terms of training transformers,

44:45

training optimizers. QR is so

44:47

fundamental in analyzing the eigen

44:48

decompose for eigen decomposition and

44:51

many other things. So it is a thing that

44:53

I think also if you think about it, only

44:57

few people have the skill set to and

44:59

they're very much not at the same place.

45:02

It's like one person here, one person

45:03

there. And

45:05

it would be ideal if models had those

45:07

abilities.

45:09

>> Maybe I'll I'll

45:10

summarize a little bit and talk from the

45:12

high level of what we want. Kernel

45:13

automation is a lab created to

45:17

build models that continuously learn and

45:19

then learn from deployment. We believe

45:21

as

45:23

I mentioned that transformers are

45:25

incapable of continual learning. There

45:27

there's no way how to put continual

45:28

learning on transformers. So we know we

45:31

have to find a different different

45:32

architecture. Some of anyway, our quest

45:34

is to find that that new architecture,

45:36

find that transformer replacement and we

45:39

want to build the most automated lab to

45:42

do it. We want to be able to build

45:45

experiments at scale the quickest we

45:47

can, iterate on them, try a lot of new

45:49

architectural ideas, have strong priors

45:51

of what we want to do to search the

45:54

space of architectures efficiently to

45:56

find to go to that go to that place

45:58

fastest than than anyone else. That's

46:00

what we what we want to do and all the

46:02

work we are doing on kernels, on large

46:04

scale training, on trying new

46:06

architectural ideas is is exploring that

46:10

space.

46:10

>> So you can experiment your way

46:13

into finding a superior architecture.

46:16

How will you know when you've found it?

46:18

What are you looking for to say, "Aha,

46:21

this is the one."

46:23

>> That's that's a great question. There

46:24

are there always two angles. In in in my

46:27

mind and my experience, every successful

46:30

research had a plot that shows something

46:33

that's that other other plots don't

46:34

show. There there is a one line that is

46:36

a little bit bending in a

46:37

different different way and you're

46:39

saying this is this is what you want.

46:41

But at least it is my experience with

46:43

research always has been that plot is

46:45

already quite late in a journey where

46:48

most of the time you already know what

46:49

you want and already know what you are

46:51

up to. I am a bit a bit joking, but it's

46:54

actually true that all the best plots in

46:56

my life I have done in a dream before

46:58

before actually they were they were

47:00

real. I kind of knew I was looking for.

47:03

Just just the question is is like when

47:04

when it actually clicks if you if you

47:06

know what I mean because most of the

47:08

time you kind of know what you are

47:09

looking for, but you are you are not

47:11

finding it. You you try one thing and it

47:13

doesn't work. Try second thing and it

47:15

doesn't it doesn't work. But eventually

47:17

there all the all the right pieces fall

47:19

into it and most of that deep learning

47:21

systems are very intricate. So usually

47:22

you have to get five things right in a

47:24

row for the for the thing to start

47:26

working and then eventually you get the

47:28

plot that looks like like like like you

47:30

want and then and then you know. So So I

47:32

think

47:34

what we are what we are looking for is

47:35

systems that learn and test time and if

47:39

we see meaningful

47:41

long-term adaptability of our of our

47:44

systems and like we are we are we are

47:46

joking, but it's it's it's a it's a real

47:48

we want to be evaluating our systems of

47:50

our of our

47:51

everyday work. Will they get better at

47:54

doing the work of core automation

47:56

scientists each day?

47:57

>> Yeah, we like go on a vacation as a team

48:00

and see if the lab produces something

48:01

better.

48:02

For the week, give the

48:04

>> Then what you do when you get back?

48:05

>> Um, we'll see what

48:07

>> Extend Extend the vacation two times,

48:09

three or four times, until we are on

48:11

permanent vacation.

48:14

>> Uh, that is a beautiful note to end on.

48:17

Um, Rohan, Jerry, thank you so much for

48:20

joining us. You've both worked on like

48:22

really, really

48:24

transformative work for where we are

48:26

today. And I'm so excited to see you

48:29

starting a lab on this is new journey.

48:32

Uh, and very excited to see what you're

48:34

able to come up with. Thank you for

48:36

joining us.

48:36

>> It's an It's a great to be here and to

48:38

chat with you.

48:39

>> Thank you.

48:43

>> [music]

48:57

[music]

49:05

[music]

Interactive Summary

Jerry Rohan and Rohan Anil, founders of Core Animation, discuss the limitations of current Transformer models and the future of AI. They argue that while Transformers are economically valuable and have scaled pre-training and reinforcement learning (RL), they are not the ultimate solution due to their inability to learn continuously at test time on real-world data. They highlight that current approaches like limited in-context learning and fine-tuning with catastrophic forgetting are insufficient. Core Animation's mission is to develop new architectures that can meta-learn and adapt over longer horizons, as RL is just one form of learning from experience. They started their company to pursue this radical research, which larger labs avoid due to short-term market pressures. Their goal is to build the 'most automated lab' to rapidly iterate on architectural ideas, focusing initially on automating kernel generation for vastly improved computational efficiency. They define Artificial General Intelligence (AGI) as a model capable of self-improvement without human intervention.

Suggested questions

10 ready-made prompts