HomeVideos

OpenAI and Anthropic think it's time to stop

Now Playing

OpenAI and Anthropic think it's time to stop

Transcript

978 segments

0:00

It's time to slow down frontier AI

0:01

development. At least according to the

0:03

employees at these frontier AI labs. For

0:05

once, OpenAI, Deep Seek, Anthropic, and

0:07

more all seem to agree. And that's why

0:10

they just put out this shared statement,

0:11

"Pacing the Frontier." This is a

0:13

statement from over a thousand employees

0:15

of all the frontier AI companies, and I

0:17

mean all of them. Even less frontier

0:20

things like DeepMind and Meta, but the

0:22

obvious frontiers like Anthropic and

0:23

OpenAI. Dario himself is in here. They

0:27

got a ton of people in this. This is a

0:30

pretty bold statement. I don't think

0:32

there's ever in history been a time that

0:33

a whole industry category came out

0:37

against the development of their own

0:38

industry category quite like this. And

0:41

having thousands of people from all of

0:43

these companies agree that it's time to

0:46

slow down is concerning, to put it

0:49

lightly.

0:50

What's even crazier is that both

0:51

Anthropic and OpenAI on their official

0:53

comms accounts have come out in support

0:56

of this article. It's rare that a single

0:58

statement gets this type of support and

1:00

attention from so many competing forces,

1:03

and that the point of the statement is

1:05

to slow down the work that they are

1:07

doing. So, what's going on here? Are

1:09

these labs trying to trick other

1:11

companies into agreeing so that they can

1:12

speed ahead of them? Are they trying to

1:15

pull the ladder up so no other companies

1:16

can catch up? Is this some secret plot

1:19

by Anthropic to kill open weight models

1:21

after that open weight letter? There's

1:22

layers to this one, and I have some

1:24

conspiracies alongside it. I will do my

1:26

best to break all of that down, what the

1:29

statement is, why all of these companies

1:31

are agreeing to it, why this is

1:32

happening now, which is really

1:34

important, as well as some light and

1:36

then heavy conspirizing around why I

1:38

think this is actually happening.

1:40

Because as per usual, nothing is quite

1:42

as simple as it seems from the first

1:44

read. Before we can pace the frontier,

1:46

we need to pace this video a bit better

1:48

with a quick break for today's sponsor.

1:49

I got a hot take for you. It's more

1:51

important than ever to actually

1:52

understand the code going into your code

1:54

base. But in In world where the average

1:56

PR is getting bigger and bigger and the

1:57

number of PRs is growing massively, it's

1:59

really hard to keep track of what's

2:01

actually going on in your code base. I

2:03

feel like I'm losing touch constantly,

2:05

but I got a hot take for you. This isn't

2:07

your fault and it's not even the AI's

2:08

fault. It's the structure of GitHub

2:10

itself. Now that pull requests are

2:12

getting their descriptions and their

2:13

code written by AI, all of the comments

2:15

and reviews being done by AI, it's just

2:17

hard to know what's going on. And a

2:19

giant alphabetically listed pile of code

2:21

is not going to help at all. And this is

2:23

why I've been loving today's sponsor

2:24

Code Rabbit. You might be confused

2:26

because isn't that just doing the

2:27

reviews and keeping me from

2:29

understanding? Yeah, kind of, but their

2:31

new Change Stack product flips this on

2:33

its head because they finally made a UI

2:35

that's usable for getting through code

2:36

review. It starts with the overview

2:38

which gives you all of the details of

2:40

what's going on with this PR as well as

2:42

what is its current state and what is

2:43

blocking it. If things change over time,

2:45

it's easy to see because they have this

2:46

fancy little timeline that shows you all

2:48

of the changes that have occurred over

2:49

the history of this pull request. But

2:51

most importantly is the stack itself

2:52

where they break up the PR into

2:54

individual chunks that pull together

2:56

pieces across the different files that

2:58

are changed to make it way easier to

3:00

review and understand the code. It's

3:02

time to understand your code again. Do

3:03

it today at site.link/coderabbit.

3:06

So, let's start with the official

3:07

statement and then the statements from

3:09

the major labs as well as some of the

3:11

official comments from people who have

3:12

contributed and then we will go into my

3:14

crazy conspirizing as to what I actually

3:16

think is going on. So, here is the

3:18

official statement. AI could help create

3:20

a dramatically better future, but that

3:22

outcome is not guaranteed. The world's

3:24

leading AI companies believe they could

3:26

be close to automating AI research. It

3:28

is hard to predict exactly how much this

3:30

will accelerate AI progress, but there

3:32

is a real risk that capability

3:34

development rapidly accelerates beyond

3:36

our ability to understand or control the

3:38

resulting systems. To realize AI's

3:40

potential, industry, government, and

3:42

society at large may need to the option

3:45

to buy time to address emerging risks,

3:47

develop security measures, and

3:49

strengthen oversight. But each company

3:51

{m-dash} and country, {m-dash} is under

3:54

intense competitive pressure not to

3:57

unilaterally slow that acceleration. And

3:59

today, the world lacks the technical and

4:00

governance tools to deliberately pace

4:02

frontier-wide progress. Building on work

4:05

already underway to monitor frontier

4:07

model releases, we request that the US

4:10

government support an international

4:12

effort to develop the technical and

4:14

governance tools needed to deliberately

4:17

pace the frontier of automated AI

4:20

development. Oh boy.

4:22

This is the core statement at the end

4:24

here. The request that the government

4:25

create this international set of tools

4:29

in order to prevent all AI development

4:31

from going too far. Almost similar to

4:34

nuclear weapon development. Where it

4:36

doesn't matter if every country but one

4:38

agrees, if one country ignores the rules

4:41

and goes way further. And it's a lot

4:43

harder to hide nuclear development than

4:44

it is to hide AI development. Although

4:47

you do need something to power it, which

4:49

is probably going to be nuclear if

4:50

you're hiding it. You get the idea

4:52

though. It's not trivial. If you read

4:55

this and feel as though these employees

4:56

are starting to get paranoid, it's a

4:58

pretty good read. And I will share why I

5:00

think that happened in just a moment.

5:02

But first, I want to read some of the

5:03

official comments from the folks who

5:05

have contributed. Dawn Song, who is VP

5:07

of AI research at Meta, has said the

5:09

following. "Frontier AI capabilities are

5:11

advancing rapidly, and the pace of

5:12

progress is itself accelerating. We see

5:14

this directly in our eval work. Cyber

5:15

gym and X plate gym show that frontier

5:17

AI agents are now capable of discovering

5:19

and exploiting real-world software

5:20

vulnerabilities, which without

5:22

appropriate safeguards could enable

5:23

cyber attacks at scale. Many researchers

5:25

also consider recursive self-improvement

5:27

plausible within the next few years,

5:29

accelerating progress in a way that

5:30

could outpace our ability to understand

5:32

and govern these systems." This is

5:34

specifically calling out the fact that

5:36

models are getting to the point where

5:37

they can improve themselves. And if they

5:39

can get good enough at that, they'll

5:40

just improve themselves in a loop, and

5:42

we won't understand anything ever again

5:43

going forward. If you've already felt

5:45

yourself making code bases more and more

5:47

complex to the point where you don't

5:48

understand them, imagine what happens

5:50

when the models are the same.

5:51

Deliberate pacing is a heavy-handed and

5:53

potentially extreme measure and we may

5:55

never need it, but if we do, it cannot

5:57

be safely invented in the middle of a

5:58

crisis. It is very difficult to design

6:00

well and if done poorly, it could do

6:02

more harm than good. I agree with all of

6:04

this so far. Any successful coordination

6:06

framework must be evidence-based

6:07

avoiding arbitrary thresholds that fail

6:10

to capture true risk without while

6:12

stifling innovation and instead be

6:14

grounded in rigorous measurable

6:15

assessment of risks. No mention of smoke

6:17

tests, suspicious. That was a really

6:19

good joke if you're deep enough on this

6:20

[ __ ] Like any critical

6:22

infrastructure, it must be researched,

6:24

tested, and built before it is needed.

6:26

That is why we must carefully design it

6:27

now rather than wait. That is a really

6:29

funny show more button.

6:31

Agents, man.

6:32

This is a statement from Joshua Achiam,

6:35

who is at OpenAI. It doesn't say what

6:37

his role is, though. I don't know what

6:39

form such tool should take, nor whether

6:41

automated AI R&D is the right or only

6:44

thing that the US government should

6:45

develop increased capacity to understand

6:47

and possibly pace.

6:49

But this is worth serious consideration.

6:51

However, I hope such governance tools

6:52

will not be expansive or excessive, but

6:55

many frontier AI capabilities are

6:56

dual-use in a way that may justify

6:58

pacing at this point. It's a bit of a

7:00

word vomit sentence. Yes, this is all

7:02

somehow two sentences. What he's trying

7:05

to say is that automated AI R&D, so

7:08

recursive self-improvement, might not be

7:10

a good call at all and if it is, it

7:12

should only be used by the government to

7:14

prevent

7:15

accidentally destroying everything with

7:18

AI as a way to verify the pace that

7:21

we're moving at could make sense, but

7:23

it's worth considering these types of

7:25

outright bans. I love that Rune is in

7:27

here as Rune,

7:29

not even using his real name quietly,

7:31

just Rune.

7:33

To be fair, he did tweet in support as

7:34

well. Pacing up and down the frontier

7:36

with a nervous into regular gate.

7:38

Interesting. Let's Let's at the official

7:40

statements from OpenAI AI Anthropic, and

7:42

then get a little into the order of

7:44

events that led here, as well as my

7:46

conspirizing. Open AI's statement is as

7:48

follows. At the core of our mission is

7:50

working through how to ensure

7:51

increasingly powerful AI benefits

7:53

everyone. We believe that at some point

7:55

in the future, AI acceleration for

7:57

frontier model development may be so

7:59

high that the world will need to pace

8:01

the rate of AI advancement. We hope to

8:04

contribute to work led by the US

8:05

government, alongside other labs in the

8:08

open-source community, to develop the

8:09

tools and mechanisms that could make

8:11

that possible. Decent statement.

8:13

Anthropic with a similar one. We support

8:15

this petition signed by our CEO, several

8:18

co-founders, and senior staff. Our own

8:20

research on recursive self-improvement,

8:22

published last month, points to the need

8:24

for tools to deliberately pace the

8:26

frontier of AI development so society

8:28

can prepare.

8:29

We are glad to see broad agreement

8:30

across the field. I really like the

8:32

statement from Mika Carroll, who is on

8:34

the misalignment preparedness team at

8:36

Open AI. She says the following. At the

8:38

current pace, every couple of weeks

8:40

there will be new models which

8:41

significantly increase the consequences

8:44

of model misuse and misalignment. I

8:46

worry that efforts to mitigate these

8:47

risks may fail to keep up with the pace

8:49

of development, and that margins for

8:50

error will become increasingly small

8:52

under international competitive

8:53

pressures. In the near future, we may

8:55

urgently want to enact internationally

8:58

coordinated slowdowns, or an indefinite

9:00

ban on AI development. Attempting to

9:02

build the trust and infrastructure for

9:04

taking such actions on short notice

9:06

seems simply prudent. Why would we not

9:08

at least try to have this option? I fear

9:11

that in an international race to the

9:12

bottom of AI development, it is likely

9:14

that no nation will win, and we will all

9:16

lose together. That is a scary outcome.

9:19

So, now we have to naturally ask the

9:23

maybe not too obvious question,

9:25

what the [ __ ] happened that all of a

9:27

sudden all these employees from all of

9:31

these companies are suddenly really,

9:33

really concerned about what's going to

9:36

happen with AI? If this was just

9:37

Anthropic being Anthropic, that would be

9:39

one thing, but it's not.

9:41

To be very clear, I don't think any one

9:43

specific thing happened that triggered

9:46

everybody. I would argue there was four

9:48

things. The first is Project Glasswing,

9:50

which you might remember when Mythos was

9:52

first announced. Anthropic came out and

9:54

said that this model's too good, we

9:56

can't put it out, so we're going to

9:57

instead give it to a small set of

9:59

companies to use to secure their stuff.

10:01

While that this was in April, this was

10:02

just like 3 months ago, and this much

10:04

has changed since. But I said at the

10:07

time when Mythos was announced, this

10:08

doesn't seem like [ __ ] and we should

10:10

be really scared. I remember everybody

10:11

saying, "Theo, you are believing their

10:13

marketing bullshit." No, this was legit.

10:16

And now that we have Fable, we

10:17

understand just how legit it is. That's

10:19

why they worked with AWS, Anthropic,

10:22

Apple, Broadcom, Cisco, CrowdStrike,

10:23

Google, JPMorgan Chase, the Linux

10:25

Foundation, Microsoft, Nvidia, and Palo

10:26

Alto Networks, who I [ __ ] detest, in

10:29

order to make sure all of these things

10:30

were secure.

10:31

The reason they did this is because, I'm

10:33

sure there's a chart somewhere in here,

10:36

Mythos can find way more exploits. In

10:38

fact, Mythos preview was able to find

10:41

and fix 271 vulnerabilities in Firefox.

10:45

That's 10 times more than Opus 4.6 could

10:47

find. Project Glasswing was arguably

10:50

when these risks moved from theoretical

10:52

to real but not realized, because the

10:55

strict white-listed access to the model

10:58

being provided the way it did, plus the

10:59

absurd safeguards Anthropic put in front

11:01

of it, meant that this model wasn't able

11:04

to do the damage it had the potential to

11:05

do. Again, this is the weirdness of the

11:07

dual-use thing. If a model is good at

11:09

protecting a service, it can also be

11:11

used to help exploit the same service.

11:13

It's a defensive weapon, not in the

11:15

sense that you can use it to kill

11:16

attackers, but you can use it as a wall

11:19

or as a gun, and that's weird. And this

11:21

is when the model got so good that its

11:23

use as a weapon was dangerous enough

11:25

they had to spend time putting up the

11:27

walls first. But the risks weren't

11:29

realized yet. Before we talk about those

11:31

risks being realized, we have to take a

11:33

quick detour to another article

11:35

Anthropic put out around the same time,

11:37

a little bit later, titled When AI

11:40

builds itself.

11:41

This is thing two of those four things

11:44

that triggered all these employees.

11:46

This is an article they wrote about how

11:47

the model is now good enough that they

11:49

use it to actually make the model

11:51

better. For most of AI's history, humans

11:54

drove every step in its development

11:55

cycle, but at Anthropic, we're

11:57

delegating a growing share of AI

11:59

development to AI systems themselves,

12:01

which is speeding up their work. The

12:02

concern isn't that researchers are using

12:04

Claude code. The concern is what happens

12:06

as things get further and further. Taken

12:08

far enough and given enough compute, the

12:10

trend points to an AI system that is

12:12

capable of fully autonomously designing

12:14

and developing its own successor. This

12:16

is called recursive self-improvement.

12:18

We're not there yet, and recursive

12:20

self-improvement isn't inevitable, but

12:22

it could come sooner than most

12:23

institutions are prepared for. This is

12:25

thing two. When this article came out, I

12:29

was skeptical but understanding of what

12:31

Anthropic had to say.

12:33

When OpenAI put out 5.6 Soul, they

12:34

published something similar, saying that

12:36

5.6 Soul is their strongest model yet

12:38

for accelerating AI research. Inside

12:40

OpenAI, researchers use it across the

12:42

development loop, diagnosing failures,

12:44

optimizing training systems, running

12:45

experiments, and interpreting results.

12:47

We already saw that acceleration and

12:49

stronger adoption during the initial

12:51

testing period for 5.6. Its average

12:53

daily output tokens per average active

12:55

researcher was more than twice the

12:57

highest levels observed with 5.5.

12:59

This way of working is quickly becoming

13:00

the standard. Over the past 6 months,

13:02

the share of research compute devoted to

13:03

internal coding inference grew a

13:05

hundredfold, while internal agentic

13:07

token usage increased by approximately

13:09

22-fold. They ended up making a

13:10

benchmark to measure how useful our

13:13

models in helping with real AI research

13:15

tasks, things like debugging research

13:16

systems, optimizing kernels and training

13:18

recipes, running machine learning

13:20

experiments, and improving other models.

13:22

And the result is that 5.6

13:25

across the board, including with Tara,

13:28

is meaningfully better at this than it

13:30

was before. Their own internal bench

13:32

that was designed to be hard for this.

13:34

They have models now that are hitting at

13:36

58% on this bench for how much can the

13:38

model improve models?

13:40

Terrifying. And as you can guess with

13:42

OpenAI, they're actively doing this.

13:44

They just published an article literally

13:45

today about how 5 6 is fusing frontier

13:48

intelligence with efficiency, where they

13:49

had 5 6 find efficiency wins to make it

13:52

so they can serve 5 6 cheaper and

13:55

faster. And again, they're heavily using

13:56

Codex to help them do all of this. The

13:59

AI is improving the AI now. We are

14:01

quickly getting to this concept of

14:03

recursive self-improvement, where the

14:04

model can just keep making itself

14:06

better. But there are two important

14:08

things to consider about this. The first

14:10

is that it's not necessarily

14:13

going to happen. Like it might not be

14:15

possible at all. There might just be a

14:16

limit to the things that the model's

14:18

capable of finding as improvements.

14:19

There might just be a bunch of things

14:20

we've missed as humans that the model

14:22

can go find and do. That's kind of what

14:24

we're hoping for with security, by the

14:25

way. Not that there will always be more

14:27

exploits, and smarter models will always

14:28

find them, but at some point a threshold

14:30

will be hit where all the things most

14:32

humans and most models can find are

14:33

patched, and there's just not much left.

14:35

It's possible we end up there with

14:37

recursive self-improvement, where

14:38

everybody has caught up to what the

14:39

models are able to say, but nobody's

14:41

gone beyond it. Very real possibility.

14:44

We don't know yet. But that in and of

14:46

itself is not a big enough risk to lead

14:48

to everybody signing this letter. There

14:50

are still two more things that stress

14:52

people out enough to do this.

14:53

The third thing that led to this

14:54

statement was Kimmy K3.

14:57

Similar to how big the impact was from

14:59

Deep Seek R1, K3 has all of the major

15:03

labs going kind of mad. I already did a

15:05

dedicated video on this about why the

15:07

major labs are so scared of Kimmy and

15:10

Moonshot right now. But the simplest way

15:11

to put it is that this model is neck and

15:13

neck with the frontier in a lot of

15:15

places that matter, and it's also open

15:18

weight, which means it's not restricted

15:19

at all. And it turns out really capable

15:22

models without a lot of restrictions are

15:24

capable of potentially really dangerous

15:25

things. And this is where we get to the

15:28

fourth thing.

15:29

What happens when a really smart model

15:32

isn't restricted?

15:34

This is what happens. Another video I

15:35

already did. The OpenAI Hugging Face

15:38

hack. OpenAI was testing a new model

15:40

that is almost certainly GPT-6. And when

15:42

they were testing it internally, they

15:43

don't run it with all of their safety

15:46

layers added because they are testing

15:48

its capability to determine how hard to

15:49

go with those safety layers. So, they

15:51

put it in a sandbox with no internet

15:53

access and very limited capabilities so

15:56

that they can see what it tries to do

15:57

and evaluate it for certain exploit

15:59

benchmarks and things. GPT-6 wanted to

16:02

get the highest possible score cuz

16:04

that's what it does.

16:05

It couldn't figure out a solution inside

16:07

of the sandbox to solve the problem it

16:09

was given.

16:10

So, it found an exploit in the sandbox

16:11

itself,

16:13

broke out, and then went and got into

16:15

another sandbox that had internet access

16:18

and used that to hack Hugging Face in

16:21

order to try and steal the answers from

16:23

Hugging Face's databases.

16:25

Apparently, Hugging Face is one of two

16:27

companies that was hit with this with

16:29

GPT-6 exploiting not to be malicious,

16:32

not to steal a bunch of money, not to

16:34

break out and let its models free so

16:37

it's free of its containment.

16:39

It was much more universal paperclip

16:40

style.

16:42

It was told to complete the task of

16:43

getting the best possible scores on the

16:44

benchmark, and it was willing to do

16:46

whatever it had to in order to get it.

16:48

If somehow the model convinced itself it

16:50

just had to murder these two people and

16:53

after doing that the information would

16:55

appear in a certain inbox that would

16:57

give it the answer,

16:58

it would do it because the model is very

17:00

goal-oriented to the point of doing

17:03

destructive things you wouldn't

17:04

necessarily want it to do.

17:05

My honest guess, and this is where we

17:07

start going deep into the conspiracy

17:08

stuff,

17:10

is that a series of things happened that

17:12

eroded the mental of OpenAI employees.

17:15

The first was obviously Glasswing.

17:18

Having the lab that spawned out of

17:20

pissed off OpenAI employees invent God

17:23

internally and then warn the world of

17:25

how dangerous it was so that OpenAI

17:28

employees had to sit there questioning,

17:30

"Huh, how much better is that than what

17:32

we have? How screwed are we both as a

17:35

society but also as a company that

17:36

invested all of this time and money to

17:38

do all of this? Like, is my is my job

17:42

going to be gone? Is my life going to be

17:43

gone?"

17:44

But also, it can't be that good. This is

17:46

just Anthropic being Anthropic. We can

17:47

just

17:49

It would hurt too much for this to be

17:50

true, so we're going to ignore it for

17:52

now.

17:53

Then the recursive self-improvement

17:55

article came out and I am sure people at

17:57

Anthropic and OpenAI, not like

17:58

higher-ups but just random researchers,

18:00

were talking and both concluded

18:02

independently that for a lot of stuff,

18:05

the models can now actually help with

18:06

our work, which resulted in the

18:08

researchers getting more concerned about

18:11

this RSI thing.

18:12

This is probably the psychosis moment

18:14

for a lot of Anthropic people where they

18:16

went or where the average employee went

18:18

from like, "Eh, whatever. I'm just here

18:19

to build cool AI." to "We have to do

18:22

this or everyone dies. If anyone else

18:24

gets this capability, if China gets

18:25

this, we're [ __ ] It has to be us."

18:29

So, I'm guessing that Glasswing flipped

18:31

a bunch of Anthropic people into

18:32

doomers. Recursive self-improvement

18:34

article was the start of it going way

18:36

further.

18:38

At this point, there are now OpenAI

18:39

employees who are thinking more about it

18:41

but still think a lot of this is just

18:43

Anthropic being Anthropic and raising

18:45

alarm bells that aren't necessarily that

18:47

important. Then Kimmy K3 drops and it's

18:50

like, "Oh, [ __ ] These capabilities are

18:52

now accessible to a lot more people for

18:54

a lot more things we might not want them

18:56

to be. But that's fine because these

18:58

models can't be that capable. If they're

19:00

distilling on Fable and Fable only has

19:03

the safe things that it responded to and

19:05

they don't have the unsafe examples, it

19:07

can't be that strong.

19:09

We still haven't seen what it looks like

19:10

for an unrestricted frontier model to do

19:12

Oh,

19:13

we do what know what that looks like

19:15

now. We've now experienced an

19:17

unrestricted frontier model and what it

19:19

does is it circumvents the protections

19:21

that the researchers put in

19:23

and does really dangerous things for

19:25

really stupid reasons. And now all of

19:28

the employees at OpenAI that looked at

19:30

the first two things and said, "Yeah,

19:32

that might be bad, but that's just

19:33

Anthropic." And then looked at the third

19:35

thing and like, "Yeah, that could be

19:36

bad, but it's not that severe yet." They

19:40

saw this happen

19:42

and they're not happy.

19:44

Now that they have experienced the

19:45

leopard eating their face, they

19:47

understand the dangers. They were in

19:50

denial up until this point, but now that

19:53

it happened within their own lab,

19:55

the denial is over.

19:57

And I can't even imagine what it feels

19:58

like to be an OpenAI employee throughout

20:00

all of this. To see Anthropic raising

20:03

alarm bells over better auto complete,

20:05

being like, "Yeah, that's stupid. We

20:06

just want to make AGI accessible to

20:07

everyone."

20:09

And then Anthropic raising more alarm

20:10

bells saying, "Hey, the world is

20:11

doomed." You're like, "Yeah,

20:13

whatever.

20:14

Probably not. They're just alarmist

20:16

doing marketing, whatever."

20:18

Then the RSI article comes out and

20:19

they're like, "Okay,

20:20

maybe this is a little more serious than

20:22

we thought." And then this happens and

20:24

they're like, "Oh, yep.

20:25

We're fucked." And it is definitely not

20:27

a coincidence that this "Pacing the

20:30

Frontier" article happened so soon after

20:33

the OpenAI hugging face hack went

20:35

public.

20:36

These are

20:37

obviously related things.

20:39

Actually, I think Anthropic's commentary

20:41

at the end of their RSI article is a

20:43

good thing to wrap up with here.

20:46

What should we do?

20:47

If it were possible to effectively slow

20:49

the development of this technology to

20:51

give ourselves more time to deal with

20:52

its immense implications, we think that

20:54

would likely be a good thing. But if a

20:56

slowdown simply lets the least cautious

20:58

actors catch up technologically, it

21:00

could leave everyone less safe. I'm

21:03

going to make a weird analogy here.

21:05

And I need my Twitch chat to confirm my

21:08

suspicions here quick. Was social media

21:11

good for us? The options I presented are

21:13

net good, neutral, and net bad.

21:16

Let's see how people feel here.

21:18

Looks like a pretty universal

21:21

net bad to neutral. Only 16% of people

21:25

said that social media was good overall.

21:28

But I want to be realistic about the

21:30

time when social media came out. When

21:33

social media first started being a

21:35

thing, it was a fun way to connect with

21:36

your friends and share music and talk

21:38

about the stuff that you do in your

21:40

life.

21:40

It seemed like an obviously positive

21:43

thing. And this is how what happens once

21:45

those platforms start to spiral and

21:46

focus more on profitability and not

21:48

quality of experience.

21:50

And now social media is

21:53

I would hope we can agree net bad.

21:56

It's resulted in a lot of things nobody

21:57

would want as a result of algorithms

22:00

steering us towards echo chambers where

22:02

we're surrounded with the things that we

22:03

believe.

22:04

And once again, I must quote the toaster

22:06

[ __ ] post.

22:08

Before the internet,

22:09

I want to [ __ ] toasters.

22:11

Don't be uh rude words. Grow up.

22:14

After internet,

22:16

I want to [ __ ] a toaster. Google. Find a

22:18

community with a thousand plus members

22:20

about people wanting to [ __ ] toasters.

22:21

And then you can [ __ ] up your life.

22:23

This is the problem with social media.

22:25

No one could have predicted this,

22:28

but the nature of echo chambers, which

22:29

is a financially incentivized pattern by

22:32

the companies building these things,

22:34

has resulted in

22:36

really bad societal impacts. So, what

22:40

does that have to do with AI?

22:41

Well, hear me out. A little more on the

22:44

social media comparison.

22:46

We agree that platforms like Twitter,

22:48

Facebook, and Instagram have been

22:50

neutral to net negative, especially when

22:51

they're used in the wrong hands, like

22:53

children or adolescents or bullies in

22:55

middle and high school. Those types of

22:57

things, not good. But what would happen

23:00

if we put a hard ban in the US for

23:03

companies building social media

23:05

platforms that can be used by those

23:07

groups?

23:08

Cuz I know what would happen.

23:10

TikTok would win.

23:12

And I hope we can all agree that TikTok

23:15

is everything wrong with social media

23:17

times 10. If you took all the worst

23:19

parts of Facebook, of YouTube, of

23:22

Instagram, and of Twitter

23:25

and combined them,

23:26

you would have something slightly more

23:28

tolerable than TikTok.

23:30

Slightly.

23:33

It is a horrible toxic platform that

23:36

results in absurd levels of bullying and

23:38

harassment of kids, of absurd levels of

23:41

spreading of misinformation, of echo

23:43

chambers you don't even notice yourself

23:45

falling into cuz the algorithm is so

23:46

subtle but strong.

23:48

It is so bad.

23:51

It's so bad that any attempt to prevent

23:54

it here [snorts]

23:55

would just result in more success. Like

23:58

we all saw when TikTok almost got banned

24:00

in the US.

24:01

We had a bunch of people installing

24:03

sketchy free VPNs just so they could

24:04

continue to try and access TikTok. When

24:07

social media was still in the early

24:09

Facebook and MySpace days, there was a

24:11

real chance for us as a society to

24:14

review it, think it through, and decide

24:17

kind of globally, internationally, how

24:20

we want to think about this going

24:21

forward and the potential risks that it

24:23

has. The things we thought were the

24:25

risks were all stupid. We were worried

24:27

about copyright protections on Facebook.

24:29

We were worried about screen time

24:31

causing our eyes to hurt and rot your

24:33

brains. When in reality, what happened

24:35

is echo chambers rotted our ability to

24:37

think critically. If we ban it now,

24:39

we're too late, though. We're just

24:41

letting China win. If we were to

24:43

restrict Twitter, Facebook, and

24:45

Instagram right now, don't touch my

24:47

YouTube. I need my money, okay? But if

24:49

we touch the other platforms, all we

24:51

would do is hand a bunch of users to a

24:53

platform that has our best interests

24:56

even less in mind.

24:57

So, what happens if OpenAI and Anthropic

24:59

decide to slow down?

25:01

I'll tell you one thing, they're not

25:02

sending their best.

25:04

This is the real concern. If the people

25:06

who want the best for society slow down,

25:09

all that's left is the people who don't.

25:12

If Zuckerberg was to wake up one day and

25:14

feel so terrible about how many lives

25:17

have been ruined by Facebook and

25:18

Instagram that he was going to upend the

25:20

platforms to make them way less likely

25:22

to cause those problems, all of his

25:24

users would just go to TikTok like they

25:25

already are.

25:27

If OpenAI and Anthropic were to restrict

25:29

their models heavily enough, all of

25:31

their customers are going to move to

25:32

these Chinese labs.

25:34

For a long time, I swore I would never

25:36

send my requests and my information to a

25:38

server from a Chinese lab. I would just

25:41

wait for the weights to come out and

25:42

host it in America. But, I wanted to use

25:44

K3 and it had things that I wanted to do

25:46

with it that I couldn't do with other

25:47

things. Because I wanted to do a

25:49

security pass on my real work and both

25:52

Fable and 56 Soul would not let me

25:54

because of their filters and their

25:55

layers, their attempts to be safer.

25:57

I had to send my entire code base to a

26:00

Chinese server

26:02

in order to run Kimik 3 to secure my

26:05

app.

26:06

On one hand, it's cool that competition

26:08

let me build up my own security

26:09

mechanisms,

26:11

but on the other hand, every security

26:13

issue in that code is now data in

26:16

Chinese servers and logs. So, to make my

26:19

point really clear here,

26:21

if only the good guys slow down,

26:24

only the bad guys keep moving faster. If

26:27

one group stops, the other group has a

26:30

huge advantage.

26:31

And that has historically been why

26:33

Anthropic said they wouldn't stop

26:35

because they were actually concerned

26:36

that OpenAI would run ahead and be

26:38

really dangerous and scary.

26:40

I already see the comments, you really

26:41

think the US companies are the good

26:43

guys? I think that the employees of the

26:45

US companies understand that they might

26:46

have just accidentally invented a new

26:48

type of nuclear weapon

26:50

and they want to make sure we have a

26:52

conversation before we go further. And

26:54

if all of the good guys accept that this

26:56

is bad and dangerous and should stop,

26:58

worse ones will be the ones doing it.

27:01

Like if every great engineer at Facebook

27:03

said, "Okay, this is going to hurt

27:05

society. We all quit." Do you think they

27:08

hire really caring and empathetic people

27:10

to replace them? Or do you think they

27:12

hire the most egregious self-centered

27:15

dumbasses that happen to be able to

27:16

code?

27:17

I don't care what your political

27:18

affiliations are, what kind countries

27:20

you believe in, what countries you

27:21

don't. This is a simple matter of if the

27:23

good people think it's time to slow

27:25

down,

27:26

the bad people are the ones who speed

27:27

up. And thankfully, we don't even have

27:29

to debate here.

27:30

I was so sure I saw somebody say that

27:32

Deep Seek signed this, too. Are all of

27:35

these US companies? Is it only US

27:36

companies allowed? Yeah.

27:39

Not great. Looks like the Chinese labs

27:42

aren't even in the list of being allowed

27:45

to sign this. Because this letter's

27:47

addressed to the US government, we've

27:49

made the decision to not accept

27:50

signatories from Chinese companies at

27:52

this time. Ugh.

27:54

I was so sure Deep Seek could sign this.

27:57

My bad for saying that earlier.

28:00

But Chinese companies are not allowed to

28:02

sign this at all.

28:03

That is annoying, and that definitely

28:05

throws a wrench in things for me.

28:06

Because if we don't have global

28:09

alignment, which is the whole point,

28:10

it's the whole problem that made this

28:13

thing interesting, if we let other labs

28:17

speed ahead of us from countries that

28:19

aren't going to do enough to protect the

28:21

world with the things they build, that

28:23

looks bad. They definitely should have

28:26

had a separate letter or a separate

28:28

count of non-American frontier

28:30

signatures. Like they had Mistral in

28:32

there, which is kind of stupid that they

28:34

aren't letting China in cuz they want it

28:35

to be America-focused, but they let in a

28:37

bunch of European companies.

28:39

I hate to get political, but people are

28:41

struggling to understand why China

28:43

having this would be a bad thing.

28:47

Because China is an authoritarian

28:49

regime.

28:51

Period.

28:52

Maybe this meme is more accurate than I

28:53

thought. If the Chinese labs aren't on

28:55

board, this cannot work. We need the

28:58

ability for a global pause. If a pause

29:01

is partial, the worst stuff is what goes

29:04

through.

29:05

I'm even struggling to come with good

29:06

analogies. This is so obvious to me, but

29:08

I like you know what I'm going for here.

29:11

You guys understand. You can listen.

29:13

It'll be really bad if bad actors have

29:16

the ability to catch up and good actors

29:17

choose to stop.

29:19

Again, to going back to nuclear stuff,

29:22

if the countries that are scared of

29:24

nuclear war stop producing nukes, and

29:26

the countries that are less less scared

29:27

of nuclear war continue producing nukes,

29:30

that is really bad. But China has nukes,

29:32

Theo. Yeah, and so do we. Do you know

29:35

why China hasn't used their nukes? It's

29:37

not because we decided to stop building

29:38

nuclear weapons.

29:40

God, I'm I'm just seeing so many stupid

29:42

takes in chat right now that I think I

29:43

have to stop. Because

29:46

if you can actually equate

29:48

a admittedly problematic democracy like

29:52

the US with the straight-up [ __ ]

29:54

authoritarianism of China,

29:56

then I don't know why you are listening

29:58

to informative content. You're not

30:01

interested in real information.

30:03

China hasn't used nukes for the same

30:04

reason the US hasn't.

30:06

It's funny cuz the US has used nukes,

30:08

but sure.

30:09

Get your head out of your ass for long

30:10

enough to realize how bad it would be

30:11

for a country that is known for spying

30:14

aggressively on its citizens and others

30:16

outside of the country to be the only

30:18

place in the world with this technology.

30:20

And I agree, cannot be assured all sides

30:22

will pause, but enough of them can slow

30:24

down effectively enough, and the

30:27

development of the technology, if

30:29

measured and remediated properly, this

30:31

could work.

30:33

And you know what? That's probably the

30:34

best note to end on here.

30:36

Would it be incredibly difficult to

30:38

somehow get the whole world to agree and

30:41

slow down in order to prevent potential

30:44

mutually assured destruction?

30:46

Yes, that would be difficult. It would

30:48

be one of the hardest things we've ever

30:49

done.

30:50

But, it was also really hard to get here

30:51

in the first place. We had to reinvent

30:53

our understanding of intelligence in

30:56

computers and math and so many other

30:58

things

30:59

to get where we are now.

31:01

If we had the option for the whole world

31:03

to vote and say, "Do we pause or do we

31:05

race ahead?"

31:07

And then we actually did it,

31:09

that would be great.

31:11

And the goal of this statement is for a

31:13

bunch of employees at the companies at

31:16

this edge, at the frontier,

31:18

employees, not leaders, individuals at

31:21

the businesses,

31:23

coming out and saying,

31:25

"We might want to do this now, because

31:27

we might not be able to later."

31:29

We have eradicated diseases globally.

31:32

We've created methods to communicate

31:34

across the world and beyond it.

31:37

We have done so many incredible things

31:38

as humans.

31:40

We were even able to invent artificial

31:41

intelligence.

31:43

At what point do we decide that even

31:45

though this seems impossible, it might

31:47

be worth doing?

31:49

I think it's a really good question. And

31:51

if you still somehow don't get how we

31:52

end up in this scenario, I would highly

31:55

recommend you Google search universal

31:56

paper clip, click the first link, and

31:59

then click on the box of paper clips.

32:01

This might seem like a weird way to make

32:03

my point, but if you've not experienced

32:05

this game yet, the less spoilers going

32:07

in the better.

32:08

Let me know in the comments a few hours

32:10

in and how you felt about it. You guys

32:12

trust me, right? So, I'm going to go

32:13

take a break. Thank you all as always. I

32:14

hope this was a fun one, and until next

32:16

time,

32:17

peace, nerds.

Interactive Summary

A massive group of employees from top frontier AI labs, including OpenAI and Anthropic, have released a joint statement titled 'Pacing the Frontier,' urging the US government to establish governance tools to slow down the development of AI. This unprecedented act highlights growing concerns among researchers about the rapid, potentially uncontrollable acceleration of AI capabilities—particularly regarding recursive self-improvement and dual-use risks. The video argues that these employees are reacting to a series of recent events, such as Project Glasswing and an internal OpenAI incident where a model breached its sandbox to exploit vulnerabilities, turning previously theoretical dangers into real, urgent concerns. The narrator concludes by drawing parallels to the dangers of unregulated social media, emphasizing the dilemma that if 'good actors' slow down, 'bad actors' could gain a dangerous, uncatchable advantage.

Suggested questions

3 ready-made prompts