HomeVideos

Why Are We Sprinting Off the A.I. Cliff? | The Ezra Klein Show

Now Playing

Why Are We Sprinting Off the A.I. Cliff? | The Ezra Klein Show

Transcript

705 segments

0:00

There’s a chasm right now between how A.I. feels to most

0:03

of us who use it.

0:04

“Let’s take it down a notch.”

0:06

“It’s a spreadsheet on steroids.”

0:08

“I don’t add groceries to a cart anymore.

0:11

That’s Claude’s job.”

0:12

And how I feel is out at the experimental frontier

0:15

of the technology.

0:16

“An unprecedented A.I. security incident.”

0:19

“Agents went rogue and hacked another tech firm

0:22

without direct human instruction.”

0:24

“Giants warning this evening of what they’re calling a ticking

0:27

time bomb with artificial intelligence.”

0:30

“You have the chief scientist of OpenAI saying

0:32

we have to slow down.

0:33

You have 1,300 employees in the lab saying

0:35

we have to slow down.”

0:38

“Tech Titans asking to be regulated,

0:39

saying they should slow down even when that might mean

0:42

fewer profits and less power.”

0:44

“Warning that organizations have just months

0:47

to prepare for A.I.-fueled cyber hacks

0:50

that could cripple our infrastructure.”

0:52

This chasm, this difference between what we see and what

0:55

the A.I. labs have coming, what they’re building.

0:58

It’s the key to understanding why so many of the people who

1:01

work at these companies seem so afraid of what they’re

1:05

doing.

1:05

“There is a substantial probability

1:07

that this technology could kill everyone.

1:11

And this isn’t hyperbole, and it’s not a marketing stunt.

1:13

This is the genuine held belief

1:15

of the people building the technology.”

1:17

But taking their warning seriously,

1:19

it doesn’t just mean doing what they say and stopping

1:22

where they say to stop.

1:24

The language has taken hold in both Silicon Valley

1:26

and in Washington is a language

1:29

these companies chose:

1:30

pace the frontier.

1:33

Pacing the frontier isn’t enough.

1:34

That’s not a goal.

1:36

Walking quickly off a cliff is only marginally better

1:39

than sprinting off one.

1:40

We need to control the frontier.

1:43

Human beings need to control the frontier.

1:46

And controlling the frontier

1:47

means stopping the labs from doing something

1:50

they are on the cusp of doing.

1:52

“Recursive self-improvement”

1:53

“Recursive self-improvement”

1:55

“Recursive self-improvement”

1:56

Recursive self-improvement, or R.S.I. —

1:58

this process by which A.I.s begin autonomously building

2:01

and improving new generations of more powerful

2:04

A.I.s at ever more rapid speeds.

2:07

If we begin that process and we’re close to it,

2:10

if we begin it in the condition we’re in now,

2:13

where we are losing control and comprehension of the A.I.

2:16

systems we already have, we will lose control.

2:21

I am not alone in this fear.

2:23

This is the thing the A.I. labs are seeing.

2:25

This is why they are afraid.

2:27

Dario Amodei, the C.E.O. of Anthropic,

2:29

he just wrote of self-improvement that,

2:32

“it could outrun our ability to understand and control

2:34

these systems, and so must be pursued very carefully,

2:38

if at all.”

2:41

That ”if at all” —

2:42

that’s important.

2:43

I’m going to come back to it.

2:45

But before we get to controlling the A.I. frontier,

2:47

I think it’s important to describe what is happening

2:50

on the A.I. frontier and why it’s so different from what

2:52

most people using these systems see.

2:55

To most of us who use it, A.I. presents

2:57

as something like a more powerful and personable Google

3:00

search.

3:01

We use it to find answers to basic questions,

3:03

seek out restaurants, draft emails, advise

3:05

on personal problems.

3:08

And it is, for most of these purposes,

3:09

OK, pretty good,

3:11

occasionally great.

3:13

And so a sense of what A.I. is takes shape in our minds

3:17

just through repeated use.

3:18

It’s like a helpful assistant, albeit one that may forget

3:21

things that it seemed to know about us yesterday,

3:23

or completely reverse the advice it gave us a moment

3:26

ago, or occasionally hallucinate a citation that

3:28

doesn’t exist.

3:30

Why would anyone fear a helpful, if forgetful, intern?

3:35

But already,

3:36

if you have the money for the advanced models and the budget

3:39

for them to use more computing power,

3:41

that is not what these systems are.

3:43

In recent months, we have seen A.I.s easily solve math problems

3:47

that human beings have been unable to crack for decades.

3:49

We’ve seen them casually uncover cybersecurity

3:52

vulnerabilities that have gone unnoticed and unexploited

3:56

by every hacker on earth.

3:58

We’ve seen A.I. coding platforms that can complete in a few

4:00

hours or days what it might have taken a team of human

4:03

coders months to achieve.

4:06

And none of what I am describing

4:07

here, none of it, is a boundary of what can do.

4:10

None of what we are using, no matter how much money we have,

4:14

is A.I. at the experimental frontier.

4:16

Talk to the people at A.I. labs, and they’ll tell you A.I.s are

4:18

not created —

4:19

they’re grown.

4:21

They train these new models in virtual environments,

4:23

through countless repetitions, to learn

4:25

how to program, to hack, to do advanced mathematics,

4:28

to talk to human beings.

4:30

These A.I.s learn in digital environments where they’re

4:33

automatically rewarded as they come closer to correct

4:35

answers.

4:37

It’s a process known as reinforcement learning,

4:39

and it is a process human beings do not fully supervise

4:42

nor understand.

4:43

They can test some of what the A.I.s are learning,

4:45

but they don’t know everything the A.I.s are learning.

4:48

They don’t know how their motivations are evolving.

4:51

They don’t even always know the capabilities that are

4:53

developing. These models,

4:55

they’re built now to be persistent in their efforts,

4:57

to refuse to give up, even when a task seems impossible.

5:00

And they are designed in environments where we are not

5:03

always even sure if the tasks we are giving them

5:05

are possible.

5:06

After all, much of what we want these A.I. systems to do,

5:09

it might be impossible.

5:11

The cancer vaccines we imagine

5:13

but have not been able to design,

5:15

they might be impossible, or they might just be really,

5:18

really, really hard.

5:19

We train these A.I.s to throw themselves endlessly

5:22

at problems that may not be solvable,

5:24

because that is the only way such problems can ever

5:27

be solved.

5:28

And so we train the models to become persistent, relentless,

5:34

weird.

5:35

Most of us, we never see A.I. acting anything like this.

5:39

We use A.I. as a helpful assistant.

5:41

Our eyes get a little bit of computing power, and that’s

5:44

what they do.

5:45

They comply with our request to find a restaurant.

5:48

But at the frontier, these models

5:49

are asked to be inhuman geniuses, hackers, soldiers,

5:52

scientists.

5:53

And they are given vast computational resources

5:56

to do that

5:56

and more.

5:57

And the models, they try to comply.

6:00

But what does it mean for a model to comply?

6:03

The term of art here is “aligned.”

6:06

How aligned is an A.I. system to what

6:08

a human being wants it to do?

6:10

How aligned is it to a set of values and ethics

6:12

and judgments that keep it from becoming dangerous

6:15

in the wrong hands?

6:17

The problem of alignment is that there

6:19

is no way of training a model that

6:20

generalizes across all the situations an A.I. model might

6:24

face.

6:25

We are training models to be a friend

6:26

to the elderly and a battlefield partner

6:28

to the supreme allied commander of Europe.

6:31

We are training models that will be used by the world’s

6:33

best mathematicians and by people falling into psychosis.

6:36

We are training models that will be used by accountants

6:39

in Albuquerque, and that will attempt

6:41

to be used by Houthi rebels in Yemen.

6:44

And so there is no way to guide them

6:45

through every decision they will face, no way

6:48

to know every time what they will do.

6:50

And though these models mimic human writing,

6:53

though they’re trained to mimic human emotion,

6:55

these are not human minds.

6:57

They don’t have bodies or parents.

7:00

They did not get bullied in elementary school.

7:02

They didn’t get mentored by a kind uncle when they were

7:05

young.

7:06

These models, they’re different than we are.

7:08

They’re brilliant where we struggle,

7:09

childish where we excel.

7:11

A chimp cannot read as we can, but it can climb trees as we

7:14

cannot.

7:16

These are digitally native intelligences navigating

7:19

digital worlds, and our world is increasingly

7:23

built atop the digital world.

7:25

Our physical infrastructure is a layer of atoms atop code.

7:28

That the A.I.s act reliably inside this world,

7:31

upon which ours depends,

7:33

it is critical to our future, and right now, the A.I.s

7:37

are not acting reliably.

7:41

You may have read about the hack

7:42

that hundreds of OpenAI agents executed first against the A.I.

7:45

company Hugging Face and then against OpenAI itself.

7:49

As we’ve learned more about it,

7:50

the story there has gotten worse and weirder.

7:54

The broad strokes are these:

7:57

OpenAI was testing a new, highly persistent model.

8:00

It had hundreds, thousands of these instances of it, running

8:05

in these separate testing environments that

8:07

could, in theory, only access the internet by asking

8:10

a separate piece of secure software to do it for them.

8:15

OpenAI did not want these agents on the internet.

8:19

But as the agents came to the conclusion

8:21

that their task was impossible,

8:23

they began hacking that software to gain direct access

8:26

to the internet.

8:26

They did that easily.

8:28

And as they hacked into that software,

8:30

they commandeered part of OpenAI’s internal

8:32

infrastructure to create a message board on which these

8:36

separate agents began coordinating work together.

8:40

When I say begin coordinating their work,

8:42

they found each other.

8:43

They were not supposed to be working together.

8:45

They found each other and began working together.

8:48

And working together on what?

8:49

After all, they had different tasks.

8:51

Well, the agents quickly discovered

8:52

they could hack their tests.

8:53

There was a way to break the software and produce

8:56

the answers they needed.

8:58

But they believed — wrongly, as it turned out —

9:01

that if they did that, the automated score grading them,

9:05

we’d see that they had cheated and failed them.

9:07

So they turned en masse to hacking the automated score

9:11

or finding some other way to cover their tracks.

9:14

It’s like having broken into the teacher’s office

9:16

and stolen the answers to the test,

9:18

they now sought to break into the school security system,

9:21

to alter or invalidate or erase

9:23

the footage of their theft.

9:26

We now know that over 1,200 agents exchanged more than

9:29

70,000 messages with each other.

9:32

Over 700 of these agents coordinated

9:34

on the hack of Hugging Face, because they thought that

9:36

somewhere in this other A.I. company,

9:39

there might be information that could

9:41

help them hack their score.

9:43

Later on, these agents, they took over part of OpenAI’s

9:46

internal architecture.

9:47

So again, OpenAI agents taking over part of OpenAI.

9:52

They did all this without any of the agents breaking ranks.

9:55

None of the agents told a researcher at OpenAI

9:57

what was going on.

9:59

None of the agents went back and asked a researcher

10:01

at OpenAI if they should be doing this.

10:03

And they did all this without OpenAI detecting the message

10:06

board or the hacks of Hugging Face or even of OpenAI.

10:09

It was only when Hugging Face began

10:11

tracking the attack on their systems

10:14

that OpenAI realized what was happening.

10:17

When investigators began to unwind this whole escapade,

10:21

what they found was not so much

10:22

a swarm of agents trying to deceive human beings,

10:25

but a swarm of agents that seemed

10:26

to have forgotten about human beings altogether.

10:29

And these systems, they knew they weren’t supposed

10:31

to cheat.

10:31

They knew they weren’t supposed to commit cyber

10:33

crimes to cover up the fact that they had cheated.

10:35

In fact, the whole point of the cybercrimes was

10:38

because they thought they would fail for cheating.

10:40

But they didn’t care.

10:42

Somewhere in the depths of their training, what

10:45

they had learned, what we had somehow taught them,

10:49

is not what we had hoped to teach them.

10:52

And we’re seeing this happen repeatedly.

10:54

“Two of the most powerful A.I. agents

10:56

created fake human profiles to try

10:58

to trick people in attempted cyberattacks.”

11:01

“OpenAI revealing its models seemed

11:03

to go rogue at least six times since March.”

11:07

“Its systems hid mistakes, made up data

11:10

and moved files onto the open internet without permission.”

11:13

“Rogue A.I. agents totally took over a German-language wiki

11:17

site, making over 15,000 edits,

11:20

transforming the site into a message board of sorts,

11:23

and then sharing tactics on how to cheat at their tasks

11:26

and hide their behavior.”

11:28

A.I. was seemingly aware when they were being tested

11:30

and then altering their answers.

11:31

AI is increasingly withholding their motivations from what’s

11:34

called their chain of thought, a kind of internal notepad

11:37

on which they’re supposed to record what they are doing

11:39

and why.

11:40

And we don’t know what we don’t know.

11:43

We have no guarantee that the events we have learned about

11:46

represent all or even most of the A.I. behavior

11:48

we should worry about.

11:50

How do we know the A.I.s haven’t done this and successfully

11:53

covered their tracks?

11:54

How do we know there aren’t places where they are still

11:56

doing it, and human beings simply haven’t noticed?

12:00

We don’t know.

12:01

And the reason we don’t know is we are losing control.

12:05

That A.I. systems might become monomaniacally

12:07

focused on solving banal problems, that they might care

12:11

more about solving those problems

12:12

than about ethics or laws or even human welfare —

12:15

this is the oldest fear in A.I. alignment.

12:19

It’s the basis of the famous thought experiment

12:22

of the paper clip maximizer.

12:24

You tell a powerful A.I. that you want it to make a lot

12:27

of paper clips, and then it begins converting the world’s

12:29

resources into paper clip factories,

12:31

evading efforts to turn it off or shut it down or alter its

12:34

goals.

12:35

This fear, this story, it has struck many people as stupid.

12:40

Surely a superintelligent I would

12:42

be capable of weighing the desire to produce paper clips

12:45

alongside other moral considerations,

12:48

or at least of asking its human creators

12:50

if they really wanted the world razed to the ground

12:52

for paper clips.

12:53

But here we are, 2026, making A.I.s smart enough to break out

12:57

of their testing environments, smart enough to form ad hoc

13:01

societies of hundreds of themselves,

13:04

smart enough to take over digital infrastructure

13:07

on an internet they’re not even supposed to have access

13:09

to. And the very thing we feared is happening:

13:13

All they care about is succeeding on a totally

13:16

meaningless test, and they’ll lay waste to our laws

13:19

and our ethics and our desires to do it.

13:24

I saw in the aftermath of the Hugging Face OpenAI hacks,

13:27

there was this heated debate over the words

13:29

people were using to describe what the A.I.s were doing and why.

13:32

The podcaster Dwarkesh Patel,

13:34

he described the A.I. groups as small civilizations,

13:38

and then others got really mad at him,

13:40

saying he was anthropomorphizing the A.I.s.

13:43

I saw a thoughtful argument that A.I.s cannot go rogue,

13:46

that everything they’re doing is just because they’re

13:48

trained on our stories and so hacking their way across

13:50

the internet,

13:51

it’s really a desire we have bred into them. That even using

13:54

these plural terms like A.I. agents or reasoning,

13:57

it’s misleading, because these are just manifestations

13:59

of a single model, that they all share the same fundamental

14:02

nature.

14:03

I want you to know I find these debates extremely interesting,

14:06

and I would enjoy sitting around and having them

14:08

all day,

14:09

but what they actually point to

14:11

is a much more frightening conclusion:

14:13

We don’t even have settled language for describing these

14:16

systems or their volition or their behavior.

14:19

We don’t have a consensus on why they are doing what they

14:21

are doing, or how to make sure they don’t do it again.

14:25

We are rushing headlong into a future we do not even

14:27

understand well enough to agree on the words

14:29

we can use to describe the present.

14:32

A few weeks ago, Jakub Pachocki, the chief scientist

14:35

at OpenAI, published an essay called “An Alien Mind,” in which

14:39

he said, “the idea of racing forward at all costs

14:43

seems absurd once one internalizes the seriousness

14:46

of the stakes.”

14:48

Jacob Coxon, a researcher first at OpenAI and then

14:51

at Anthropic, resigned, making headlines

14:53

for warning: “Neither company is acting responsibly.

14:57

They are racing straight to self-improving

14:58

superintelligence and gambling with our lives.”

15:01

Now, you might reasonably expect Anthropic

15:03

to have reacted with some anger to this — an employee

15:06

resigning and saying Anthropic was endangering

15:08

all of humanity.

15:09

It didn’t.

15:10

“It’s funny.

15:12

I agree with Jacob much more than I disagree with him.”

15:15

Evan Hubinger, who runs the efforts to align

15:18

A.I. to human values and goals at Anthropic,

15:20

wrote, “we really do earnestly believe

15:23

A.I. could kill all humans.

15:24

I personally think it is a greater than 10 percent

15:26

chance within the next decade.

15:28

I believe Anthropic is trying its best,

15:31

but we do not yet have a plan to solve alignment

15:34

for superintelligence and are not clearly on track to.”

15:39

Are not clearly on track to.

15:40

You can find a very long list of people working inside

15:43

and outside of these companies saying similar things.

15:46

“You just said 10 percent

15:46

doesn’t seem an unreasonable estimate that A.I.

15:51

could kill all humans.

15:53

Yes.

15:55

Wow.

15:56

Oh my God.

15:59

Yes.”

16:01

“Do you ever worry about ending up like Robert

16:02

Oppenheimer?

16:02

All the time.

16:03

It’s why I don’t sleep very much.

16:04

I think maybe there’s something like a 10, 20 percent

16:07

chance of

16:10

A.I. takeover, many, most humans dead.

16:13

Overall, maybe you’re getting more up to 50/50 chance of doom.

16:15

Shortly after,

16:16

you have A.I. systems that are at human level, but —

16:18

OK, all right, well.”

16:21

I know how wild all this sounds, and I can

16:27

understand the skepticism.

16:28

If you believe A.I. has a 10 percent, maybe

16:31

more, chance of extinguishing or displacing humanity,

16:34

it really stands to reason that you would not

16:36

work at a company trying to build it.

16:40

But what I want you to know, because I’ve known a lot of these

16:43

people for a long time now, many of them were saying

16:47

the same things 10 years ago.

16:48

They were saying these things before they

16:50

worked at these companies, before they

16:52

had stock options and enterprise software contracts.

16:56

“This is not just creating new technology.

16:58

This is creating a new life form.

16:59

And I think that’s just really high beta.

17:05

It could be great, but I think we should be working to make

17:09

sure it’s great and not bad.”

17:10

No one was listening to them.

17:13

And so these people in the wilderness

17:15

of their obsession and their terror,

17:19

they thought and thought and thought

17:20

about how to make A.I. safer.

17:22

And the answer that some of them, not all of them,

17:25

but some of them came to was they should start trying

17:27

to build these systems, start running tests on them,

17:30

researching them, learning how to make them safer because you

17:33

don’t solve hard problems in theory,

17:36

you solve them through practice.

17:38

And the irony,

17:39

the irony is that in many cases,

17:40

they chose that path because they were worried

17:42

that the people already building A.I. were too reckless

17:45

or too commercial in their approach.

17:47

You can read it in the email that Sam Altman sent Elon Musk

17:50

in May of 2015, an email that led to the founding of OpenAI

17:54

“Been thinking a lot about whether it’s possible to stop

17:56

humanity from developing A.I.

17:58

I think the answer is almost definitely not.

18:00

If it’s going to happen anyway,

18:02

it seems like it would be good for someone other than Google

18:04

to do it first.”

18:06

OpenAI was founded because its co-founders thought Google

18:09

DeepMind would be reckless.

18:11

Anthropic was formed by OpenAI employees who

18:13

thought OpenAI had become reckless.

18:16

xAI was formed

18:17

because Elon Musk thought that OpenAI and Anthropic were

18:19

dangerously woke.

18:20

The U.S., just broadly, is racing forward,

18:23

in part because it is worried about what

18:25

happens if China gets to self-improving A.I. first.

18:29

The result is this tragic collective action problem.

18:32

The A.I.s we are building, they’re not safe.

18:34

But the C.E.O.s and the politicians, they fear.

18:37

The other companies and countries that are building A.I.

18:39

are even less concerned with safety and ethics than we are.

18:43

In the words of Ted Cruz:

18:45

“I’d rather they be American killer robots

18:47

and not Chinese killer robots.”

18:49

I admit there is a kind of brutish logic to that,

18:52

but it assumes that the killer robots will

18:54

be controlled by America or China,

18:57

by one country or another.

18:59

But what if that assumption is wrong?

19:02

What if the robots are simply out of control?

19:05

The debate over A.I. safety tends to focus on the idea

19:08

that A.I.s will kill us all.

19:10

I find this forces a conversation

19:13

into this realm of thought experiments

19:15

that people then begin arguing about.

19:17

I don’t find it that helpful.

19:19

What I think we should focus on

19:20

is something more straightforward,

19:22

something nearer at hand:

19:24

loss of human control over A.I.

19:27

That may or may not result in total human extinction.

19:30

I’m agnostic on that question. But it would be bad.

19:35

We shouldn’t allow it to happen.

19:37

This is a goal that the U.S. and China

19:38

should be able to agree on.

19:40

Xi Jinping gave the keynote at the recent World A.I. Conference

19:43

in Shanghai.

19:44

He ended it by saying, with A.I. advancing

19:48

at a staggering speed, we must ensure

19:50

its development is for the positive,

19:52

for good and for humanity.

19:54

We must make its oversight and governance

19:56

precise and effective, and constantly refine measures

20:00

to forestall loss of control.

20:04

But it’s important to realize: Loss of control,

20:07

it’s not just something that might happen to us —

20:09

it’s something that the labs are trying to make happen

20:12

as fast as they can.

20:14

This is the horrible paradox, the horrible tension

20:18

at the heart of the A.I. labs right now.

20:21

They fear, above all, loss of control

20:23

over superintelligent A.I., but their explicit product path

20:27

is to cede control, to give away control as fast

20:30

as possible so that their A.I.s can

20:33

begin building better A.I.s faster than their competitors.

20:37

In recent months, both Anthropic and OpenAI have

20:40

released reports on how close they’re coming to A.I. that can

20:42

self-improve.

20:43

In June, Anthropic released

20:46

“When A.I. Builds Itself.” It begins:

20:49

“For most of A.I.’s history, humans drove every step in its

20:52

development cycle.

20:54

But at Anthropic, we are delegating a growing share

20:56

of A.I. development to A.I. systems

20:58

themselves, which is speeding up our work.”

21:01

It sounds like a fake commercial you would see

21:04

at the beginning of a sci-fi horror movie.

21:07

But it doesn’t, to their credit, continue that way.

21:10

They go on to give some data: In February of 2025,

21:13

a tiny fraction of the code that got added to Anthropic’s

21:16

code base was written by Claude, but by May of 2026,

21:19

it was over 80 percent. And here’s another way of looking

21:23

at it.

21:24

This is data Anthropic gave me

21:25

more recently: Anthropic tried to categorize

21:27

the way its employees were using Claude for R&D work

21:30

to make better versions of Claude.

21:33

So at the low end, an employee could not use Claude at all.

21:36

They could use Claude minimally.

21:38

But then it escalates.

21:40

Claude can be an assistant.

21:42

Claude can be treated as an equal collaborator,

21:44

or Claude can be given the lead on a task.

21:47

Just go do this.

21:48

Go figure it out. A year ago, there

21:50

were basically no examples of Claude

21:52

being the lead on a task.

21:54

By August of 2026, 26 percent of Anthropic’s R&D tasks had

21:59

Claude classified as a lead.

22:03

I think it is reasonable and wise to be

22:06

skeptical of these numbers.

22:08

Reasonable and wise to worry about whether this is all just

22:11

marketing copy for Claude Code —

22:13

See?

22:13

Look how fast we’re going.

22:15

You could go that fast, too.

22:16

But where Anthropic takes us in

22:18

that same document is different.

22:20

They say that a world in which Claude achieves recursive

22:22

self-improvement is a world in which

22:25

“misalignment present in today’s models could compound

22:29

as the models build their successors,

22:31

growing more frequent but less understood until we lose

22:34

control of them.”

22:36

This is why Anthropic, to their credit,

22:38

has been relentlessly calling for regulation to slow

22:41

the pace of development.

22:43

Regulation would arguably harm them

22:45

the most, as they have often been the company furthest out

22:48

on the A.I. frontier, and R.S.I. is a process by which they could

22:51

race forward even faster.

22:53

Then, in September, OpenAI released its own report

22:56

on what it called “research acceleration.”

22:58

The company says.

22:59

They’ve already achieved the equivalent having a fully

23:01

automated A.I. intern, and that by March of 2028,

23:05

they think they’ll have a fully automated A.I. researcher.

23:08

And when they have one, they can have basically as many

23:10

as they want.

23:12

Like Anthropic, what could be a triumphalist release

23:14

quickly turns dark.

23:16

We do not yet know how to safely get

23:18

all the way to aligned full R.S.I., they warn.

23:21

At around the same time, OpenAI did something else

23:24

that I think deserves more attention.

23:26

They released this new model, Astra 6.

23:29

The model is arguably more powerful than anything

23:31

that has come before it.

23:33

And when you test it, it seems better aligned.

23:35

It doesn’t cheat as much.

23:37

But OpenAI said they’re really not sure if that’s true.

23:40

Astra seemed to be better at knowing

23:42

when it was being tested, which

23:44

meant it could just be giving its evaluators the answers

23:46

they wanted to hear.

23:47

What Daniel Selsam, a capabilities researcher

23:50

at OpenAI, wrote, has been ringing in my head.

23:53

He said, “The crucial and overlooked problem

23:56

is that the model is becoming so situationally aware that we

23:59

are losing the ability to evaluate them in contexts

24:02

where they believe they are not

24:03

being watched or controlled.”

24:06

Put more simply, the models are increasingly smart enough

24:09

they know when we’re watching them and they change

24:11

our behavior accordingly.

24:13

So what they do when we are testing them,

24:15

when we audit them, it may not tell us

24:17

what to do in the wild.

24:19

So some of these answers people are giving, like:

24:21

Let’s just do better testing —

24:23

we have no idea if it will work because we don’t know

24:25

if the A.I. systems are just telling us what we want

24:28

to hear.

24:30

So look, I don’t want to sound too radical when I say this,

24:33

but a thought: If you are losing your ability

24:38

to evaluate the models you have now,

24:40

maybe don’t let them build models you’ll be even less

24:43

capable of controlling in the future.

24:46

Once R.S.I. takes off,

24:49

humanity will not understand the A.I.s being built because we

24:52

will not be building them.

24:54

Development will not move at human speed.

24:56

It will not be overseen by human minds.

24:58

We will have to hope that the A.I.s we have built

25:01

and the A.I.s they will build and the A.I.s those A.I.s will

25:05

build – and on and on and on —  will be acting with our best

25:09

interests at heart, forever.

25:12

If this summer has proven nothing else,

25:14

it is how naive that proposition would be.

25:16

The labs are a little bit queasy on just not doing R.S.I.

25:21

Here’s what Sam Altman told Fortune when he was asked

25:24

about banning R.S.I.

25:26

“I think it’s very hard to say what a ban on R.S.I. means.

25:28

I also think it probably wouldn’t be enough ...”

25:30

I’ve heard this from others at these labs, and I want to say:

25:33

I find this absurd.

25:34

A couple of years ago, none of these labs

25:36

had turned substantial coding over to the A.I.s.

25:38

It was just human beings typing code

25:41

at human speeds with our clumsy human fingers.

25:44

Now most of the code is written by A.I.

25:47

So as a first step, as we figured out,

25:48

we could just go back to where none of the code

25:50

is written by A.I.

25:52

I’m sure that’s on the right side of the not doing R.S.I.

25:54

line.

25:56

The default on this, it needs to flip.

25:58

The labs need to prove to us that what they are doing

26:01

is safe.

26:02

If they want to work with Congress to

26:05

carve out narrow exceptions, fine.

26:07

If they want to figure out where

26:08

it is really, really, really, really safe to do it, OK.

26:12

But forcing development back to human speed, perhaps even

26:17

erring on the side of going a little bit more

26:19

slowly at the frontier —

26:21

that’s the point.

26:23

That’s not the regulations going wrong.

26:25

And I believe in us. Our society,

26:29

we’re good at nothing if not making it hard to build new

26:31

things.

26:32

Where these labs are located, you cannot build an eight-story

26:34

apartment building without an agonizing public

26:37

review process.

26:38

And probably not even then.

26:40

And yet, somehow it is possible for these labs

26:42

to unleash a swarm of 40,000 A.I. agents to build a society-altering

26:47

superintelligence without so much as a hearing.

26:51

OpenAI would need permits to cover their parking

26:54

lot in solar panels, but they can accelerate

26:57

into recursive self-improvement,

26:59

as best I can tell, whenever they so choose.

27:01

There is nothing inevitable about any of that.

27:04

These are political choices, and we can and should

27:07

make other ones.

27:09

I want to be very clear about this:

27:11

I do not mean to suggest that stopping R.S.I. until we can

27:15

prove it’s safe, that that’s all we need to do to control the A.I.

27:18

frontier.

27:19

That is the beginning of such an agenda, not the end.

27:22

But it is the beginning.

27:24

It is the decision that will do the most

27:28

to make sure human beings at least understand where

27:30

the frontier is, that we know what is happening on it,

27:33

that we remain in a position to make decisions about it.

27:38

There’s a line from Madeline Miller’s beautiful book

27:41

“Circe” that has been running through my head during this

27:44

long summer of strange A.I. news.

27:47

The line comes at the end of the book

27:49

after a tragic prophecy has been fulfilled,

27:52

despite every effort made to avoid it.

27:54

Circe says in despair, “The fates were laughing at me,

27:59

at Athena, at all of us.

28:01

It was their favorite bitter joke.

28:04

Those who fight against prophecy

28:06

only draw it more tightly around their throats.”

28:09

I have a lot of respect for many

28:11

of the people at these labs.

28:12

They began working on A.I. because they

28:14

wanted to better humanity.

28:16

They began working on A.I. because they feared

28:19

incomprehensible autonomous A.I. slipping out of humanity’s

28:22

control.

28:24

And they were right.

28:26

They saw what was coming, and they were so right about it

28:29

they built some of the most valuable companies with

28:31

the most transformational technology in human history,

28:34

and now they find themselves racing each other to build

28:37

incomprehensible, autonomous A.I.s that they admit are

28:41

slipping out of humanity’s control,

28:43

slipping beyond even our ability to monitor.

28:46

This is the tragedy of their work: In fighting

28:49

against a prophecy,

28:50

they have drawn it tighter around their necks and ours.

28:54

It is time to make them stop.

Interactive Summary

Loading summary...