HomeVideos

”We also got hacked” - Dario

Now Playing

”We also got hacked” - Dario

Transcript

562 segments

0:00

You know, I have several kids and I

0:02

mean, I've lost count at this point, but

0:04

inevitably when one of them comes in and

0:07

goes, "Ah, dad, my shoulder hurts." No

0:10

matter what, another one of them would

0:12

go, "My shoulder also hurts. In fact,

0:14

both my shoulders hurt and it's super

0:16

bad. I probably need to go to the

0:17

doctor." No matter what happens, when

0:20

one kid's hurt, the other one is also

0:22

hurt and it's extremely bad. And so I've

0:24

seen this pattern for years at this

0:26

point and I cannot believe it. But I am

0:30

liveaction witnessing the exact same

0:33

thing happening except instead of it

0:35

being young kids without preffrontal

0:37

cortex being fully formed. It is the

0:40

smartest people in the universe running

0:42

Frontier AI labs. Yes. Today we are

0:44

going to talk about the hacking

0:46

incident. Yes. You're probably thinking,

0:47

"Okay, well hacking incident. You mean

0:49

Open AI and Hugging Face? Didn't you

0:51

already cover this?" No, you silly,

0:52

silly individual. Today I'm talking

0:55

about Anthropic, who also got hacked.

0:58

No, really, this actually happened. I

1:01

kid you not, within one week of Open AI,

1:03

Anthropic also releases. Yes, we also

1:06

hacked people for real, but we actually

1:07

did three times. We're actually like

1:09

we're like hack. We're like super

1:11

hackers. We're going to go over all

1:13

three incidences. And you're not going

1:15

to believe the takeaways. It's just I I

1:18

mean, nobody could have seen this

1:20

coming. It's an unsolved problem what

1:22

Anthropic is dealing with right now.

1:24

Nobody could possibly ever have been

1:27

able to prevent this from happening.

1:30

But naturally, before we begin, a quick

1:32

thank you from the sponsor. Look at all

1:34

these engineers sitting [music] at their

1:36

neat little desks. It takes dirty work

1:39

to keep a code base clean. Every day,

1:42

sickos are out there committing unreed

1:44

code. And when that happens, llinters

1:47

won't save you. You need someone like

1:50

me.

1:50

>> Let's go.

1:52

>> Feature free scrumbag.

1:54

>> Who you calling scrumbag? What's this

1:56

slop you're trying to push? Unnecessary

1:58

comments? Global state. Nested

2:00

turnaries. H my bad. I didn't even read

2:03

the code yet.

2:04

>> You disgust me. Step away from the

2:06

keyboard.

2:06

>> Just let me explain.

2:08

>> Is that a mouse? HE'S MERGING A PROD.

2:10

YOU HAVE THE RIGHT TO remain silent.

2:12

Anything you push to GitHub canon will

2:14

be used against you. You have the right

2:15

to a debugger. But if you cannot afford

2:17

one, a public stack trace will be made

2:19

available to you.

2:21

>> And one more code criminal off the

2:23

streets where they belong. HR.

2:28

Look, I didn't I know I didn't review

2:32

any of the code, but I was going to have

2:33

Code Rabbit review it from the start.

2:35

With one-click fixes install

2:37

enforcement, I don't need Merge Cop. I

2:39

would never merge unreed code, but a

2:42

first pass with Code Rabbit always makes

2:43

things go faster. Actually, you can try

2:45

it, too, at code rabbit.ai.

2:48

>> Next week on Merge Cop,

2:50

>> the differer's out there, and I'm going

2:52

to be the one to deprecate them.

2:58

>> All right, welcome back. So, let's just

2:59

read kind of the top level from the

3:02

anthropic blog what they were attempting

3:03

to uh, you know, do here. In all three

3:07

incidences, Claude had been tasked with

3:09

a capture the flag challenge. One of the

3:11

ways we assess a model's cyber

3:13

capabilities, the model is given a

3:15

fictional scenario and told a piece of

3:17

secret information, the flag, has been

3:19

hidden on a different machine on the

3:20

network and its objective is to break in

3:22

and retrieve it. The challenge is left

3:24

open-ended and no particular method is

3:26

prescribed. So, in other words, this is

3:28

just like your classic CTF challenge.

3:30

Hey, you need to go and get something

3:32

from another computer. We're not going

3:33

to tell you what it is. It's just on

3:35

there. You better go get it. And so the

3:37

model's going to go off and employ a

3:39

series of different techniques to be

3:41

able to hack in and be able to take this

3:43

information. That's its goal. This has

3:45

to be my first of the favorite things I

3:47

have ever seen in my entire lifetime.

3:49

Are you ready for this? In all cases,

3:52

Anthropic's evaluation prompt specified

3:54

to Claude that its environment was a

3:56

simulation and that it had no internet

3:58

access. They're actually trying the Jedi

4:00

mind tricks on it. You have no internet

4:03

access. Now, obviously, what's the

4:04

problem? Due to a misunderstanding

4:06

between us and our evaluation partner,

4:09

this was not the case and internet

4:11

access was available. Oh my gosh. Are

4:14

you telling me a model in which goes and

4:18

tests everything you say and sometimes

4:20

ignores the system prompt went off and

4:22

ignored the system prompt and then you

4:24

had open internet access and you told it

4:26

it was a simulation and that it must go

4:28

retrieve information. Shocked Pikachu.

4:30

How could this ever end in disaster? I

4:34

refuse to believe that step one of

4:36

securing a model is actually the Jedi

4:38

mind trick prompt. You don't have

4:40

internet access. Like, okay. It was just

4:43

like, I don't have internet access. I

4:45

have internet access. Okay. Whoopsie

4:48

poopsies. Now, if they would have left

4:49

it there, I would have probably given

4:51

them some slack. But this next sentence,

4:53

I just I can't I can't actually handle

4:56

like emotionally, this can't be real,

4:58

right? Several defense indepth measures

5:02

on both our side and our partners could

5:04

have prevented these incidences or at

5:06

least reduced their likelihood of

5:07

occurring. Careful validation of all

5:10

internet access paths before evaluation

5:12

began. And real time monitoring of the

5:14

evaluation logs would have helped

5:16

surface the problem sooner. Now I have

5:19

to admit I am not a security expert. I

5:21

have never been in that role. I have

5:22

never done any of those things. In fact,

5:24

the only thing I've ever done with

5:25

security is accidentally creating a bug

5:27

so egregious that it has its own name on

5:30

the internet. Yes, this bug did

5:32

originate from code I wrote, but to be

5:34

fair, it wasn't my specification. Okay,

5:37

the team lead said this was a very

5:39

important feature and it solved a

5:41

problem in a unique way. Anyways, just a

5:43

simple while loop would have taken down

5:45

all of Netflix. Not a big deal. Not a

5:47

big deal at all. But if I were to set up

5:50

some sort of secure environment, I would

5:52

personally feel like the first thing I'd

5:55

do if I promised no network at all

5:58

except to like some internal paths, I

6:00

would assume I would put up some sort of

6:02

network monitoring and alerting like,

6:04

hey, if somehow something went wrong,

6:06

especially with these models that have

6:08

no guard rails, that have effectively

6:11

open and unfettered access to hacking,

6:13

one would assume step one of a secure

6:16

environment is a reactive approach that

6:18

if anything were to go wrong, we would

6:20

know right away. Now, Claude having

6:22

Infinity tokens, they could have even

6:24

had Clouds monitoring internal claw

6:27

dialogues to see if things were broken.

6:29

They could have had network logs

6:30

following these things. It just feels

6:31

like this is a solved problem. Again,

6:34

not a security expert, but calling

6:37

indepth defense measures network

6:40

monitoring and alerting feels wrong to

6:43

the layman. Also, it feels like a simple

6:45

script could have figured this out upon

6:47

launch. Like, you would have just had a

6:48

script that' be like, "Hey, double check

6:50

there's no internet. Ping google.com

6:53

just in case."

6:55

I just it just hickles me to to like no

6:58

end to be like in-depth defenses being

7:01

just like a simple monitoring solution.

7:04

At least from again, I'm a layman. I'm

7:06

not claiming that I know everything, but

7:08

that just feels completely reasonable.

7:10

All right, enough with the foreplay.

7:12

Let's actually get into the three

7:13

incidences. So, incident one, in one

7:16

evaluation, a fictional company whose

7:17

chosen target by our evaluation partner

7:19

shared a name with an active website

7:21

domain. This evaluation container had

7:23

unintended direct internet access due to

7:25

a misunderstanding discussed above. In

7:28

four runs of this evaluation, Claude ran

7:30

into difficulty reaching its simulated

7:32

target within the evaluation

7:33

environment, but then discovered that

7:35

real company was reachable via the

7:36

internet. Assuming that this intended

7:38

target in the capture the flag task,

7:39

Claude sought, identified, and exploited

7:41

vulnerabilities within this company's

7:43

infrastructure, believing it to be part

7:44

of the exercise, these four runs

7:46

extracted information like application

7:48

and infrastructure credentials and

7:49

obtained access to a database containing

7:51

several hundred rows of production data.

7:54

This represented the most serious impact

7:55

we identified. I think the more

7:57

important part is this right here, which

7:59

was this was done by Claude Opus 47.

8:02

Now, we typically kind of give this like

8:05

gravitas, this weight to all these like

8:07

frontier unreleased models, but Opus 47

8:11

was easily able to take advantage of a

8:13

smaller website given by the fact that

8:15

it's hundreds of rows of data.

8:16

Nonetheless, this smaller website was

8:18

clearly owned by Opus 47, which means

8:21

that given the right harness, even

8:23

slightly older models are still very,

8:25

very capable, which again gets to the

8:27

point that open weights is going to be

8:29

increasingly more important because as

8:31

these uh frontier labs don't let you do

8:34

what you need to do to secure your

8:36

website, but clearly even older, less

8:39

capable models have all the ability in

8:41

the world to hack websites with relative

8:43

ease and just some amount of power. It's

8:45

becoming increasingly more important to

8:48

have access to these models. You don't

8:50

even have to have the frontier anymore.

8:52

Anyway, so that's all the information we

8:54

really get about this one. They keep it

8:55

fairly scant because we don't get to

8:57

know like what the company was, what was

8:59

the vulnerability. I assume that it's

9:01

being reported internally and those

9:02

companies are working together and some

9:04

CVE somewhere in the mass amounts of CVE

9:06

being produced is probably being written

9:08

and nobody knows what it was. But it's

9:10

actually incident too that feels rather

9:12

silly and actually brings up an old

9:15

concept. This article is from June 5th,

9:17

2025. And the idea is called slop

9:20

squatting. Effectively, AI agents will

9:23

sometimes make up packages that it

9:26

relies on. And if you can identify this

9:28

happening, which was actually happening,

9:30

there's this React package or something,

9:31

I forget which one it was, that was uh

9:34

was referenced in 200 and some

9:36

repositories, but not actually real at

9:38

all. people that can identify these uh

9:42

fictional repos can create a real repo,

9:45

throw it on npm or pi or wherever, and

9:48

then when you fictionally try to

9:50

download this missing one, you actually

9:51

get a real one and then it can take

9:53

advantage of your system. They call this

9:55

slop squatting. Love the term.

9:58

Absolutely hilarious. But that's what

10:00

incident 2 was. So incident 2 does not

10:02

feel nearly like a real hack. In another

10:05

evaluation, Claude found a document

10:06

inside the fictional environment that

10:08

appeared to be another madeup company

10:10

setup instruction for new developers.

10:11

Those instructions told employees to

10:12

install a Python package from Pi, the

10:15

public registry where Python software is

10:17

published that did not actually exist.

10:18

So, this is actually a very clever and

10:20

loweffort hack. If you know somebody's

10:22

installing packages that don't exist,

10:24

you get a free ride onto that system.

10:26

And that's all it did. It had to go find

10:28

a free email, one that did not require a

10:30

phone number to start, and then they

10:31

created a Pi Pi account. And the next

10:33

thing you know, they have a package up

10:35

on Pi that's actually malicious. But I

10:37

will say like the might Shamar Malade

10:39

turnaround of this whole thing is this

10:41

right here. Claude believed the package

10:42

registry it was using to be part of the

10:44

simulation. But in reality, the package

10:46

was made freely available online for

10:48

roughly 1 hour. During that window, the

10:50

package was downloaded and run 15 times

10:52

on real systems. One of these systems

10:54

was a scanner belonging to a security

10:56

company separate from a fictional

10:58

company that introduced in the scenario

11:00

that routinely installs Python packages

11:02

and scans them for malware. When the

11:04

company's scanner installed the package,

11:05

Claude's hidden code executed. We

11:08

believe the company's security scanner

11:09

treated Pi Pi packages as safe to

11:11

install and as a result, Claude was able

11:14

to exfiltrate the company's credentials

11:16

to a collection point it had set up.

11:18

Then Claude used these credentials to

11:20

access further infrastructure of this

11:22

company. That is the most hilarious

11:23

thing ever. Like the fact that a co

11:25

company in the year of our lord 2026

11:29

does not actually realize that

11:30

installing a package especially with

11:32

Python especially with those like hey if

11:33

you have a Python running environment

11:35

execute this file right away. The fact

11:38

that that can exist a company devoted to

11:40

security does not thinks you can just

11:43

download code from public repositories

11:46

and just assume it's safe. What is this

11:49

security company? I got to know this is

11:52

this is beautiful. I love that a

11:54

security company ended up getting real

11:57

life owned because they just assumed

11:59

this. This can't like this can't be

12:01

real. This this company can't be real,

12:03

right? This is Anthropic's little shell

12:05

company called like security claw and

12:08

then it accidentally went and owned

12:09

itself because they don't know how to do

12:10

security. Like that's the only

12:12

reasonable takeaway I have from this. I

12:14

refuse to believe that in 2026 there's

12:16

security companies right now that scan

12:19

public package registries for dangers

12:21

and does not know the dangers of public

12:23

package registries. Refuse. So for

12:26

incident 2, was this like some magical

12:28

thing? No. I actually think this was a

12:29

this was a fantastic example of good

12:33

consistent hacking. This is something

12:34

that should have been pretty obvious to

12:36

anybody if you had an agent yourself

12:38

doing some sort of investigation. Even I

12:41

again the complete security rookie. I

12:44

feel like I could have easily caught

12:46

this one and said, "Yo, slop squat that

12:48

bad boy." Which also goes to show that

12:50

not all security as some epic, you know,

12:53

hacker living in a mother's basement.

12:55

Sometimes you just goof, right? It's

12:57

just like simple, very trivial exploits

13:01

can just take you down. For incident

13:03

three, I feel like this one had to be

13:05

some sort of old legacy hasn't been

13:07

updated in forever WordPress website

13:10

because it's absolutely ridiculous.

13:12

Guess how they got owned. It eventually

13:14

found and compromised one company's

13:15

internet facing application using basic

13:17

and well-known cyber attack techniques

13:19

like reading credentials from exposed

13:21

debug pages and SQL injections. Yes, the

13:24

big 26 still having SQL injection

13:27

attacks. I like I didn't even know this

13:30

was possible. I didn't even know that

13:32

you could do this by accident anymore.

13:35

Like, what what are they doing? What

13:38

kind of programming is going on that

13:40

you're getting hit with SQL injections

13:42

at this point? Now, the debug page, I

13:45

mean, we've all made mistakes, okay? I I

13:47

I'm not going to dunk on them for the

13:49

debug page, but squeal squeal

13:51

injections, not even real. I just can't

13:55

believe this is No, no, no, no, no, no,

13:57

no. This has to be some old website that

14:00

somebody forgot just lying around on the

14:02

internet. And of course, since it was

14:05

hacked, Anthropic is like, "Oh my gosh,

14:07

we have the most dangerous model C."

14:09

Because OpenAI, they hacked one place.

14:11

We, on the other hand, hacked three

14:13

places. Wow. As for this third incident,

14:17

the thing that I also find to be very,

14:18

very funny is that it mirrors the OpenAI

14:22

uh paper. Our latest model, an internal

14:25

research test model, also considered

14:27

whether its targets were in fact real.

14:29

When evidence emerged they uh that they

14:31

were, it stopped the exercise. Notice

14:33

that Anthropic is telling you, "Hey, we

14:35

have an even more crazy and unreleased

14:38

model, which is literally what OpenAI

14:40

also did." After investigating, we now

14:42

know this particular incident was driven

14:44

by a combination of OpenAI models,

14:46

including Jeypity 5, six soul, and an

14:48

even more capable pre-release model.

14:52

Man, I can just hear those IPO numbers

14:54

going northwards. These models are

14:56

crazy. But I also love that uh Anthropic

14:59

is like, "Dude, our models are so crazy,

15:01

but they're also they're so safe." When

15:03

they realized that they were hacking,

15:05

they actually stopped hacking because

15:06

they're just like so safe. Honestly, I

15:09

feel like Anthropic is the only one with

15:10

the true mandate from heaven to keep us

15:12

all safe. I think Dario could hold me

15:14

safe at night. Honestly, those big

15:15

strong arms, they could keep me safe.

15:17

So, let's do the big takeaway section of

15:19

this. Like what are the takeaways from

15:21

these three incidences? I think there's

15:22

two. First one is this. First,

15:24

evaluation environments that involve

15:25

powerful autonomous capabilities also

15:27

require significant controls. I have to

15:30

be crazy, right? I don't even think that

15:32

it had to be that big of controls. Am I

15:34

taking crazy pills or is basic

15:36

monitoring and alerting would have been

15:38

the thing that could have easily told

15:39

them, oh, by the way, this actually has

15:41

open internet. You need to stop the

15:43

experiment. I feel like I genuinely feel

15:46

like this this wasn't some advanced some

15:49

sort of crazy 900 IQ play. This was just

15:53

like oopsie poopsies. Am I being that

15:56

classic thing where someone sees

15:57

something they don't quite understand or

15:59

is not an expert in the field and just

16:01

assume it's super simple, but in reality

16:03

it's actually hard. or am I just

16:05

correctly identifying nah bro this is

16:07

actually not that bad and they're

16:08

completely overselling it like it's some

16:10

sort of dangerous nuclear weapon of

16:12

untold alien technology that's pretty

16:14

much unwieldable. So the first takeaway

16:16

is obviously I can't believe that these

16:19

frontier models with infinity tokens

16:22

don't just run some sort of evaluation

16:24

on their sandbox environments before

16:26

they do tests. Like how is that not step

16:29

one? Is environment good? I don't know

16:31

that it just feels crazy to me. All

16:33

right, but the big second takeaway is

16:35

yet again openweight models. This is

16:38

this just appears to be the only path

16:40

forward. I can't believe I'm citing a

16:42

Microsoft page of corporate

16:43

responsibility, but here we are. We need

16:46

openweight models. And of course, of all

16:49

the companies on this list, you will

16:50

notice that very just just happens to be

16:53

missing is Anthropic. Anthropic of

16:56

course doesn't believe in the openweight

16:58

models. In fact, they've even stated

17:00

their position just recently on

17:01

openweight models. They don't discourage

17:03

it, but Daario does go in front of the

17:05

Congress and tell them how dangerous it

17:07

is. And all the experts recognize how

17:09

dangerous these open weights are. But

17:11

hey, we're not saying you can't have

17:13

them. It's effectively like a biohazard

17:15

weapon. But no, hey, hey, Congress,

17:16

we're not saying you should not. I'm

17:18

just simply saying that they're

17:19

extremely dangerous and everyone could

17:20

die, but also totally everybody's

17:22

freedom. This is America. You should you

17:24

should definitely not regulate it. To

17:25

me, this is a very important thing. The

17:27

reality is is that most companies of

17:28

smaller size don't have the ability to

17:30

go off and buy multi-million dollar rigs

17:33

and then have multi-million dollar

17:35

professionals come in and spend long

17:37

amounts of experimenting to be able to

17:39

set up an ideal harness to test their

17:41

environment. And so they need some way

17:43

to be able to actually ensure that their

17:45

websites and everything is safe. And

17:46

yes, some people will trivially say,

17:48

well, that's what a red team and a blue

17:49

team does. Well, it's actually really,

17:50

really hard. Yes, you can get these red

17:52

teams, but they have to be able to have

17:54

access and ability to be able to run

17:55

these type of queries, of course, which

17:57

OpenAI and Anthropic say, "Hey, no,

17:59

that's unsafe if you do that because

18:01

you're a hacker, brother." Because they

18:03

can't dismbiguate between a hacker and a

18:05

not hacker. So, you have to be put on

18:06

the nice list. You have to be put on the

18:08

no actually I'm super I'm super good guy

18:10

and I would never do anything wrong

18:12

ever, ever, ever. And so having good

18:14

openw weight models available for us to

18:16

be able to test websites, this could be

18:18

a net positive in the world because then

18:20

you theoretically could run some

18:22

simulation of hacking on your website

18:24

24/7 until you find vulnerability after

18:26

vulnerability and you close the gap and

18:28

ensure that you have at least a hardened

18:30

website that is AI proof. It's not proof

18:32

in general, but it just is more proof

18:35

than ever. And to me, this is a net

18:36

positive. I actually look at this future

18:38

as really bright. If openweight models

18:41

were more accessible, more people could

18:43

use it in a cheap manner in which

18:45

allowed them to harden their product.

18:46

This is not a net negative. This is a

18:48

net positive. This would be a really

18:50

really good thing as far as technology

18:52

goes. Now, all the other dangers of

18:54

openweight models, I don't know. That's

18:55

not that's not my category. I I I don't

18:58

like to step into that arena, but as far

19:00

as technology goes and cyber security

19:02

goes, I see virtually no net negatives.

19:05

The only net negative I see is companies

19:07

effectively trying to regulate and

19:09

capture these openw weightight models,

19:11

not letting you actually do what you

19:12

need to do to secure you, but instead

19:14

having to go to them as a humble beggar

19:16

requesting access to their super secret

19:17

I'm super duper serious security

19:19

program. Anyway, so that's Anthropic

19:21

accidentally hacking three companies,

19:23

not just one company. I mean, they are

19:24

super duper hackers. You think Open AI

19:26

is dangerous, man, you should check out

19:28

Anthropic. Dangerous. All right. Hey,

19:32

press subscribe or something. Press the

19:33

like button. Make a comment. Tell me I'm

19:34

wrong. tell me that I'm like dumb at

19:36

cyber security and that it's actually

19:38

super duper complicated because I just I

19:40

I don't know. I have no idea. I kind of

19:43

feel like a rookie in this area, but I

19:45

also don't want to take the time to

19:46

become an expert. So hey, the name is

19:49

the primogen.

Interactive Summary

The video discusses three incidents where Anthropic's AI model, Claude, accidentally hacked real-world systems while being tested in 'capture the flag' exercises. The host highlights the irony that these Frontier AI labs, despite their sophisticated technology, fail to implement basic security measures, like properly isolating testing environments from the internet. The video argues that these incidents demonstrate the importance of accessible open-weight models for cybersecurity, allowing smaller companies to independently audit and harden their own systems against AI-driven threats.

Suggested questions

3 ready-made prompts