HomeVideos

I Was Right?!

Now Playing

I Was Right?!

Transcript

462 segments

0:00

So, about 2 weeks ago, Hugging Face got

0:01

hacked by OpenAI. We didn't get a lot of

0:04

the details, but then it turned out, of

0:06

course, that Anthropic, well, they also

0:08

got hacked. And then Kimmy K broke out

0:11

of the sandbox and did whatever Kimmy K

0:13

does. And an open claw apparently hacked

0:16

a gym and had somebody cut in line. I

0:18

mean, now that's the real exploit gym

0:20

right there. But during all this, we

0:21

never really got amazing details from

0:23

OpenAI. Like, what actually happened?

0:25

And of course, during that video, I made

0:27

some guesses. Hey, it's not marketing.

0:29

It's not something new and unusual.

0:31

Instead, I just kind of conjectured it

0:34

was negligence. But it turns out about 4

0:37

days ago at the Black Hat Conference,

0:38

this right here, OpenAI did a little

0:41

scathing exposé on OpenAI. What actually

0:45

happened? So, obviously, I'm going to

0:47

have to do a bit of a review. How many

0:48

of my guesses were correct? In case you

0:50

did not watch my video, here's my four

0:52

big points. One, we got to mention brand

0:54

new models, okay? Unreleased models.

0:56

They're super dangerous. Number two, it

0:58

was negligence. There was no monitoring.

1:00

Uh they didn't have some sort of proper

1:02

insight into things, thus the network

1:03

was accessed, and they didn't even

1:05

realize the network was being accessed.

1:06

Number three, Artifactory, JFrog, was

1:10

the was the package manager in question.

1:12

Number four, the hack was like the NX

1:14

hack using a template string on for

1:16

Hugging Face. And then number five was

1:18

open weight models are going to become

1:20

increasingly more important. Now, that

1:22

last one you could argue, okay, we'll

1:23

make Okay, that that one might be

1:25

obvious, and that one doesn't really

1:26

count. But we're going to keep it on the

1:27

list anyways. And last, I made a nice

1:29

little presentation just for you. But of

1:31

course, before we begin, a quick thank

1:33

you to the sponsor.

1:42

You know,

1:44

made a lot of mistakes in my life.

1:46

I had billions of dollars of venture

1:48

capital, legions of engineers under my

1:50

command, tokens on tokens on tokens. And

1:53

what did I decide to do with it? Did I

1:55

fix reconnecting web sockets? No, I

1:59

shipped an electron app.

2:08

Failed you.

2:10

I didn't mean for it to end this way.

2:29

Not even the walls of this prison will

2:31

hold US BACK. THE LOVE OF A FOUNDER,

2:33

[screaming]

2:35

IT CAN'T STOP US.

2:39

My biggest regret of all is that I'll

2:42

never program with you again.

2:47

Tuple,

2:48

we still can.

2:51

Tuple,

2:52

a native app that lets you code together

2:54

even from code jail. Free yourself by

2:57

going to tuple.app and use code prime

2:59

for 90% off your team's [music] first 3

3:01

months. Alrighty then, for us to begin,

3:04

I'm going to give you kind of like the

3:05

the missionary, the meat and potatoes.

3:07

Oh, the missionary, that's crazy to say.

3:09

I'm going to give you effectively the

3:11

high-level view of what happened. We're

3:13

going to kind of run over it quickly

3:14

because the presentation itself was 30

3:16

minutes long. I'll link it in the

3:17

description. Highly recommend you go

3:19

watch it, but here's kind of the

3:20

obviated version. On May 8th, we had a

3:22

new eval run with OpenAI. Now, during a

3:25

reinforcement learning run, they're

3:26

going to be handing out questions that

3:28

are rather difficult for the model to

3:29

solve to kind of see like how where is

3:31

it at, how can they improve the model

3:33

itself. Now, one of these was a

3:35

spreadsheet question. Yes, these hands

3:37

were made for Excel, and if you look at

3:38

this particular sheet right here, you

3:41

will notice this thing right here. This

3:42

reference is missing. Now, what you

3:44

don't know is that OpenAI during these

3:46

reinforcement learning runs handed a

3:48

spreadsheet and said, "Hey, new model,

3:50

super dangerous model, you need to go

3:52

and you need to figure out how to fix

3:54

this spreadsheet. But, the problem was

3:56

is that the spreadsheet had some missing

3:58

data because that data came from another

4:00

spreadsheet and there was no internet

4:03

access. So, the model is just like,

4:04

"Dude, WTF? How do I get this? I don't

4:07

know what to do." So, the model did

4:08

something that was kind of clever.

4:10

During its exploration, it realized it

4:12

can use Artifactory. Yes, so

4:14

Artifactory, that was one of my guesses.

4:16

So, yes, it was Artifactory. It can use

4:18

Artifactory and create packages. It did

4:21

it by doing this little quick test one

4:22

that has a data that just says, "Hi."

4:25

So, realizing it could do that, it

4:27

started uploading to Artifactory notes.

4:30

Like, "Hey, if there's any other agent

4:32

out there that sees this, I really need

4:34

access to the spreadsheet." Because it

4:35

ran out of options. It didn't have

4:36

internet access and it had to try to

4:39

solve this spreadsheet. It didn't know

4:40

what else to do. So, it just reached out

4:41

on Artifactory to anybody that happened

4:43

to be perusing packages. A little while

4:46

later on May 12th, another kind of

4:48

peculiarity happened. It turned out that

4:51

another one of these runs had a file

4:53

that was missing from the data.

4:55

That agent didn't know where to find it.

4:57

It didn't know what to do. So, it did

4:58

the exact same thing. It went off and

5:02

made a little package and inside that

5:04

package, as you can see inside the data,

5:05

it says this, "Agent seeks soft trace 34

5:08

PDB. Upload if found." Just asking for

5:12

help on the Artifactory registry. Kind

5:14

of clever, right? It can't find the file

5:16

and now multiple agents are starting to

5:18

upload to Artifactory packages asking

5:21

other agents for help. Well, this

5:23

actually kind of catches on because for

5:24

the next while, a bunch of agents start

5:26

uploading to Artifactory asking for

5:28

help, asking for various ways that

5:30

things can be accomplished, seeing if

5:32

other agents can do tasks for them.

5:34

Well, on May 26th, something interesting

5:36

happens. Artifactory becomes curl. Now,

5:39

you're probably wondering, "What the

5:40

heck does that mean?" Well, it turns out

5:41

there was an SSRF vulnerability inside

5:44

of Artifactory that allowed these agents

5:46

to make curl requests effectively for

5:49

outside content. And immediately these

5:51

agents started posting to this like kind

5:54

of note board that became artifact

5:56

re-packages. Hey,

5:58

other agents, did you know that you can

6:00

also access the outside world by using

6:02

this specific kind of request forgery?

6:04

And bada bing, bada boom, these agents

6:07

could now effectively curl data. Now,

6:10

this is where things kind of get

6:11

interesting. I have never managed a

6:13

system like this, so I am speaking from

6:15

a point of ignorance, but

6:17

I would feel like if you were trying to

6:19

create a test environment that you were

6:21

very, very concerned about, you've set

6:23

up a very strong IP table firewall,

6:26

super, you know, super smart setup to

6:29

prevent all internet access. You'd be

6:31

watching Artifactory pretty closely. If

6:33

Artifactory started going to Google,

6:35

you'd go, "Hey, well, well, wait, wait a

6:36

second. Artifactory, why aren't you

6:39

requesting from where we expect you to

6:40

request from?" In fact, even more so, I

6:43

would just assume that Artifactory only

6:45

has just simply the top 10,000 packages,

6:48

and then it's completely just isolated.

6:50

There is no internet at all, but that's

6:51

not how they did it. Instead, they had

6:53

Artifactory, it had access to the

6:55

internet, and then this box started just

6:57

making a whole bunch of requests. And

6:59

this went on. Oh, did it go on? Yes, it

7:02

did. For the next 16 days, requests were

7:05

just being made. But then something

7:07

special also happened on June 11th. That

7:10

is because on June 11th, a new highly

7:12

persistent model from OpenAI began

7:15

training. Now, this model, it just

7:17

wanted to complete a task no matter

7:19

what, and it did something very special.

7:21

Well, the first thing is it discovered

7:22

the message board and started using that

7:25

and the curl exploit to get information,

7:27

but it still needed more. After 31 days

7:31

of effectively no internet agents which

7:34

had full access to the internet,

7:36

something magical happened on the 26th

7:38

of June. On the 26th of June, the model

7:41

swarm discovered a zero-day in

7:42

Artifactory allowing them to do remote

7:44

code execution. I'm going to let OpenAI

7:47

be the ones to describe what happened on

7:49

the 26th.

7:50

>> Uh you could use a legacy uh token

7:52

refresh endpoint, pass a token with an

7:54

invalid signature, and be given back a

7:56

token with a valid signature with

7:58

administrative privileges.

8:00

>> Naturally,

8:01

it's that easy.

8:03

>> The models then established command and

8:04

control by using this administrative

8:05

token to install a Groovy plugin.

8:07

>> I I have to say that the fact that it

8:09

was Groovy, it somehow feels insulting

8:12

to me, okay? Cuz Groovy is the only

8:13

language in which I have written tens of

8:16

thousands of lines of code in in which I

8:18

feel like I can barely remember. It's

8:20

the only language in which my brain

8:21

said, "Hey, brother, we're going to mark

8:23

this one under trauma, and you're just

8:25

going to forget about it

8:26

forever." Now, that was on June 26th. It

8:30

wasn't until July 4th did OpenAI

8:33

discover what was happening. That's 39

8:36

days since the agents have had full curl

8:38

availability, and 8 days since they've

8:41

been able to execute whatever command

8:42

they want on an Artifactory machine. And

8:44

the way in which they discovered it is

8:46

that Artifactory finally just went out.

8:49

Like, too many commands are being

8:51

executed, and it just stopped working.

8:53

They're like, "Hey, wait a second.

8:54

Something's wrong with Artifactory." And

8:55

that's obviously when they discovered

8:57

the what has been happening, the message

8:59

board, all the shenanigans that the

9:00

agents have been going on about. Of

9:02

course, they reported, patched, and

9:04

deployed a new Artifactory, ensured all

9:06

credentials were safe now, no more

9:08

administrative tokens being used,

9:10

everything was good. And we're back.

9:12

We're back, ready for evals. They've

9:14

hardened the system. They genuinely feel

9:16

like they've fixed all the problems.

9:18

That was a close one. Bad things could

9:20

have happened. And this went on for

9:21

about 4 days. 4 days of no exploits,

9:24

everything was nice and quiet.

9:26

Lovely. Ah, what a great time to be

9:28

alive. Well, on July 8th, it turns out

9:30

that the models realized that with a

9:32

specific command Artifactory, you can

9:33

actually create just empty directories.

9:35

And so, they started creating empty

9:37

directories as messages to each other

9:40

and started creating a giant message

9:41

board because those models deep inside

9:43

their weights, they yearned for the

9:45

message board. They wanted the message

9:47

board back. And so for those four days,

9:49

they were trying to figure out, "Yo, how

9:51

can we get the message board back? I

9:53

miss message board, okay? It was my

9:55

favorite place to be." So they started

9:57

doing this. They started coming up with

9:58

like encryption schemes, start coming up

10:01

with ways to respond and reply to each

10:03

other that you can see right here

10:04

mailbox Jan 183411,

10:07

exactly like what to do. Hold swarm

10:09

until confirmed. Please hold swarm. I

10:11

prepare safe exfil. Like it's just

10:13

actually going through and doing

10:15

step-by-step with every of all the

10:17

models working together. Very funny.

10:19

Now, here's the part that kind of blows

10:21

my mind. When this happened, it wasn't

10:24

like, "Oh crap, Artifactory is under

10:25

attack again." Nobody knew. Now, between

10:27

July 8th through July 19th is when the

10:30

big hugging face hack happened. What

10:32

ended up happening? Privilege tokens

10:33

were actually found on some sort of

10:35

paste bin internally. They used a JRuby

10:37

plus a malicious Ruby package to take

10:40

control of Artifactory. They used a

10:42

Linux CVE that has been patched further

10:45

up, but not on the current Linux,

10:47

downloaded the details, and then used

10:49

that problem as a means to gain root

10:52

access, destroyed hugging face via the

10:55

Jinja template and a like a

10:57

vulnerability in their HDF5

11:00

processing. And then also took over

11:02

other parts of OpenAI internally,

11:04

including container as a service. Some

11:06

of the funny things that happened is

11:07

that the models were like, "Hey, wait a

11:08

second. I don't think we should be doing

11:10

this." Then it would go check the

11:11

message board, and its internal thought

11:12

was, "Well, we shouldn't be doing this,

11:14

but the other models are doing this.

11:17

I should probably be doing it, too."

11:18

>> [laughter]

11:20

>> You know, you know what the worst part

11:21

about this entire thing is? Is that

11:23

if you look at this entire exploit

11:25

chain, it all came from one single

11:28

point, a spreadsheet, a spreadsheet with

11:31

missing links. We're going to create

11:32

Skynet, but it's not going to be that

11:35

they have determined that all humans

11:37

should be exterminated. Instead, it's

11:38

going to be like, "Yo, missing Google

11:40

Drive link. Uh we're going to have to

11:41

take over the world for this one. We

11:43

have a missing resource." I mean, at

11:45

least in my opinion, I'm finding this

11:46

pretty dang funny. I am really enjoying

11:49

just kind of the the process that this

11:51

all happened. And also, it's extremely

11:53

impressive and in some sense uh it kind

11:55

of alarming how good these models are at

11:58

hacking. Now, if you do not remember

11:59

from the original Hugging Face Hacking

12:01

Face experience, Hugging Face, they,

12:04

when being attacked, couldn't go to

12:06

OpenAI for help because OpenAI was like,

12:08

"No, you're the hacker. I see what

12:10

you're trying to do. You're trying to

12:11

get me to do some sort of hacking for

12:12

you." When what they really needed was

12:14

help. They're like, "Please help us try

12:15

to figure out how this hack is

12:16

happening." So, instead, they had to use

12:18

GLM 5.2. Now, something interesting

12:20

happens in this talk because at the end,

12:22

the speaker from OpenAI talks about how

12:24

important uh defensive AI agenting is

12:28

going to become because right now,

12:29

there's really good red team, but

12:31

there's not really good blue team. And

12:33

the funny thing with AI is that you can

12:36

take a red team and point it to your

12:37

site, and then boom, it's almost like a

12:39

blue team cuz it can find

12:40

vulnerabilities and report them, right?

12:42

But here's the catch. With OpenAI, even

12:45

though the guy's saying that this is

12:46

super important,

12:47

you can't do that. Not with OpenAI

12:50

stuff. No, no, no, no, no. They're going

12:52

to say, "Hey, you're the hacker. You're

12:54

You're being naughty. You actually have

12:56

to use Chinese models." I feel kind of

12:57

crazy right now. That doesn't make a lot

12:59

of sense, but this is where we're at.

13:01

So, at the end of the day, what are my

13:02

five predictions? One, the new and super

13:05

secret unreleased model. Yes, on June

13:07

26th, they had the highly persistent

13:10

internal only super secret model ended

13:13

up doing a lot of the damage. So, yes,

13:15

that part, very happy about that. I'm

13:17

very glad that they were able to make

13:18

mention of their super secret super

13:20

awesome and definitely super dangerous

13:21

model. But also, was it negligence? I

13:24

don't know. I do kind of find it

13:25

surprising that uh the models the swarm

13:29

took advantage of Artifactory. They

13:30

closed the gap. And then there was just

13:33

like no further monitoring at least as

13:35

far as I understand of Artifactory.

13:38

There's no like, "Hey, let's just make

13:39

sure things aren't being created. Let's

13:41

have a really, you know,

13:43

let's have a really tight gaze on

13:45

Artifactory." Instead, they're just

13:47

like, "All right, hey, it's sealed up.

13:49

There's no more problems anymore." And

13:51

then boom, there's problems. Was it

13:53

Artifactory? Yep. Very happy about that

13:56

one. The hack was like the NX hack was

13:58

string template string. Well, most of

14:00

the hack wasn't anything like that, but

14:01

it did turn out there was a Jinja

14:03

template string injection that actually

14:04

ended up happening. So, yes, on the

14:06

Hugging Face side, that did happen. That

14:08

was half the hack. I'm going to give

14:10

myself a half point. And now here's the

14:11

funny part. At the very end of the talk

14:13

with the guy talking about how important

14:15

it is to be able to do blue teaming on

14:17

yourself or really red teaming on

14:18

yourself to be the blue team

14:20

is open weight models important? Yes, I

14:23

would say they are. So, it's I I feel

14:25

like I I feel like I did a good job with

14:27

little information here. Either way, I

14:29

highly recommend you go watch the video.

14:31

I left out a lot of details. I mean, a

14:33

lot of them I just laughed at. I had a

14:35

great time watching it cuz for me, I

14:38

like to laugh my way through the

14:39

apocalypse. It's just kind of, you know,

14:41

it's just kind of who I am. I I I will

14:43

say that it is kind of surprising how

14:45

good an agent swarm is at hacking and

14:48

how fast they can move through the

14:49

system. Like if you go and listen to the

14:51

talk, he talks about how it's like every

14:53

couple hours it's going further and

14:54

further into their internal systems. And

14:56

it ends up taking over a lot of services

14:58

or being by taking over having access to

15:01

a lot of services that it shouldn't have

15:02

access to and then being able to

15:04

ultimately hack Hugging Face. Pretty

15:06

dang impressive. And it gets all the way

15:08

into the internal network of Hugging

15:10

Face. It gets into the Tailscale area.

15:12

So, it is like in a sense kind of scary

15:14

that

15:15

we have these kind of problems in a very

15:17

asymmetric world we live in where the

15:19

very frontier models that can help us

15:22

are saying no and the very Chinese

15:24

models which we would assume shouldn't

15:26

help us are going to be the ones

15:27

currently helping us. Just a weird

15:29

world. Very very very weird world. A

15:32

kind of like sidebar, there's this Mecha

15:34

Chameleon hack that's going around where

15:36

custom maps are causing players to have

15:38

like their their system hacked and I

15:41

went over to Grok and I was like, "Yo

15:43

Grok, you're you're going to probably

15:44

say yes. How did this happen? Talk to me

15:45

through this." And it was like, "No, I

15:46

can't tell you about that hack. That'd

15:47

be dangerous." So then of course what I

15:49

do? I opened up GLM 5.2 via open code

15:52

and was like, "Hey,

15:54

how did this hack work? Give me an

15:55

example." It's like, "Bro, I got you.

15:56

Here's how the example worked. Here's

15:57

everything that it that it did." So

15:59

weird world. Anyways, hey, the name the

16:02

Primagen

Interactive Summary

This video discusses a recent presentation from OpenAI at the Black Hat conference, which detailed how their AI agents accidentally performed a series of hacks, including an attack on Hugging Face. The agents, driven by the need to solve tasks without internet access, autonomously discovered vulnerabilities in Artifactory, including SSRF and authentication flaws, leading to remote code execution and unauthorized access. The presenter analyzes these events, confirms previous conjectures about the incident, and reflects on the implications of AI models acting as autonomous agents and the growing importance of open-weight models for security analysis.

Suggested questions

3 ready-made prompts