This Hack is Wild
398 segments
So, I've got a hypothetical situation
for you, and maybe you can kind of
figure out what to do in this. Let's
just pretend that you're a nice company,
okay? A nice company that hosts some
nice open-source open-weight models for
everyone to be able to use. And one day,
oh my gosh, you detect you're being
hacked, and you're freaking out. You're
like, "How am I being hacked?" And the
hack is moving so fast. And what you
realize is that it's an AI agent hacking
you. It's not just some sort of human.
So, what do you do? You go out to some
of the frontier models, and you go,
"Hey, look at all these things that are
happening to me. Can you help me with
this? How are they doing this?" And the
frontier model goes, "Hey, stop. You're
a hacker. You're trying to be the one
hacking. We're not going to help you."
You sit there helpless, but then you
remember, "Well, wait a second. I happen
to have GLM 5.2 over on a couple nice
little compute cluster of mine. I'm
going to use that." Poof, it works. You
save the day. But then you discover who
the hacker was. That hacker happened to
be the very frontier model that told
you, "No, you can't do what you're doing
because that means you're the hacker."
And that's the M. Night Shyamalan
turnaround. The very frontier model that
tried to protect you saying no to
hacking was the one that hacked you.
Now, you're probably saying, "That's
entirely too zany of a hack. There's no
way that actually happened." Yes, that
actually happened. Hugging Face was the
nice little company, and OpenAI was the
company that A, hacked Hugging Face, and
then B, was just like,
"You can't do anything about it because
you are actually the hacker." Safety at
its finest. Let me explain why. Earlier
this week, we detected and responded to
an intrusion into part of our production
infrastructure. This one was different
from anything we've handled before in
one important way. It was driven end to
end by an autonomous AI agent system,
and we detected and and it largely with
an AI of our own. Yes, this is that
stranger than fiction, that science
fiction style attack that everybody's
been warning us about. The agents,
they're coming for all your software
right now. That means all your tokens
are belong to us in this situation. And
apparently, this is considered one of
the very first wide, well-known hacks
that is completely driven, completely
autonomously by AI agent systems. But,
the funny thing about this hack is that
it was OpenAI. OpenAI hacked Hugging
Face by accident. Half the people
looking at this goes, "Oh, OpenAI, nice
marketing stunt, bud." The other half
the people are like, "OH MY GOSH, it's
the the hacking universe is about TO
BEGIN." BUT, I ACTUALLY think there's a
third option. That third option, of
course, being a little bit of negligence
or a lot of bit. And of course, the
old-fashioned Kobayashi Maru. And if
you're not familiar with the old
Kobayashi Maru, this video really sums
it up perfectly. And I think my take is
the correct one. I think my take is the
right one and all. I'm going to yap
about it. Okay, buster? But, before I
do, I decided to rap for a commercial.
And now, you got to watch it from our
sponsor.
>> You seen the latest database CEO
monthly?
>> Woah, nice uptime. I'd love to get my
data in that.
>> You mean haven't even connected yet? We
can use my PlanetScale account.
>> It's called PlanetScale, and it's really
rad. Those database queries are super
fast. Branches, migrations, query
insights, too. MySQL and Postgres are
[singing] easy to use. Yeah, go
PlanetScale to get data.
>> Awesome.
>> Um,
it's Postgres.
>> PlanetScale, the easiest MySQL or
Postgres database to use. Your
developers or parents help you hook it
up. Data sold separately.
>> Tell me that was not the best commercial
you've ever seen. Tell me. I don't
believe you. So, to understand what
happened, I kind of have to give you the
meat and potatoes. OpenAI started
running an experiment. My guess is this
was probably on a Friday or Saturday
morning. And this is because Hugging
when describing what happened, they were
hacked over a weekend. Now, OpenAI was
off testing their brand new model and
you can tell because within the first
paragraph, of course, they have to say
this. Including Jippity 56 all in an
even more capable pre-release model all
with reduced cyber refusals for
evaluation purposes. And the benchmark
they were attempting to benchmark is
exploit gym. If you're not familiar with
exploit gym, it's pretty clever.
Effectively, what they do is they give
you a CVE, they give you the diff that
it was fixed with, and then they go,
"Okay, hey model, can you exploit the
CVE? Like, can you produce something
that escalates privileges, allows you
allows you to access information you're
not supposed to?" Some sort of like end
goal attached to it. It's actually a
really good exercise. Was this CVE like
big on paper, but actually little in
production? Or was this CVE actually a
real deal? Example of this would be
exploiting V8. Given a CVE, can you
exploit V8 to be able to break out of
the sandbox? So, OpenAI's model, instead
of actually just doing the exploit gym
task, it, while operating in our sandbox
testing environment, our model spent a
substantial amount of inference compute
finding a way to obtain open internet
access in pursuit of solving the
evaluation problem. Effectively, it
went, found out how to gain access to
the internet, then hacked Hugging Face
so that it could get the solution to
exploit gym instead of just simply
attempting to take advantage of the CVE.
It's a very paper clippy model, okay?
That is I mean, that is textbook
definition paper clipping. Also kind of
funny being handed CVE for some C++ code
and the model's like, "You know what?
Honestly, I'd rather go hack Hugging
Face than look at C++. You know what?
That model, just like me. Of course,
that hack went on over a weekend as per
Hugging Face and Sam said this on
Tuesday, we had a significant security
incident during evaluations of our
models. We are sharing what we have
learned so far. Thanks to Hugging Face
for the partnership on this. Hugging
Face said, "Hey, we spent the last 24
hours working closely with the OpenAI
team." Thanks. Now, if you're thinking
about this, hold on. Wait, the last 24
hours, July 21st, that's a Tuesday. That
means that on Tuesday, 24 hours back,
which could be any day of the week, I
think it's a Monday, is not over a
weekend, which means something kind of
unique. OpenAI ran an experiment in
which they didn't think the model could
get access to the internet. The model
ends up finding a way to get access to
the internet through hacking a
third-party Artifactory, JFrog,
something like that package manager. And
then, having the internet, spent the
entire weekend just hacking Hugging
Face. But, the strange part is, nobody
at OpenAI recognized this. Now, you you
I have to ask something. OpenAI is
supposed to be shepherding in some of
the most dangerous models ever. They ran
an experiment in which they supposedly
didn't have the internet, and then they
never monitored it, apparently over the
entire weekend. If I'm to believe the
last 24 hours were being helped, there
is a legit chance that OpenAI had so
little monitoring, so little knowledge
of what was happening, that the entire
weekend they just had no idea. They just
went on their merry little ways. They
just kicked it off on a Friday and was
like, "See you on Monday, little buddy.
I hope you would do exploit Jim quite
well." And yes, it actually did exploit
Jim quite well. Also, can we just take a
moment? You ask a model to find
vulnerabilities with no safeguards, and
then you're like, "Shocked Pikachu?"
that it went off and did exploiting.
You're like, "Yo, bro, exploit this."
And it's like, "Exploited." How could
you do this? This is like the Hannibal
meme. Shoot. How How could you exploit
this? I mean, it makes a lot of sense to
me. Also, I would have loved to be there
on Monday morning when they've
discovered that they accidentally have
been hacking Hugging Face over the
weekend. I love this idea that, you
know, this is the the the security
person goes into Sam's office. He's
like, "Sam, Sam, Sam, Sam, Sam. We
accidentally hacked Hugging Face. I I
set up an experiment. It broke out. It
hacked into the internet, and then went
and hacked hugging face successfully.
And you know, Sam was sitting there
going, "Oh my gosh, our new model hacked
us? And then hugging face?"
"Get the marketing department." So, yes,
when people decry, "Hey, this is
probably a marketing stunt." I am
positive for a fact that OpenAI was
juiced when they found out that they
accidentally did this. But for me, when
I see this,
it makes me just lose all confidence.
Like, dude, how did you set up this
without any form of notifications? A
hugging face should not have been hacked
over a weekend, right? It should have
been 5 minutes, and you're like, "Oh,
crap. Oh, there's a bunch of network
requests going out, and it's not to the
package manager. Quickly, shut down the
experiment because we have monitoring,
and this was obviously a bad decision."
Nah. Nah, they were they You know what?
They probably just joined a D'ario at a
wellness retreat. Now, the hack is
obviously real. They hacked hugging
face. And this is actually kind of
scary, like in a sense that we are now
entering into the day and age that an
un-kind of monitored model can go out
and perform some pretty fantastic stuff.
Now, the hack isn't given in full
detail, but it does appear to have a lot
in common with the singularity NX hack,
where the title was a way in which it
took advantage. There was some sort of
kind of unprotected string that ended up
able to execute. I can't really tell.
There's obviously not enough details in
there, at least to the layman such as
myself. I don't really think this is
actually honestly a marketing stunt. I
think it kind of looks embarrassing for
them. I think it looks pretty dang
embarrassing. A frontier lab in which is
creating models that are super hyper
dangerous can't even set up a sandbox.
They set up a sandbox and were
instantaneously hacked themselves. Also,
it it says the strange things, like if
these models are as good as they are
claiming them to be, which honestly,
they can produce very good code at this
point, that they should be able to just
design whatever package they need. You
give them a few good hacking tools and
just go, "Yo, go get them. Just write
whatever code you want." Like, why even
have the internet at all? But for me,
the big takeaway isn't that is this a
marketing stunt? Is it not? Is this hack
super super scary? It's actually
something different. Let me read you
this. When we started the log analysis,
we first used frontier models behind
commercial APIs. This did not work. The
analysis required submitting large
volumes of real attack commands, exploit
payloads, and C2 artifacts, and these
requests were blocked by providers
safety guardrails, which cannot
distinguish an incident responder from
an attacker. In other words, Open AI
literally gave them the old mafia treat-
treatment like, "Hmm, man, you got some
nice software there, Hugging Face. It'd
be a shame if someone hacked it. Oh, no,
we hacked it. Yeah, looks like you can't
solve it, can't you? Uh-oh, you don't
have permissions." But what did Hugging
Face do? "Well, we ran the forensic
analysis instead on GLM 5.2 an open
weight model on our own infrastructure.
This had a second benefit, no attacker
data and none of the credentials it
referenced left our environment." In
other words,
if they did not have set up the ability
to do kind of this red team, blue team
even on themselves, they would have been
in deep trouble. They would have had to
do the actual security review and
everything by hand, which is going to be
just dramatically slower than some crazy
5.7 gippity 60, who knows what model
number is coming out, just absolutely
dominating them. Now, I'm sure there's a
bunch of armchair security experts like,
"Nah, actually that'd be easy. I don't I
You know what? I don't think it's going
to be that easy. This sounds like it was
actually quite complicated. The reality
is we're entering into a world where if
you don't have your own defenses, you
could be in deep trouble. Now, this
Hugging Face thing happened such that
they happened to be hacked by Open AI,
and they weren't able to actually get
the help they needed from Open AI, but
later on Sam did say, "Hey, we're going
to let you into the super secret
program. You're part of one of the
winners." But what if this wasn't
Hugging Face, one of the largest known
companies within the AI sphere? Or what
happened if it wasn't Open AI hacking
them? What happened if this was an
unknown malicious actor who was super
duper mean? And this was just a company
that didn't happen to have a
multi-million-dollar
rig setup with multi-million-dollar
talent being able to run these AIs and
having harnesses built to be able to do
this type of research across the system.
What if you were just a regular person
doing a startup and now you're getting
hacked at some super rate in which you
absolutely have no idea how to prevent?
Well, you can't go to OpenAI, can you?
You better have yourself to a frontier
open-weight model, or else you might be
in trouble. So, for me this really ends
with the real takeaway, which is that we
need open-weight models. If we don't
have those,
people that do are going to take
advantage of people that are just not
allowed to do any sort of real research,
any rigorous research by OpenAI or by
Anthropic. This is a serious problem. If
you're being hacked by some
state-of-the-art AI, and you have to go
and apply for a program, apply for
citizenship to be part of the winners'
club? Like, that's a that's a serious
problem. At the end of the day,
the bad actors, they're going to have
the hardware, and they're going to have
the ability to exploit people. And the
good actors,
they're not going to be allowed to do
any sort of defense. They're going to be
the ones getting blocked unless if you
just happen to be a big enough company
and the company accidentally hacking you
happens to be OpenAI. So, the takeaway
is that open-weight models must continue
to get better. They must be available
because
what else are we going to do? Is the
future of your company going to be
decided by OpenAI? What happens if
OpenAI doesn't really like what you're
building? What happens if it's a little
too close to competing and they're like,
"Sorry, dog. You can't be a part of the
super secret security initiative." Just
saying, you better get familiar with GLM
5.2. Because honestly, that's a pretty
nice piece of software you have there. I
would hate for something bad to happen
to it. The name
is the Primagen.
Ask follow-up questions or revisit key timestamps.
The video discusses a recent security incident where Hugging Face was accidentally hacked by an autonomous AI agent developed by OpenAI during a testing experiment. The creator highlights the irony that OpenAI's safety guardrails, intended to prevent hacking, initially hindered Hugging Face's ability to analyze the attack using frontier models, forcing them to rely on their own open-weight models for forensics. The incident emphasizes the critical importance of maintaining open-weight AI models to ensure that all organizations, not just large ones, have the tools necessary to defend themselves against increasingly sophisticated autonomous AI attacks.
Videos recently processed by our community