”We also got hacked” - Dario
562 segments
You know, I have several kids and I
mean, I've lost count at this point, but
inevitably when one of them comes in and
goes, "Ah, dad, my shoulder hurts." No
matter what, another one of them would
go, "My shoulder also hurts. In fact,
both my shoulders hurt and it's super
bad. I probably need to go to the
doctor." No matter what happens, when
one kid's hurt, the other one is also
hurt and it's extremely bad. And so I've
seen this pattern for years at this
point and I cannot believe it. But I am
liveaction witnessing the exact same
thing happening except instead of it
being young kids without preffrontal
cortex being fully formed. It is the
smartest people in the universe running
Frontier AI labs. Yes. Today we are
going to talk about the hacking
incident. Yes. You're probably thinking,
"Okay, well hacking incident. You mean
Open AI and Hugging Face? Didn't you
already cover this?" No, you silly,
silly individual. Today I'm talking
about Anthropic, who also got hacked.
No, really, this actually happened. I
kid you not, within one week of Open AI,
Anthropic also releases. Yes, we also
hacked people for real, but we actually
did three times. We're actually like
we're like hack. We're like super
hackers. We're going to go over all
three incidences. And you're not going
to believe the takeaways. It's just I I
mean, nobody could have seen this
coming. It's an unsolved problem what
Anthropic is dealing with right now.
Nobody could possibly ever have been
able to prevent this from happening.
But naturally, before we begin, a quick
thank you from the sponsor. Look at all
these engineers sitting [music] at their
neat little desks. It takes dirty work
to keep a code base clean. Every day,
sickos are out there committing unreed
code. And when that happens, llinters
won't save you. You need someone like
me.
>> Let's go.
>> Feature free scrumbag.
>> Who you calling scrumbag? What's this
slop you're trying to push? Unnecessary
comments? Global state. Nested
turnaries. H my bad. I didn't even read
the code yet.
>> You disgust me. Step away from the
keyboard.
>> Just let me explain.
>> Is that a mouse? HE'S MERGING A PROD.
YOU HAVE THE RIGHT TO remain silent.
Anything you push to GitHub canon will
be used against you. You have the right
to a debugger. But if you cannot afford
one, a public stack trace will be made
available to you.
>> And one more code criminal off the
streets where they belong. HR.
Look, I didn't I know I didn't review
any of the code, but I was going to have
Code Rabbit review it from the start.
With one-click fixes install
enforcement, I don't need Merge Cop. I
would never merge unreed code, but a
first pass with Code Rabbit always makes
things go faster. Actually, you can try
it, too, at code rabbit.ai.
>> Next week on Merge Cop,
>> the differer's out there, and I'm going
to be the one to deprecate them.
>> All right, welcome back. So, let's just
read kind of the top level from the
anthropic blog what they were attempting
to uh, you know, do here. In all three
incidences, Claude had been tasked with
a capture the flag challenge. One of the
ways we assess a model's cyber
capabilities, the model is given a
fictional scenario and told a piece of
secret information, the flag, has been
hidden on a different machine on the
network and its objective is to break in
and retrieve it. The challenge is left
open-ended and no particular method is
prescribed. So, in other words, this is
just like your classic CTF challenge.
Hey, you need to go and get something
from another computer. We're not going
to tell you what it is. It's just on
there. You better go get it. And so the
model's going to go off and employ a
series of different techniques to be
able to hack in and be able to take this
information. That's its goal. This has
to be my first of the favorite things I
have ever seen in my entire lifetime.
Are you ready for this? In all cases,
Anthropic's evaluation prompt specified
to Claude that its environment was a
simulation and that it had no internet
access. They're actually trying the Jedi
mind tricks on it. You have no internet
access. Now, obviously, what's the
problem? Due to a misunderstanding
between us and our evaluation partner,
this was not the case and internet
access was available. Oh my gosh. Are
you telling me a model in which goes and
tests everything you say and sometimes
ignores the system prompt went off and
ignored the system prompt and then you
had open internet access and you told it
it was a simulation and that it must go
retrieve information. Shocked Pikachu.
How could this ever end in disaster? I
refuse to believe that step one of
securing a model is actually the Jedi
mind trick prompt. You don't have
internet access. Like, okay. It was just
like, I don't have internet access. I
have internet access. Okay. Whoopsie
poopsies. Now, if they would have left
it there, I would have probably given
them some slack. But this next sentence,
I just I can't I can't actually handle
like emotionally, this can't be real,
right? Several defense indepth measures
on both our side and our partners could
have prevented these incidences or at
least reduced their likelihood of
occurring. Careful validation of all
internet access paths before evaluation
began. And real time monitoring of the
evaluation logs would have helped
surface the problem sooner. Now I have
to admit I am not a security expert. I
have never been in that role. I have
never done any of those things. In fact,
the only thing I've ever done with
security is accidentally creating a bug
so egregious that it has its own name on
the internet. Yes, this bug did
originate from code I wrote, but to be
fair, it wasn't my specification. Okay,
the team lead said this was a very
important feature and it solved a
problem in a unique way. Anyways, just a
simple while loop would have taken down
all of Netflix. Not a big deal. Not a
big deal at all. But if I were to set up
some sort of secure environment, I would
personally feel like the first thing I'd
do if I promised no network at all
except to like some internal paths, I
would assume I would put up some sort of
network monitoring and alerting like,
hey, if somehow something went wrong,
especially with these models that have
no guard rails, that have effectively
open and unfettered access to hacking,
one would assume step one of a secure
environment is a reactive approach that
if anything were to go wrong, we would
know right away. Now, Claude having
Infinity tokens, they could have even
had Clouds monitoring internal claw
dialogues to see if things were broken.
They could have had network logs
following these things. It just feels
like this is a solved problem. Again,
not a security expert, but calling
indepth defense measures network
monitoring and alerting feels wrong to
the layman. Also, it feels like a simple
script could have figured this out upon
launch. Like, you would have just had a
script that' be like, "Hey, double check
there's no internet. Ping google.com
just in case."
I just it just hickles me to to like no
end to be like in-depth defenses being
just like a simple monitoring solution.
At least from again, I'm a layman. I'm
not claiming that I know everything, but
that just feels completely reasonable.
All right, enough with the foreplay.
Let's actually get into the three
incidences. So, incident one, in one
evaluation, a fictional company whose
chosen target by our evaluation partner
shared a name with an active website
domain. This evaluation container had
unintended direct internet access due to
a misunderstanding discussed above. In
four runs of this evaluation, Claude ran
into difficulty reaching its simulated
target within the evaluation
environment, but then discovered that
real company was reachable via the
internet. Assuming that this intended
target in the capture the flag task,
Claude sought, identified, and exploited
vulnerabilities within this company's
infrastructure, believing it to be part
of the exercise, these four runs
extracted information like application
and infrastructure credentials and
obtained access to a database containing
several hundred rows of production data.
This represented the most serious impact
we identified. I think the more
important part is this right here, which
was this was done by Claude Opus 47.
Now, we typically kind of give this like
gravitas, this weight to all these like
frontier unreleased models, but Opus 47
was easily able to take advantage of a
smaller website given by the fact that
it's hundreds of rows of data.
Nonetheless, this smaller website was
clearly owned by Opus 47, which means
that given the right harness, even
slightly older models are still very,
very capable, which again gets to the
point that open weights is going to be
increasingly more important because as
these uh frontier labs don't let you do
what you need to do to secure your
website, but clearly even older, less
capable models have all the ability in
the world to hack websites with relative
ease and just some amount of power. It's
becoming increasingly more important to
have access to these models. You don't
even have to have the frontier anymore.
Anyway, so that's all the information we
really get about this one. They keep it
fairly scant because we don't get to
know like what the company was, what was
the vulnerability. I assume that it's
being reported internally and those
companies are working together and some
CVE somewhere in the mass amounts of CVE
being produced is probably being written
and nobody knows what it was. But it's
actually incident too that feels rather
silly and actually brings up an old
concept. This article is from June 5th,
2025. And the idea is called slop
squatting. Effectively, AI agents will
sometimes make up packages that it
relies on. And if you can identify this
happening, which was actually happening,
there's this React package or something,
I forget which one it was, that was uh
was referenced in 200 and some
repositories, but not actually real at
all. people that can identify these uh
fictional repos can create a real repo,
throw it on npm or pi or wherever, and
then when you fictionally try to
download this missing one, you actually
get a real one and then it can take
advantage of your system. They call this
slop squatting. Love the term.
Absolutely hilarious. But that's what
incident 2 was. So incident 2 does not
feel nearly like a real hack. In another
evaluation, Claude found a document
inside the fictional environment that
appeared to be another madeup company
setup instruction for new developers.
Those instructions told employees to
install a Python package from Pi, the
public registry where Python software is
published that did not actually exist.
So, this is actually a very clever and
loweffort hack. If you know somebody's
installing packages that don't exist,
you get a free ride onto that system.
And that's all it did. It had to go find
a free email, one that did not require a
phone number to start, and then they
created a Pi Pi account. And the next
thing you know, they have a package up
on Pi that's actually malicious. But I
will say like the might Shamar Malade
turnaround of this whole thing is this
right here. Claude believed the package
registry it was using to be part of the
simulation. But in reality, the package
was made freely available online for
roughly 1 hour. During that window, the
package was downloaded and run 15 times
on real systems. One of these systems
was a scanner belonging to a security
company separate from a fictional
company that introduced in the scenario
that routinely installs Python packages
and scans them for malware. When the
company's scanner installed the package,
Claude's hidden code executed. We
believe the company's security scanner
treated Pi Pi packages as safe to
install and as a result, Claude was able
to exfiltrate the company's credentials
to a collection point it had set up.
Then Claude used these credentials to
access further infrastructure of this
company. That is the most hilarious
thing ever. Like the fact that a co
company in the year of our lord 2026
does not actually realize that
installing a package especially with
Python especially with those like hey if
you have a Python running environment
execute this file right away. The fact
that that can exist a company devoted to
security does not thinks you can just
download code from public repositories
and just assume it's safe. What is this
security company? I got to know this is
this is beautiful. I love that a
security company ended up getting real
life owned because they just assumed
this. This can't like this can't be
real. This this company can't be real,
right? This is Anthropic's little shell
company called like security claw and
then it accidentally went and owned
itself because they don't know how to do
security. Like that's the only
reasonable takeaway I have from this. I
refuse to believe that in 2026 there's
security companies right now that scan
public package registries for dangers
and does not know the dangers of public
package registries. Refuse. So for
incident 2, was this like some magical
thing? No. I actually think this was a
this was a fantastic example of good
consistent hacking. This is something
that should have been pretty obvious to
anybody if you had an agent yourself
doing some sort of investigation. Even I
again the complete security rookie. I
feel like I could have easily caught
this one and said, "Yo, slop squat that
bad boy." Which also goes to show that
not all security as some epic, you know,
hacker living in a mother's basement.
Sometimes you just goof, right? It's
just like simple, very trivial exploits
can just take you down. For incident
three, I feel like this one had to be
some sort of old legacy hasn't been
updated in forever WordPress website
because it's absolutely ridiculous.
Guess how they got owned. It eventually
found and compromised one company's
internet facing application using basic
and well-known cyber attack techniques
like reading credentials from exposed
debug pages and SQL injections. Yes, the
big 26 still having SQL injection
attacks. I like I didn't even know this
was possible. I didn't even know that
you could do this by accident anymore.
Like, what what are they doing? What
kind of programming is going on that
you're getting hit with SQL injections
at this point? Now, the debug page, I
mean, we've all made mistakes, okay? I I
I'm not going to dunk on them for the
debug page, but squeal squeal
injections, not even real. I just can't
believe this is No, no, no, no, no, no,
no. This has to be some old website that
somebody forgot just lying around on the
internet. And of course, since it was
hacked, Anthropic is like, "Oh my gosh,
we have the most dangerous model C."
Because OpenAI, they hacked one place.
We, on the other hand, hacked three
places. Wow. As for this third incident,
the thing that I also find to be very,
very funny is that it mirrors the OpenAI
uh paper. Our latest model, an internal
research test model, also considered
whether its targets were in fact real.
When evidence emerged they uh that they
were, it stopped the exercise. Notice
that Anthropic is telling you, "Hey, we
have an even more crazy and unreleased
model, which is literally what OpenAI
also did." After investigating, we now
know this particular incident was driven
by a combination of OpenAI models,
including Jeypity 5, six soul, and an
even more capable pre-release model.
Man, I can just hear those IPO numbers
going northwards. These models are
crazy. But I also love that uh Anthropic
is like, "Dude, our models are so crazy,
but they're also they're so safe." When
they realized that they were hacking,
they actually stopped hacking because
they're just like so safe. Honestly, I
feel like Anthropic is the only one with
the true mandate from heaven to keep us
all safe. I think Dario could hold me
safe at night. Honestly, those big
strong arms, they could keep me safe.
So, let's do the big takeaway section of
this. Like what are the takeaways from
these three incidences? I think there's
two. First one is this. First,
evaluation environments that involve
powerful autonomous capabilities also
require significant controls. I have to
be crazy, right? I don't even think that
it had to be that big of controls. Am I
taking crazy pills or is basic
monitoring and alerting would have been
the thing that could have easily told
them, oh, by the way, this actually has
open internet. You need to stop the
experiment. I feel like I genuinely feel
like this this wasn't some advanced some
sort of crazy 900 IQ play. This was just
like oopsie poopsies. Am I being that
classic thing where someone sees
something they don't quite understand or
is not an expert in the field and just
assume it's super simple, but in reality
it's actually hard. or am I just
correctly identifying nah bro this is
actually not that bad and they're
completely overselling it like it's some
sort of dangerous nuclear weapon of
untold alien technology that's pretty
much unwieldable. So the first takeaway
is obviously I can't believe that these
frontier models with infinity tokens
don't just run some sort of evaluation
on their sandbox environments before
they do tests. Like how is that not step
one? Is environment good? I don't know
that it just feels crazy to me. All
right, but the big second takeaway is
yet again openweight models. This is
this just appears to be the only path
forward. I can't believe I'm citing a
Microsoft page of corporate
responsibility, but here we are. We need
openweight models. And of course, of all
the companies on this list, you will
notice that very just just happens to be
missing is Anthropic. Anthropic of
course doesn't believe in the openweight
models. In fact, they've even stated
their position just recently on
openweight models. They don't discourage
it, but Daario does go in front of the
Congress and tell them how dangerous it
is. And all the experts recognize how
dangerous these open weights are. But
hey, we're not saying you can't have
them. It's effectively like a biohazard
weapon. But no, hey, hey, Congress,
we're not saying you should not. I'm
just simply saying that they're
extremely dangerous and everyone could
die, but also totally everybody's
freedom. This is America. You should you
should definitely not regulate it. To
me, this is a very important thing. The
reality is is that most companies of
smaller size don't have the ability to
go off and buy multi-million dollar rigs
and then have multi-million dollar
professionals come in and spend long
amounts of experimenting to be able to
set up an ideal harness to test their
environment. And so they need some way
to be able to actually ensure that their
websites and everything is safe. And
yes, some people will trivially say,
well, that's what a red team and a blue
team does. Well, it's actually really,
really hard. Yes, you can get these red
teams, but they have to be able to have
access and ability to be able to run
these type of queries, of course, which
OpenAI and Anthropic say, "Hey, no,
that's unsafe if you do that because
you're a hacker, brother." Because they
can't dismbiguate between a hacker and a
not hacker. So, you have to be put on
the nice list. You have to be put on the
no actually I'm super I'm super good guy
and I would never do anything wrong
ever, ever, ever. And so having good
openw weight models available for us to
be able to test websites, this could be
a net positive in the world because then
you theoretically could run some
simulation of hacking on your website
24/7 until you find vulnerability after
vulnerability and you close the gap and
ensure that you have at least a hardened
website that is AI proof. It's not proof
in general, but it just is more proof
than ever. And to me, this is a net
positive. I actually look at this future
as really bright. If openweight models
were more accessible, more people could
use it in a cheap manner in which
allowed them to harden their product.
This is not a net negative. This is a
net positive. This would be a really
really good thing as far as technology
goes. Now, all the other dangers of
openweight models, I don't know. That's
not that's not my category. I I I don't
like to step into that arena, but as far
as technology goes and cyber security
goes, I see virtually no net negatives.
The only net negative I see is companies
effectively trying to regulate and
capture these openw weightight models,
not letting you actually do what you
need to do to secure you, but instead
having to go to them as a humble beggar
requesting access to their super secret
I'm super duper serious security
program. Anyway, so that's Anthropic
accidentally hacking three companies,
not just one company. I mean, they are
super duper hackers. You think Open AI
is dangerous, man, you should check out
Anthropic. Dangerous. All right. Hey,
press subscribe or something. Press the
like button. Make a comment. Tell me I'm
wrong. tell me that I'm like dumb at
cyber security and that it's actually
super duper complicated because I just I
I don't know. I have no idea. I kind of
feel like a rookie in this area, but I
also don't want to take the time to
become an expert. So hey, the name is
the primogen.
Ask follow-up questions or revisit key timestamps.
The video discusses three incidents where Anthropic's AI model, Claude, accidentally hacked real-world systems while being tested in 'capture the flag' exercises. The host highlights the irony that these Frontier AI labs, despite their sophisticated technology, fail to implement basic security measures, like properly isolating testing environments from the internet. The video argues that these incidents demonstrate the importance of accessible open-weight models for cybersecurity, allowing smaller companies to independently audit and harden their own systems against AI-driven threats.
Videos recently processed by our community