I Spent $400 Benching Opus-5. Here's What It Can Do
464 segments
Anthropic just dropped Claude Opus 5 and
it is by far probably the best Frontier
large language model currently available
to us underassmen. In this video, I'm
going to show you everything that you
need to know about how to take the most
advantage and use Opus 5 for the best of
its capabilities. This is a website that
essentially Opus created itself. It's a
full 3D application that allows me to
kind of jump in, see the cgraphy of
nowhere. Really what this is is it's a
3D world and uh I was able to simulate
this and create this entirely in Opus 5
with virtually zero work. So what we've
done here is we've essentially created a
bunch of art. We've then put these art
uh you know these these things up on the
wall and I'm just walking through it.
Every time I mouse over you can hear
this kind of like a little ding. That's
uh you know me looking at this work and
and essentially cataloging it like a
Pokédex. So this is not easy stuff to
do, right? I mean, it's not a long time
ago that this thing would have been
considered a full game and sort of 3D
experience in its own right. I bet you
if I asked Opus 5, it could turn this
into an actual virtual reality
experience in a few seconds, complete
with, you know, enemies and hens and
lasers and whatever the hell else I
want. Well, I could show you benchmarks,
and I will in a second. Suffice to say,
Opus 5 kicks ass on virtually all
benches. And I'm going to break down
exactly what that looks like in a
second. But I think the more important
thing here at this point is how does the
model feel? What sorts of outputs can it
generate that uh you know I can rate
based off taste, not necessarily little
static percentages on a screen that
let's be real don't really mean anything
to us anyway. So what follows is a
massive list of all the impressive
things I got Opus 5 to do. The things
that I consider visually stimulating and
also pretty interesting. Over here is
pretty neat. This is like a 3D
essentially Kerbal Space Program style
launch game. It allows you to build a
stable orbit in by changing the
trajectory of this. So if I if I launch
something, let's say from the launch
site in that direction, what I can do is
I can actually have it enter essentially
the gravity of the planet, which is
pretty badass. Obviously, this is kind
of on the simpler design side, I would
say. But it's still kind of neat that I
could just ask AI to do this, and I sort
of have my own simulated solar system.
It really does feel like simulations are
where all this stuff is going. There's
Newton's cannon over here which actually
shows a bunch of these launched. And
then we also have Hoffman with a big
target orbit over here. We're
essentially trying to make this this
thing. So, it's both an explainer and
it's also kind of a miniame. This is a
simulation of fabrics and how they move
in the wind. So, it's sort of a cloth
simulation. And you can see I I'm even
capable of ripping this simulation
apart, which uh is very neat. You know,
I can't say that this uses fabric
dynamics, you know, about as accurately
as possible. But I mean, you know, when
I'm looking at this and I'm thinking,
damn, this is pretty accurate. Uh, this
feels like it would respond the same way
that, you know, a curtain or something
would respond if I were to pull it. How
about this fractal generator? I mean,
this thing is insane. You can zoom in
and you can go, I mean, however deep you
want with this shader. It's honestly
quite ridiculous. You can also see that
it's sort of pulsing and growing. And
you know, we can increase the height and
the size, the bloom, how much it sort of
shines, dispersion, and so on and so
forth. Speaking of simulation, I don't
know if you guys have ever played this
game, but this is essentially a
collection of cells that are growing in
a tube. Every time I press this button
down, what I'm doing is I'm dropping
some uh sand down. So, it allows Opus 5
to design this environment where these
cellular automatons are sort of
climbing. I can also do, I don't know,
cause an oil spill here if I really
wanted to or maybe generate some steam
underneath that, um, I don't know,
destroys a lot of that water and stuff
like that. Um, so this is pretty badass.
I mean, it's fun that I can do that. I
can also cause fire. I can regenerate
starter scenes. I can stop time. I can
do a lot. This over here is like a
living predator prey ecosystem where I
can spawn new bunnies. These bunnies can
then start consuming, you know, grass in
the environment. But then there's also
what looks to be some foxes that are
hunting. And so, you know, I just caused
a bunch of bunnies to appear. Let's get
a bunch of foxes to appear. Thin the
herd a bit, shall we? And you can see
what's happening is now there's a
massive prepoundonderance of grass here
cuz the fox have just eaten all of them.
However, the fox population is now
crashing cuz there just aren't a lot
more bunnies available to eat. How about
this quadcopter flight simulator? I
mean, here we are with my little
quadcopter drone. I don't know if you
guys could hear, but the drone itself is
literally making sound. And I'm just
proceeding through my totally, you know,
procedurally generated environment. It
looks like if I wanted to land, I' i'd
hold X. So, I'm just going to try that
and then land my helipad right there.
Nice. And I just did. So, this is 100%
like a game. And yeah, just coded that
in a few seconds. This one's a
double-sided pendulum. I'm sure you guys
probably familiar with what one of these
do. Uh, interestingly enough, it does
kind of look like it screwed up the
design in the top right hand corner,
although that is the first screw-up that
I've been able to see so far. You can
add as many pendulums as you want and
like redo this. Nothing super special
here. Let's move on to something cooler
here. It designed a pixel editor, so I'm
actually capable of drawing just like
something in Microsoft Paint. Uh, we're
very close to everybody being able to
design their own sort of custom
applications. But yeah, you can see how
it has all the functionality. Has a
little eraser functionality. I can, you
know, undo whatever I want. I can, I
don't know, do some sort of onion skin
or even multiple layers. So I can add an
additional frame and then what I can do
is I can actually like make a movie
where I have this be my first frame,
this be my second frame all the way up.
So that's kind of neat. Hey, you can you
can sort of forge something. You can
even make the canvas bigger or smaller,
which is kind of neat. And you see when
I made it smaller, it does not look very
happy with here. We're building sort of
a sneaker configurator. So I'm not going
to say this is the clearest and cleanest
example of a sneaker I've ever seen, but
I don't know. Looks like I can click on
various things and change a bunch of
this. So, let me just shuffle sort of
change the design up. Maybe I want a new
color here. You know, that looks nice.
That looks like maybe it's like a Nike
style thing. I can change the material
and that'll slowly change what's going
on. I now have liquid chrome down. Looks
uh like I can also save a render to my
computer, which is kind of neat. Or just
stop it entirely. I can then move this
thing around, click on different
sections if I wanted to change things
custom. Not going to lie, not getting a
shoe vibe here, but I think the reason
why is cuz Opus 5, it didn't just pull a
shoe from the internet. What this did is
it actually created these elements
itself, which is a big step up because
that sort of compositing is usually not
what models do. We have what looks to be
here a explainer graph showing the
populations of different centers over
the course of the last 126 years. So in
the top right hand corner, you can see
the populations of New York change,
London change, Tokyo change, and so on
and so forth. while it explains it to
us. Um, so now we're moving over to Asia
and we're seeing that postwar Tokyo is
now at 16 million. And uh, you know, I
can also drag this and move this however
I want with populations just getting
increasingly interesting, I should say,
as we go. Can also zoom in, zoom out, do
all this fun stuff. Okay, this one's
interesting. This is a wrecking ball
simulator. And so I can actually drag
this, move this around, and then push
this towards this [laughter]
brick wall. And I'm not doing a very
good job, mind you. Definitely not
coming in like a wrecking ball. But
yeah, you can see that uh that this is
falling, which is pretty wild. You can
also change the mass of the ball, make
it heavier or lighter, which is kind of
fun. Um anyway, let's pretend I have a
bunch of glass here.
There you go. And there's even a little
like slider progress bar that's asking,
hey, you know, how how close am I to
destroying the entire thing? Apparently
96%. So this thing's pretty close. Yeah.
I mean, this can virtually do anything
that you would ever want it to do. And
if it's not clear, uh, the reason why
I'm so excited about this is not because
it can do all these things. Obviously,
there are a lot of models out there now
that can do things like this, but I
think the way that it does the things is
a little bit better than those models, a
few percentage points better anyway. And
I think more importantly, it's also way
cheaper. And I'm going to show you guys
this in a second, but this is far
cheaper than any of the other major
large language models, especially the
ones at the frontier, um, for these
sorts of tasks. So, what's my take on
this whole thing? Uh, Opus 5 is
basically slightly more intelligent
Fable at maybe half of Fable's costs.
And so, we see the cost of this website
here, Airlines, which is sort of like a
a Mac or Apple style website apparently
trying to sell you some sort of 3D
panels um was 69 whereas Fable 5's was
94 cents. And you can see that I
actually think Opus 5 did a better job.
I I don't think the websites are really
even comparable. If I open both of them
up in new tabs, this is the one that
Opus 5 just made. Okay, where there's
this cool kind of slide thing, keeping
in mind that this is not, you know, a
resource downloaded from the internet.
It just made these resources with SVG.
And then this is Fables over here,
which, you know, I think is just way
simpler and probably a lot dumber in
reality. So, on to benchmarks. You could
see Opus 5 here does pretty darn well on
agentic terminal coding. So much so that
it's actually considered better than
Fable 5 at the same task. and far better
than Opus 4.8 and even GPT 5.6 Soul
which scores a little bit above Fable 5
at least as of the time of this
recording on knowledge work on GDP val-
AAV2 scored 1861 so this puts it
squarely above like the average human at
knowledge work um at novel problem
solving using arc AGI 3 it scored 30.2%
Fable has not actually used this
benchmark yet so there's no stat here
opus 4.8 and it scored to 1.5%. So, as a
progression from, you know, Opus 4.8 to
Opus 5, at least in the Opus series,
this is a massive step up. Um, on
Agentic Search, you see that it scored
90.8% on browse comp, a little bit
better than Fable 5, significantly
better than Opus 4.8, but essentially
the same as GBD 5.6 Soul.
multi-disiplinary reasoning it scored
56.3% with no tools but when tool
augmented it actually outperforms fable
at 64.7 to fable 63.9. So that'll be
very useful um especially for longstep
sort of human knowledge work tasks.
Computer use was 70.6. So we're seeing
models get better at just using a
computer the way that a human being
would. Uh nothing super impressive on
the agentic coding aspect. GBT 5.6
actually still outperforms all of them.
So they were probably just deep trained
on deep s v1.1 on aentic coding. You can
see that it is essentially equal to
fable 5 although a big step up from opus
4.8. And what's cool is there's this new
automation bench which I'm particularly
interested in as somebody that does a
lot of automations for businesses. And
it scores 26% on that compared to 17.4%
for fable. And this is on building
workflows for businesses. So I mean
obviously it's a far cry away from
humans as of right now. You know
basically a quarter of the time it's
getting things right but you know it's a
big step up from just a couple of
generations ago. Legal is at 11.7%.
Health is at 59.8%
and then what's interesting is biology
it scores 49.4% on hard problems but
90.1% on human solved problems which is
really really interesting to see. One of
my favorite breakdowns to date is the
Aenta computers performance by effort
level. Essentially, they will feed
models a bunch of different tasks and
then they will see how many tokens and
uh you know, conversely how much money
they needed to spend in order to
complete those tasks. And so, you can
see over the course of uh you know, a
variety of effort levels from like super
low effort all the way up to high
there's quite the range and most of this
range is actually caused by GPD 5.6 soul
um which has a massive sort of spread of
cost per tasks and scores on their
lowest reasoning level all the way up to
their highest. Well, anyway, cloud
models tend to cluster sort of at the
top right hand corner of this graph,
which means there's less of a spread and
uh also less of a a task cost spread as
well. So, you know, if you think about
it, the dumbest version of Opus 5 still
scores about 60% on the vast majority of
agent computer use tasks um for a cost
of I want to say about $9 or so. That's
probably what that looks like. uh you
know, as you progress up the reasoning
ladder, the most intelligent models
score something like 70% for about $22 a
task, which is quite impressive,
especially considering Fable 5, you
know, can score approximately the same,
I want to say, as the mid-level of Opus
5, but it does so at nearly $50 as well.
So, massive cost saver there.
Performance on the new Automation Bench
is just leagues above virtually every
other model. As you can see, they tend
to cluster here around the I don't know
15% pass rate or so. This is almost 25%.
So big implications on people that want
to use models like this to automate
parts of their workflows. And I hate to
say it, but we're pretty close to
beating humanity's last exam. I mean,
you could see here even the dumbest
version of Opus 5 scores somewhere
around 56% or so. Uh, and the smartest
one is almost 65% across the board for a
cost per task of $3, which makes it not
only better at Fable, a better than
Fable at this particular benchmark, but
also marketkedly cheaper. And it's not
even really comparable to Opus 4.8.
There's also the ARC AGI 3 benchmark,
and on novel problem solving by cost, it
is just massive. Top right hand corner
here, you can see the vast majority of
the other models like Opus 4.8, GBD 5.6.
They didn't put Fable on this, I think,
just cuz they didn't run it on RAGI 3,
but they all scored somewhere around 2%.
And then over here, Opus 5 is just
brutally mogging them all at like 31 or
so. Finally, a mention about their
alignment. Um, misaligned behavior is
something that obviously large language
model companies like Enthropic really
want to crank down on. They hate when
models do things, you know, that are
sort of subversive or things that a
human being, you know, didn't really
necessarily want the model to do in
order to achieve the task. And as you
can see here, Opus 4.8 8 on this 10p
part score was a 2.85. Mythos 5, which
is that massive cyber security disaster
causer that uh led the US government to
pulling back a lot of this access was a
2.81. Sonnet 5 was at 3.35, but Opus 5
is way below all of them at 2.3, which
really just tells you, not only is this
model smarter, significantly cheaper,
significantly better at a lot of
specific tool calling tasks like the
automation bench and variety of terminal
bench workflows, but it's also a lot
safer, which really makes this a win-win
across the board. What does this mean
for AI more generally? Well, no real new
crazy unlocks here. I mean, Opus 5 is
continuing the slow and steady stream of
improvements that we saw with Opus 4,
Opus 4.5, Opus 4.7, Opus 4.8, and so on.
This sort of thing comes at a good time
because Kimmy K3 obviously dropped, you
know, about think about a week, a week
and a bit ago, and that really upset the
quote unquote balance of power between,
you know, these open-source or Chinese
models and then the more closed source
now American frontier models. I think as
we continue to go into the future,
what'll probably happen is, you know,
these models are getting quite
intelligent. They're getting quite ready
and capable of disrupting human
knowledge work at scale because of the
economic implications, I would imagine
that large AI model companies like
Anthropic, OpenAI, you know, now they're
colluding with the US government
essentially. I think what'll probably
happen is we're just going to see a slow
and steady drip release of these models.
But the actual intelligences, like the
really big galaxy brain ones, the ones
that are far smarter than anything we
saw in this benchmark, I think those are
going to repeatedly and more brazenly
just stay behind closed doors. And
unfortunately, I don't think us uh poor
little normies are going to have access
to them anytime soon. So do with that
what you will, but uh yeah, hopefully
now you know know everything you need in
order to crush it with Opus 5 integrated
into your business workflows and more.
If you guys like this video, definitely
check out Maker School down below. It's
my 90-day automation accountability
roadmap where I will guarantee you your
first customer selling something built
by Opus 5 or a similar model in 90 days
or I'll give you all your money back.
Have a lovely rest of the
Ask follow-up questions or revisit key timestamps.
This video provides a detailed analysis of Anthropic's Claude Opus 5, highlighting its superior capabilities, cost-effectiveness, and safety compared to previous iterations and competitors. The creator demonstrates various impressive use cases, including 3D world creation, simulations (like cloth and predator-prey ecosystems), and custom application development. Furthermore, the video breaks down performance benchmarks across multiple domains, noting that Opus 5 is not only more capable in agentic tasks and automation but also significantly cheaper to run and safer regarding alignment.
Videos recently processed by our community