HomeVideos

I Spent $400 Benching Opus-5. Here's What It Can Do

Now Playing

I Spent $400 Benching Opus-5. Here's What It Can Do

Transcript

464 segments

0:00

Anthropic just dropped Claude Opus 5 and

0:02

it is by far probably the best Frontier

0:05

large language model currently available

0:07

to us underassmen. In this video, I'm

0:09

going to show you everything that you

0:10

need to know about how to take the most

0:12

advantage and use Opus 5 for the best of

0:14

its capabilities. This is a website that

0:16

essentially Opus created itself. It's a

0:19

full 3D application that allows me to

0:22

kind of jump in, see the cgraphy of

0:25

nowhere. Really what this is is it's a

0:27

3D world and uh I was able to simulate

0:30

this and create this entirely in Opus 5

0:32

with virtually zero work. So what we've

0:35

done here is we've essentially created a

0:36

bunch of art. We've then put these art

0:39

uh you know these these things up on the

0:40

wall and I'm just walking through it.

0:42

Every time I mouse over you can hear

0:44

this kind of like a little ding. That's

0:46

uh you know me looking at this work and

0:48

and essentially cataloging it like a

0:49

Pokédex. So this is not easy stuff to

0:52

do, right? I mean, it's not a long time

0:54

ago that this thing would have been

0:55

considered a full game and sort of 3D

0:58

experience in its own right. I bet you

0:59

if I asked Opus 5, it could turn this

1:01

into an actual virtual reality

1:02

experience in a few seconds, complete

1:04

with, you know, enemies and hens and

1:07

lasers and whatever the hell else I

1:08

want. Well, I could show you benchmarks,

1:10

and I will in a second. Suffice to say,

1:12

Opus 5 kicks ass on virtually all

1:14

benches. And I'm going to break down

1:15

exactly what that looks like in a

1:16

second. But I think the more important

1:18

thing here at this point is how does the

1:20

model feel? What sorts of outputs can it

1:22

generate that uh you know I can rate

1:24

based off taste, not necessarily little

1:27

static percentages on a screen that

1:29

let's be real don't really mean anything

1:30

to us anyway. So what follows is a

1:32

massive list of all the impressive

1:34

things I got Opus 5 to do. The things

1:35

that I consider visually stimulating and

1:37

also pretty interesting. Over here is

1:39

pretty neat. This is like a 3D

1:41

essentially Kerbal Space Program style

1:44

launch game. It allows you to build a

1:47

stable orbit in by changing the

1:50

trajectory of this. So if I if I launch

1:53

something, let's say from the launch

1:54

site in that direction, what I can do is

1:57

I can actually have it enter essentially

1:59

the gravity of the planet, which is

2:00

pretty badass. Obviously, this is kind

2:03

of on the simpler design side, I would

2:05

say. But it's still kind of neat that I

2:07

could just ask AI to do this, and I sort

2:09

of have my own simulated solar system.

2:10

It really does feel like simulations are

2:12

where all this stuff is going. There's

2:14

Newton's cannon over here which actually

2:15

shows a bunch of these launched. And

2:17

then we also have Hoffman with a big

2:19

target orbit over here. We're

2:21

essentially trying to make this this

2:23

thing. So, it's both an explainer and

2:24

it's also kind of a miniame. This is a

2:26

simulation of fabrics and how they move

2:29

in the wind. So, it's sort of a cloth

2:31

simulation. And you can see I I'm even

2:33

capable of ripping this simulation

2:35

apart, which uh is very neat. You know,

2:38

I can't say that this uses fabric

2:41

dynamics, you know, about as accurately

2:43

as possible. But I mean, you know, when

2:46

I'm looking at this and I'm thinking,

2:47

damn, this is pretty accurate. Uh, this

2:49

feels like it would respond the same way

2:50

that, you know, a curtain or something

2:53

would respond if I were to pull it. How

2:54

about this fractal generator? I mean,

2:56

this thing is insane. You can zoom in

2:58

and you can go, I mean, however deep you

3:00

want with this shader. It's honestly

3:02

quite ridiculous. You can also see that

3:04

it's sort of pulsing and growing. And

3:06

you know, we can increase the height and

3:08

the size, the bloom, how much it sort of

3:11

shines, dispersion, and so on and so

3:13

forth. Speaking of simulation, I don't

3:15

know if you guys have ever played this

3:16

game, but this is essentially a

3:17

collection of cells that are growing in

3:19

a tube. Every time I press this button

3:21

down, what I'm doing is I'm dropping

3:22

some uh sand down. So, it allows Opus 5

3:26

to design this environment where these

3:28

cellular automatons are sort of

3:30

climbing. I can also do, I don't know,

3:31

cause an oil spill here if I really

3:33

wanted to or maybe generate some steam

3:36

underneath that, um, I don't know,

3:38

destroys a lot of that water and stuff

3:40

like that. Um, so this is pretty badass.

3:42

I mean, it's fun that I can do that. I

3:43

can also cause fire. I can regenerate

3:45

starter scenes. I can stop time. I can

3:47

do a lot. This over here is like a

3:49

living predator prey ecosystem where I

3:51

can spawn new bunnies. These bunnies can

3:54

then start consuming, you know, grass in

3:56

the environment. But then there's also

3:58

what looks to be some foxes that are

4:00

hunting. And so, you know, I just caused

4:02

a bunch of bunnies to appear. Let's get

4:03

a bunch of foxes to appear. Thin the

4:05

herd a bit, shall we? And you can see

4:07

what's happening is now there's a

4:08

massive prepoundonderance of grass here

4:10

cuz the fox have just eaten all of them.

4:12

However, the fox population is now

4:14

crashing cuz there just aren't a lot

4:15

more bunnies available to eat. How about

4:17

this quadcopter flight simulator? I

4:20

mean, here we are with my little

4:21

quadcopter drone. I don't know if you

4:23

guys could hear, but the drone itself is

4:24

literally making sound. And I'm just

4:27

proceeding through my totally, you know,

4:29

procedurally generated environment. It

4:32

looks like if I wanted to land, I' i'd

4:33

hold X. So, I'm just going to try that

4:35

and then land my helipad right there.

4:37

Nice. And I just did. So, this is 100%

4:39

like a game. And yeah, just coded that

4:41

in a few seconds. This one's a

4:42

double-sided pendulum. I'm sure you guys

4:44

probably familiar with what one of these

4:46

do. Uh, interestingly enough, it does

4:48

kind of look like it screwed up the

4:50

design in the top right hand corner,

4:51

although that is the first screw-up that

4:53

I've been able to see so far. You can

4:55

add as many pendulums as you want and

4:56

like redo this. Nothing super special

4:58

here. Let's move on to something cooler

5:00

here. It designed a pixel editor, so I'm

5:02

actually capable of drawing just like

5:04

something in Microsoft Paint. Uh, we're

5:06

very close to everybody being able to

5:08

design their own sort of custom

5:09

applications. But yeah, you can see how

5:11

it has all the functionality. Has a

5:12

little eraser functionality. I can, you

5:14

know, undo whatever I want. I can, I

5:17

don't know, do some sort of onion skin

5:18

or even multiple layers. So I can add an

5:21

additional frame and then what I can do

5:23

is I can actually like make a movie

5:24

where I have this be my first frame,

5:26

this be my second frame all the way up.

5:29

So that's kind of neat. Hey, you can you

5:30

can sort of forge something. You can

5:31

even make the canvas bigger or smaller,

5:33

which is kind of neat. And you see when

5:35

I made it smaller, it does not look very

5:36

happy with here. We're building sort of

5:38

a sneaker configurator. So I'm not going

5:40

to say this is the clearest and cleanest

5:43

example of a sneaker I've ever seen, but

5:45

I don't know. Looks like I can click on

5:47

various things and change a bunch of

5:49

this. So, let me just shuffle sort of

5:51

change the design up. Maybe I want a new

5:53

color here. You know, that looks nice.

5:54

That looks like maybe it's like a Nike

5:56

style thing. I can change the material

5:58

and that'll slowly change what's going

6:00

on. I now have liquid chrome down. Looks

6:03

uh like I can also save a render to my

6:05

computer, which is kind of neat. Or just

6:06

stop it entirely. I can then move this

6:08

thing around, click on different

6:10

sections if I wanted to change things

6:12

custom. Not going to lie, not getting a

6:14

shoe vibe here, but I think the reason

6:16

why is cuz Opus 5, it didn't just pull a

6:18

shoe from the internet. What this did is

6:19

it actually created these elements

6:21

itself, which is a big step up because

6:24

that sort of compositing is usually not

6:25

what models do. We have what looks to be

6:27

here a explainer graph showing the

6:30

populations of different centers over

6:32

the course of the last 126 years. So in

6:36

the top right hand corner, you can see

6:37

the populations of New York change,

6:39

London change, Tokyo change, and so on

6:41

and so forth. while it explains it to

6:43

us. Um, so now we're moving over to Asia

6:45

and we're seeing that postwar Tokyo is

6:47

now at 16 million. And uh, you know, I

6:50

can also drag this and move this however

6:52

I want with populations just getting

6:54

increasingly interesting, I should say,

6:56

as we go. Can also zoom in, zoom out, do

6:58

all this fun stuff. Okay, this one's

7:00

interesting. This is a wrecking ball

7:01

simulator. And so I can actually drag

7:03

this, move this around, and then push

7:07

this towards this [laughter]

7:09

brick wall. And I'm not doing a very

7:11

good job, mind you. Definitely not

7:13

coming in like a wrecking ball. But

7:14

yeah, you can see that uh that this is

7:16

falling, which is pretty wild. You can

7:18

also change the mass of the ball, make

7:19

it heavier or lighter, which is kind of

7:21

fun. Um anyway, let's pretend I have a

7:23

bunch of glass here.

7:26

There you go. And there's even a little

7:28

like slider progress bar that's asking,

7:31

hey, you know, how how close am I to

7:33

destroying the entire thing? Apparently

7:35

96%. So this thing's pretty close. Yeah.

7:37

I mean, this can virtually do anything

7:38

that you would ever want it to do. And

7:40

if it's not clear, uh, the reason why

7:42

I'm so excited about this is not because

7:44

it can do all these things. Obviously,

7:45

there are a lot of models out there now

7:47

that can do things like this, but I

7:49

think the way that it does the things is

7:51

a little bit better than those models, a

7:52

few percentage points better anyway. And

7:54

I think more importantly, it's also way

7:56

cheaper. And I'm going to show you guys

7:58

this in a second, but this is far

7:59

cheaper than any of the other major

8:02

large language models, especially the

8:03

ones at the frontier, um, for these

8:05

sorts of tasks. So, what's my take on

8:06

this whole thing? Uh, Opus 5 is

8:08

basically slightly more intelligent

8:10

Fable at maybe half of Fable's costs.

8:13

And so, we see the cost of this website

8:15

here, Airlines, which is sort of like a

8:17

a Mac or Apple style website apparently

8:21

trying to sell you some sort of 3D

8:23

panels um was 69 whereas Fable 5's was

8:27

94 cents. And you can see that I

8:28

actually think Opus 5 did a better job.

8:30

I I don't think the websites are really

8:32

even comparable. If I open both of them

8:33

up in new tabs, this is the one that

8:35

Opus 5 just made. Okay, where there's

8:37

this cool kind of slide thing, keeping

8:40

in mind that this is not, you know, a

8:42

resource downloaded from the internet.

8:43

It just made these resources with SVG.

8:46

And then this is Fables over here,

8:47

which, you know, I think is just way

8:49

simpler and probably a lot dumber in

8:51

reality. So, on to benchmarks. You could

8:53

see Opus 5 here does pretty darn well on

8:57

agentic terminal coding. So much so that

8:59

it's actually considered better than

9:00

Fable 5 at the same task. and far better

9:03

than Opus 4.8 and even GPT 5.6 Soul

9:07

which scores a little bit above Fable 5

9:09

at least as of the time of this

9:10

recording on knowledge work on GDP val-

9:13

AAV2 scored 1861 so this puts it

9:16

squarely above like the average human at

9:18

knowledge work um at novel problem

9:20

solving using arc AGI 3 it scored 30.2%

9:25

Fable has not actually used this

9:27

benchmark yet so there's no stat here

9:29

opus 4.8 and it scored to 1.5%. So, as a

9:32

progression from, you know, Opus 4.8 to

9:34

Opus 5, at least in the Opus series,

9:36

this is a massive step up. Um, on

9:38

Agentic Search, you see that it scored

9:40

90.8% on browse comp, a little bit

9:43

better than Fable 5, significantly

9:45

better than Opus 4.8, but essentially

9:48

the same as GBD 5.6 Soul.

9:50

multi-disiplinary reasoning it scored

9:52

56.3% with no tools but when tool

9:55

augmented it actually outperforms fable

9:58

at 64.7 to fable 63.9. So that'll be

10:02

very useful um especially for longstep

10:04

sort of human knowledge work tasks.

10:06

Computer use was 70.6. So we're seeing

10:08

models get better at just using a

10:10

computer the way that a human being

10:11

would. Uh nothing super impressive on

10:13

the agentic coding aspect. GBT 5.6

10:16

actually still outperforms all of them.

10:18

So they were probably just deep trained

10:19

on deep s v1.1 on aentic coding. You can

10:23

see that it is essentially equal to

10:25

fable 5 although a big step up from opus

10:28

4.8. And what's cool is there's this new

10:31

automation bench which I'm particularly

10:32

interested in as somebody that does a

10:34

lot of automations for businesses. And

10:36

it scores 26% on that compared to 17.4%

10:40

for fable. And this is on building

10:42

workflows for businesses. So I mean

10:44

obviously it's a far cry away from

10:46

humans as of right now. You know

10:48

basically a quarter of the time it's

10:49

getting things right but you know it's a

10:50

big step up from just a couple of

10:52

generations ago. Legal is at 11.7%.

10:56

Health is at 59.8%

10:58

and then what's interesting is biology

11:00

it scores 49.4% on hard problems but

11:03

90.1% on human solved problems which is

11:06

really really interesting to see. One of

11:07

my favorite breakdowns to date is the

11:09

Aenta computers performance by effort

11:12

level. Essentially, they will feed

11:13

models a bunch of different tasks and

11:15

then they will see how many tokens and

11:18

uh you know, conversely how much money

11:19

they needed to spend in order to

11:20

complete those tasks. And so, you can

11:22

see over the course of uh you know, a

11:25

variety of effort levels from like super

11:27

low effort all the way up to high

11:29

there's quite the range and most of this

11:30

range is actually caused by GPD 5.6 soul

11:33

um which has a massive sort of spread of

11:36

cost per tasks and scores on their

11:38

lowest reasoning level all the way up to

11:40

their highest. Well, anyway, cloud

11:41

models tend to cluster sort of at the

11:43

top right hand corner of this graph,

11:44

which means there's less of a spread and

11:47

uh also less of a a task cost spread as

11:50

well. So, you know, if you think about

11:51

it, the dumbest version of Opus 5 still

11:54

scores about 60% on the vast majority of

11:57

agent computer use tasks um for a cost

11:59

of I want to say about $9 or so. That's

12:01

probably what that looks like. uh you

12:03

know, as you progress up the reasoning

12:05

ladder, the most intelligent models

12:07

score something like 70% for about $22 a

12:11

task, which is quite impressive,

12:12

especially considering Fable 5, you

12:15

know, can score approximately the same,

12:17

I want to say, as the mid-level of Opus

12:19

5, but it does so at nearly $50 as well.

12:22

So, massive cost saver there.

12:24

Performance on the new Automation Bench

12:25

is just leagues above virtually every

12:28

other model. As you can see, they tend

12:29

to cluster here around the I don't know

12:31

15% pass rate or so. This is almost 25%.

12:34

So big implications on people that want

12:36

to use models like this to automate

12:37

parts of their workflows. And I hate to

12:39

say it, but we're pretty close to

12:41

beating humanity's last exam. I mean,

12:43

you could see here even the dumbest

12:44

version of Opus 5 scores somewhere

12:46

around 56% or so. Uh, and the smartest

12:50

one is almost 65% across the board for a

12:53

cost per task of $3, which makes it not

12:56

only better at Fable, a better than

12:58

Fable at this particular benchmark, but

12:59

also marketkedly cheaper. And it's not

13:01

even really comparable to Opus 4.8.

13:03

There's also the ARC AGI 3 benchmark,

13:05

and on novel problem solving by cost, it

13:08

is just massive. Top right hand corner

13:10

here, you can see the vast majority of

13:12

the other models like Opus 4.8, GBD 5.6.

13:15

They didn't put Fable on this, I think,

13:16

just cuz they didn't run it on RAGI 3,

13:18

but they all scored somewhere around 2%.

13:21

And then over here, Opus 5 is just

13:23

brutally mogging them all at like 31 or

13:25

so. Finally, a mention about their

13:27

alignment. Um, misaligned behavior is

13:29

something that obviously large language

13:31

model companies like Enthropic really

13:32

want to crank down on. They hate when

13:34

models do things, you know, that are

13:36

sort of subversive or things that a

13:38

human being, you know, didn't really

13:40

necessarily want the model to do in

13:41

order to achieve the task. And as you

13:43

can see here, Opus 4.8 8 on this 10p

13:45

part score was a 2.85. Mythos 5, which

13:49

is that massive cyber security disaster

13:52

causer that uh led the US government to

13:54

pulling back a lot of this access was a

13:56

2.81. Sonnet 5 was at 3.35, but Opus 5

14:01

is way below all of them at 2.3, which

14:03

really just tells you, not only is this

14:04

model smarter, significantly cheaper,

14:07

significantly better at a lot of

14:09

specific tool calling tasks like the

14:12

automation bench and variety of terminal

14:14

bench workflows, but it's also a lot

14:17

safer, which really makes this a win-win

14:19

across the board. What does this mean

14:21

for AI more generally? Well, no real new

14:24

crazy unlocks here. I mean, Opus 5 is

14:27

continuing the slow and steady stream of

14:30

improvements that we saw with Opus 4,

14:32

Opus 4.5, Opus 4.7, Opus 4.8, and so on.

14:37

This sort of thing comes at a good time

14:38

because Kimmy K3 obviously dropped, you

14:40

know, about think about a week, a week

14:42

and a bit ago, and that really upset the

14:44

quote unquote balance of power between,

14:46

you know, these open-source or Chinese

14:48

models and then the more closed source

14:50

now American frontier models. I think as

14:53

we continue to go into the future,

14:54

what'll probably happen is, you know,

14:56

these models are getting quite

14:57

intelligent. They're getting quite ready

14:58

and capable of disrupting human

15:00

knowledge work at scale because of the

15:02

economic implications, I would imagine

15:03

that large AI model companies like

15:05

Anthropic, OpenAI, you know, now they're

15:08

colluding with the US government

15:09

essentially. I think what'll probably

15:11

happen is we're just going to see a slow

15:12

and steady drip release of these models.

15:15

But the actual intelligences, like the

15:17

really big galaxy brain ones, the ones

15:18

that are far smarter than anything we

15:20

saw in this benchmark, I think those are

15:22

going to repeatedly and more brazenly

15:24

just stay behind closed doors. And

15:26

unfortunately, I don't think us uh poor

15:28

little normies are going to have access

15:29

to them anytime soon. So do with that

15:31

what you will, but uh yeah, hopefully

15:33

now you know know everything you need in

15:34

order to crush it with Opus 5 integrated

15:37

into your business workflows and more.

15:38

If you guys like this video, definitely

15:40

check out Maker School down below. It's

15:41

my 90-day automation accountability

15:43

roadmap where I will guarantee you your

15:45

first customer selling something built

15:47

by Opus 5 or a similar model in 90 days

15:49

or I'll give you all your money back.

15:51

Have a lovely rest of the

Interactive Summary

This video provides a detailed analysis of Anthropic's Claude Opus 5, highlighting its superior capabilities, cost-effectiveness, and safety compared to previous iterations and competitors. The creator demonstrates various impressive use cases, including 3D world creation, simulations (like cloth and predator-prey ecosystems), and custom application development. Furthermore, the video breaks down performance benchmarks across multiple domains, noting that Opus 5 is not only more capable in agentic tasks and automation but also significantly cheaper to run and safer regarding alignment.

Suggested questions

3 ready-made prompts