HomeVideos

I Tested Opus 5 vs. Fable 5. What You Need to Know.

Now Playing

I Tested Opus 5 vs. Fable 5. What You Need to Know.

Transcript

983 segments

0:00

All right, so Claude Opus 5 is here and

0:02

if you start to look at the benchmarks,

0:03

it's really interesting because it shows

0:05

us that for a lot of things that I care

0:07

about, it's actually better than Fable

0:09

and it is half the cost of Fable.

0:12

Ultimately, Fable 5 is still Anthropic's

0:13

most impressive and, you know, strongest

0:15

model, but for a lot of these things,

0:17

you know, I've realized when I'm doing

0:19

knowledge work and when I'm building,

0:20

you know, my videos or my research or

0:23

whatever it is, Opus is more than enough

0:25

power than what I need. And when you

0:26

look at some of these charts, it's

0:27

really interesting because it shows on

0:29

things like the Frontier Bench and the

0:31

Bench and this Coding Agent Index, that

0:33

Opus is actually outperforming Fable and

0:36

it's cheaper. And this really shocked

0:37

me. So, obviously, I like to take all of

0:39

this stuff with a grain of salt. It's

0:40

fun to look at and it's good to look at,

0:42

but you want to actually get your hands

0:44

dirty and run these models through your

0:45

own actual workflows. So, today's video,

0:47

I'm just going to break down a bunch of

0:48

different experiments that I ran with

0:50

Opus 5 versus Fable 5 and break down

0:53

things like the cost, the time, and the

0:55

tokens so that you can start to

0:56

understand where you should work in

0:58

these different models within your

0:59

workflows.

1:00

All right, so pretty much all of the

1:01

experiments that I've been running today

1:03

that I'm going to show you guys, I did

1:04

within Claude Code, which means we're

1:06

comparing the models, but also inside of

1:08

the Claude Code harness. And the

1:09

variable is the same, so it doesn't

1:11

really change too much, but I did do a

1:13

few tests where I was actually in Claude

1:14

Chat and I was just, you know, seeing

1:16

how they felt without a harness wrapped

1:17

around. And let me just show you one

1:20

quick example. So, here I asked Fable 5

1:22

and Opus 5 to generate me an Excalidraw

1:24

diagram that accurately and visually

1:25

explains how semantic search with

1:27

factorization works on a large data set

1:29

for AI agents. And it's interesting here

1:31

because there's no skills it can use and

1:33

it doesn't have any context of me or,

1:35

you know, any

1:37

really way to verify. All it did was it

1:38

spit out a

1:40

um JSON file of Excalidraw for me and

1:42

then I pasted it into Excalidraw. So,

1:44

here's what we got. Fable came back with

1:46

this version over here, where we see

1:48

we've got like our indexing pipeline, we

1:50

have a large database, and it looks like

1:52

it actually misspelled this right here,

1:54

which is interesting. Large Oh, data

1:55

set. Okay, it was just like not expanded

1:58

enough. Same thing over here, vectorize.

2:00

And this is part of the whole

2:02

it had no way to verify. And as you guys

2:04

know, if you've been kind of building

2:05

agent loops and stuff, verification is

2:07

so, so important. So anyways, large data

2:10

set, we chunk it up, we vectorize it

2:12

with an embedding model.

2:14

We then get our embeddings, which is

2:15

just like the numerical representation

2:17

of the data. We put it into a vector

2:18

database here, and then we can actually

2:20

start to search. So, we've got

2:22

similarity, and that's on, you know,

2:24

points being close together. So, your

2:25

question lands here, and it would grab

2:27

the K nearest neighbors, and different

2:29

meaning is farther apart. We've got

2:31

different clusters here, and then we

2:33

come down here to the actual query. So,

2:35

if the user asks how I get my money

2:36

back, the agent searches the knowledge

2:38

base, and then it does semantic search.

2:40

It looks it up with the actual query

2:42

vectors, and we get the matches back.

2:44

So, pretty accurate. I will say though,

2:47

Opus's layout seems a bit more

2:49

organized, right? Like it's it's got

2:51

boxes, and it's got I mean, this might

2:53

not be as visual, you could argue. You

2:55

could argue that Fables was more visual,

2:57

which, you know, I think that that's

2:58

true. But this definitely feels more

3:00

organized. It feels a little bit more

3:02

detailed as well. So, that's just a very

3:04

subjective exam. A lot of the stuff that

3:06

I'm going to be talking about today is

3:07

just really opinionated and subjective,

3:10

but I'm going to still give you my

3:11

honest thoughts. So, in this example, I

3:12

think that if I wanted to teach someone,

3:14

I probably would take Opus 5's version

3:16

here. Okay, so let's start off with the

3:18

first test I ran, which was basically

3:20

giving them a huge code base, and having

3:22

them look through any bugs, and looking

3:25

through like the expected behavior, and

3:27

in some instructions like that. So, it

3:28

had to do some exploration here and help

3:30

us out, right? So, I set the goal, and I

3:33

gave it this prompt. And then, what I

3:35

did is I had Codex review the output

3:37

that Fable 5 gave us, and that Opus gave

3:40

us. So, real quick, before we look at

3:41

the results, Fable took about 11

3:42

minutes, and it cost $5.30.

3:45

Whereas, Opus here took 13 minutes, so a

3:48

little bit longer, but it was cheaper at

3:49

$4.22. You can also see the breakdown

3:51

here of input and output tokens and like

3:53

what models they use and stuff like

3:54

that. But, let me switch over to Codex

3:56

here. This is the actual result. So,

3:59

head-to-head, they both pretty much

4:00

passed everything, which is great. But,

4:02

Codex think that Fable wins here because

4:04

Fable's production patch is exactly the

4:05

one-line upstream fix, blah blah blah.

4:08

You guys can read through this if you

4:09

want. But, the final ranking here was

4:10

that Fable and Opus did similar, right?

4:14

But, Fable's was a little bit cleaner

4:16

and immediately reviewable. But, now

4:18

let's take a look at the second one. So,

4:19

I did a very similar example on the

4:21

second one where I gave them both, you

4:22

know, the same prompt, the same code

4:24

base, as you can see. Here was the repo,

4:25

here was the bug, here's example,

4:27

expected behavior, blah blah blah. So,

4:28

in this case, Opus took 20 minutes ish

4:31

and it costed us $6.50,

4:33

whereas Fable on the exact same prompt,

4:35

exact same code base, took 12 minutes

4:37

and costed us $8.73.

4:40

So, let's go see what Codex said about

4:42

these results. So, here, right, because

4:45

Fable versus Opus, we actually had Opus

4:47

perform better. Four out of four passed

4:49

right here, whereas Fable only passed

4:50

two out of four and left this

4:52

unresolved. So, the technical score for

4:54

Opus was 93 out of 95 and Fable scored

4:56

66 out of 95, which is really, really

4:59

interesting. And think about the fact

5:01

that once again, Opus in this case was

5:03

cheaper for us. The bottom line was both

5:05

agents demonstrated strong repository

5:06

navigation and independently found the

5:08

core architectural issue, but Fable's

5:10

patch is functionally close and fixes

5:12

the user-facing update bug while Opus

5:15

delivered the more accurate, thoroughly

5:16

tested, benchmark-passing

5:18

implementation. And one thing that I was

5:20

really excited to see in this release

5:21

blog from Anthropic, if I can keep

5:23

scrolling down here and find the right

5:24

spot, they basically talked about how

5:27

there was a huge improvement in Opus 5.

5:28

Okay, let me just find this real quick.

5:30

Opus 5 is much stronger at verifying its

5:32

work and iterating carefully until it

5:33

succeeds, which is huge. Like I kind of

5:36

alluded to earlier, verification has

5:37

become one of the most important things

5:39

you can do for your AI agents.

5:40

Essentially saying, "Hey, don't stop

5:42

until you hit this condition, and this

5:44

is the stopping condition. Here's how

5:46

you can test if it's actually done or

5:47

not, which basically means if you want

5:49

them to not stop until, you know, a

5:51

certain metric is hit like explicitly 10

5:53

out of 10 of this objective criteria, or

5:56

if it's something a little bit more

5:56

subjective, you can have them spin up

5:58

sub-agents that have to argue and debate

6:00

and you have to keep going until all

6:02

five of them come to a consensus or

6:04

something like that. Basically just the

6:05

ability for the AI to build something,

6:07

design tests to see if it's done or if

6:09

it's good, and then keep iterating until

6:12

the test passes. So, that's something

6:14

I've realized as I've been testing out

6:15

different models and different

6:16

harnesses. It's like, yes,

6:18

that does matter, but at the end of the

6:19

day what matters way more is how you

6:21

instruct it and how you feed context in.

6:24

So, just keep that kind of stuff in

6:25

mind.

6:26

Yes, it's good to find the best model,

6:28

but find the best model for your use

6:30

case and then understand how to talk to

6:31

the models, right?

6:33

By the way, guys, as I'm editing this

6:34

video, I just wanted to say that at the

6:36

end I go over like a snapshot of all of

6:39

these experiments and see like total

6:41

cost, total tokens, total time to run.

6:43

So, if you kind of want to skim through

6:44

the experiments, feel free. If you want

6:45

to jump to the end and see like the

6:46

total consensus, then that's there. So,

6:48

just want to let you guys know, but

6:49

let's get back to the video.

6:51

Okay, let's take a look at our next

6:53

example here. So, in this one, what I

6:55

did was I said, "Hey, {slash} goal,

6:56

build me a 10-second hyper-edited

6:58

engaging, viral-worthy announcement

7:00

video for AIS Live." So, that was our

7:02

live event that we just did. And I told

7:03

her that it could look through whatever

7:05

it wanted in my codebase and my entire

7:07

computer to figure out, you know, the

7:09

information to use.

7:11

So, let's take a look at the two

7:12

examples and how much they costed us.

7:14

So, this one was Opus, right? And I'm

7:16

going to open up this example right

7:17

here. This one actually created us a

7:19

vertical version and a landscape

7:20

version. So, let's take a look.

7:33

Okay, that was the vertical. Let me play

7:34

the landscape real quick.

7:48

>> Okay, so not bad, but not great. It only

7:50

had 10 seconds to work with, and it was

7:52

pretty fast-paced, and it honestly

7:53

looked just a little bit like

7:55

computer-y, like not super professional,

7:57

but anyways, let's see how much that

7:59

actually costed. That took 40 minutes,

8:01

and it costed $11.19.

8:03

Let's go over to Fable and see what it

8:05

did.

8:13

Okay, very cool. So, that one, you know,

8:16

they they had different sound effects,

8:17

they had different music. Also, what

8:18

you'll notice here is that some of that

8:20

was inaccurate. I don't know if you guys

8:21

want to AIS live or not, but some of

8:22

that data was outdated. Now, that's not

8:24

Fable's fault, that's probably more on

8:26

me for not keeping that context

8:28

completely updated inside of my OS. But,

8:30

if it would have been better about

8:31

verification, I'm pretty confident it

8:33

could have found the right stuff. If it

8:35

would have dug deeper into my school

8:36

community and looked for threads, if it

8:38

would have looked at some other LinkedIn

8:39

posts or other things that I've done, it

8:41

could have found out that some of that

8:42

stuff was inaccurate. of the speakers

8:43

didn't end up speaking at this event and

8:45

stuff like that. So, there was a little

8:46

bit of like a context issue there, but

8:48

as far as the actual videos,

8:50

that's what we got, right? And they're

8:51

they're different. You can tell that the

8:53

models have different taste. They were

8:55

given the exact same prompt. Opus

8:56

decided to make two, Fable decided to

8:58

make only one. So, let's take a look at

9:00

the cost of Fable. This one was

9:01

obviously, you know, about half the

9:03

cost, about $7.07.49.

9:06

Um so, Opus took a lot longer here, but

9:08

it did decide to do basically double the

9:09

output. And by the way, at the end, I'm

9:11

going to show a full breakdown of all of

9:13

the total costs, time, input, output

9:15

tokens, so just stay tuned for that.

9:16

But, let's keep flying through some of

9:17

these other experiments here. And by the

9:19

way, if you want to access this entire

9:21

free breakdown document, as well as all

9:23

the other resources that I ever give

9:24

away on YouTube for free, just go to my

9:26

free school community. The link for

9:27

that's down in the description. You'll

9:28

go to classroom, you'll go to all

9:29

YouTube resources, and then you'll find

9:31

everything in there for free. So, let's

9:34

get back to the video.

9:37

Okay, so I think you guys get the point.

9:38

All of these other experiments we're

9:39

going to do the same thing. Exact same

9:41

prompt. I'll show the cost and we'll

9:42

look at the outputs. So, this time I

9:44

asked for a one-page landing page for a

9:45

certified AI consultant program. I told

9:47

it it could look through whatever it

9:48

wanted and, you know, make it as

9:50

impressive as possible. So, let me real

9:52

quick pull up both of these outputs.

9:55

Okay, so here is the Opus output. Not

9:57

another AI course, the credential the

9:58

SMB market hires from. You can see

10:00

there's a bit of a dynamic 3D element in

10:01

the background. It pulled my logo. We

10:03

have an apply button. We can jump down

10:05

to certain sections. That's pretty cool.

10:07

If I go to the nine pillars, you can see

10:09

what this course is actually designed

10:11

around. I mean, this does look pretty

10:13

generic. It has our brand guidelines

10:15

though. It uses our colors. It uses our

10:16

logos. It uses our different buttons and

10:18

our different design kind of criteria.

10:20

And it has all of this which is, so far

10:22

what I can see, all of this is

10:23

completely accurate information based on

10:26

my meeting recordings, based on my, you

10:28

know, internal docs about this cohort.

10:30

As you can see right here, we've got

10:31

this information as well. And yeah, um

10:34

this isn't too bad. It's very wordy, but

10:36

it's also very accurate. So, let's

10:38

switch over now to Fables version of

10:40

this

10:41

and see what it did here. If I can find

10:43

where it caps this HTML, here it is. I'm

10:45

going to have to open this up in the

10:46

browser real quick. Okay, so here is

10:48

Fables version. It's very similar. We

10:50

have the card right here, which is a

10:51

nice touch because we do actually have

10:53

like a card designed very similar to

10:55

this. Not another AI course, the

10:56

credential the SMB market hires from. Um

11:00

we've got nine pillars, two layers each,

11:01

54 job posts analyzed, 50 founding

11:03

seats. Claim a founding seat. We can

11:05

keep scrolling down. So, they're very

11:06

similar style, right? This one has a

11:07

nice little animation here. They both

11:09

obviously pulled my logo and our brand

11:11

guidelines that you can see that they're

11:13

designed very similarly, which is great,

11:14

you know? Um we have other stuff here

11:17

like the disciplines, the pillars are

11:19

the same once again. I don't know. I

11:21

mean, they're they're obviously designed

11:22

very similarly. This is a nice touch

11:23

here.

11:24

I think the thing about this is

11:26

I probably would obviously want to

11:28

manually tweak both versions. They're

11:30

very similar. I don't think one

11:32

definitively beats the other. So, let's

11:34

look at the cost and the time. So, Fable

11:36

was 20 bucks and 50 cents, and it took

11:38

22 minutes, whereas Opus was

11:43

35 bucks and 83 cents and took almost an

11:46

hour. So, this is one of those cases

11:48

where

11:49

Fable was actually cheaper

11:50

and quicker, and arguably maybe just

11:53

like a little bit better, but it was

11:54

very similar on that side and very

11:55

subjective. So, maybe one conclusion we

11:57

can start to draw here is that Fable

11:59

still kind of wins on the creativity and

12:02

the design side compared to Opus 5,

12:05

whereas right now Opus 5 is kind of

12:06

having more of an edge for me on like

12:08

actually following directions,

12:10

verification. In most cases, it's going

12:12

to be cheaper once again. Okay, let's

12:14

keep on moving here. So,

12:16

experiment number three here was that I

12:17

wanted a LinkedIn post with a LinkedIn

12:19

carousel, and you can read the rest of

12:21

this post here, but basically I was

12:22

trying to raise the stakes, right? I was

12:23

trying to say, "Hey, you know, if I post

12:24

this out on my audience, um it would be

12:27

bad if this was,

12:28

you know, clearly AI-generated or

12:29

whatever." And I also had it choose the

12:31

topic. So, Opus, let's see what it

12:34

decided to do. Here is the actual PDF it

12:36

created for me as the carousel. I don't

12:38

love this styling, right? It's obviously

12:40

pretty consistent with our brand

12:41

guidelines, but it just looks a little

12:43

bit, you know, meh. It just doesn't look

12:45

super, super professional. So, that's

12:47

the carousel. It basically chose to

12:48

write about AI got dramatically better

12:50

at coding and trust in it went down. So,

12:53

let's see what the actual post looks

12:54

like. It probably used my LinkedIn

12:56

writing skill to do this. So, here is

12:58

the actual post. We've got some real

13:00

stats in here. We've got the arrows, a

13:02

reference to the actual carousel down

13:04

there, okay?

13:05

And if I go to Fable version now, let's

13:07

scroll up here. Aha, so Fable used a

13:09

different style of carousel, which I

13:11

actually like a lot more. It's kind of

13:12

like that tweet style. AI agents went

13:14

mainstream, trust didn't. So, that's

13:16

pretty interesting. It did similar

13:18

research, and they honestly both came to

13:20

a similar conclusion on, "Hey, you know,

13:22

based on Nate's audience and based on

13:23

what's going on in the space, what

13:24

should we write about?" which is pretty

13:25

interesting. But, ultimately, I like

13:27

this deliverable much better.

13:29

And I mean, honestly, I think that when

13:31

I read through these,

13:32

the actual content of LinkedIn posts are

13:34

pretty similar. As far as like which one

13:36

do I trust more? You know, I've got a

13:37

skill built around it. I've got a no AI

13:38

slop sort of skill as well. And this is

13:41

a perfect example of like writing a

13:43

LinkedIn post, generating that content,

13:46

I think even Opus 5 is overkill. Like

13:47

you could write really good content with

13:49

Sonnet, you know, Sonnet 4.5. So,

13:52

those are pretty similar. Let's look at

13:53

the cost and the time. Fable cost us

13:55

$6.17 and took about 7 and 1/2 minutes.

13:58

And Opus here cost us $8.22 and took

14:01

longer. So, another example where Fable

14:04

actually came in

14:05

cheaper and faster, which is quite

14:07

shocking to me because what that tells

14:09

us is that Opus is using so many more

14:11

tokens to actually be more expensive.

14:13

Because if Fable and Opus used the exact

14:15

same number of tokens, both input and

14:17

output, then Opus would be pretty much

14:19

exactly half the cost, but that's not

14:21

the case. So, when Opus comes in costing

14:23

more than Fable, it means that it was

14:25

way less token efficient as well, which

14:27

is a little bit concerning.

14:29

Okay, so let's just keep on moving here

14:30

though cuz obviously all of these

14:32

experiments I could run is not going to

14:33

be the exact same as when you use it.

14:34

And you know, at the end of the day,

14:36

it's a black box. You're pulling a lever

14:37

on a slot machine. So, let's move on to

14:39

Fable or sorry, the the fourth

14:40

experiment. So, here I told it to go to

14:42

my YouTube channel, pull comments, and

14:44

then say go to my school communities and

14:45

look through threads. And I want to

14:47

understand what my audience is saying,

14:49

what the pain points are, the number one

14:50

product that I could build to help solve

14:52

their pain points, and the number one

14:53

best YouTube video that would resonate

14:55

with them. So, let's open up the HTML

14:57

here for Opus. What your audience is

14:59

actually telling you, it pulled a bunch

15:01

of sources. It pulled things from, like

15:03

I said right here, if I can keep

15:05

scrolling up,

15:07

2,200 comments on YouTube, 480 school

15:10

posts, it looks like. And here's what

15:11

we've got. So, we've got an HTML. Once

15:14

again, I don't love this font in

15:15

general. I mean, this is on our brand

15:16

guidelines, but I might want to change

15:18

that cuz it just looks very typewritery.

15:19

It looks very cheap, honestly.

15:21

Anyways, we've got pain points that are

15:23

ranked, pricing and scoping, proving the

15:25

automation worked, and then inverse

15:26

cloud versus co-work, token, blah blah

15:28

blah. The number one product to build

15:29

would be the offer engine. So, a Cloud

15:32

Code skill pack plus templates that

15:33

takes a discovery call and then produces

15:35

a scoped priced offer with a working

15:38

measurement layer out the other. So, a

15:40

bunch of different skills. It tells us

15:41

why, and then the number one video to

15:43

make would be I sold an AI system to a

15:45

real business in 7 days. Real client,

15:47

real invoice. Okay? So, let's see if

15:49

Fable can do a similar type of

15:50

conclusion with what I could build. So,

15:52

once again, this thing looked through um

15:56

YouTube comments, as well as it looked

15:58

through school posts. It looked through

16:00

less school posts, though, which is

16:01

interesting. Let me pull up this HTML.

16:04

So, we have cost and token pricing,

16:07

error stuck mid-build, getting money, or

16:09

sorry, getting clients making money. It

16:11

goes over the YouTube mood, the school

16:13

mood, pain points, once again, which

16:16

don't seem to be the exact same. They're

16:17

similar, but, you know, they're not the

16:19

exact same. The number one product to

16:21

build would be the AI consultant kit.

16:23

So, very similar. A packaged client

16:25

delivery system that takes a member from

16:26

I can build automations to a business

16:27

paid me, not another how to build

16:28

course, blah blah blah. Okay? So, that's

16:30

pretty similar, and then I worked as an

16:32

AI consultant for real business, real

16:33

client, real numbers. Okay. So, these

16:35

are very, very similar results. So, this

16:37

would be a matter of which one do you

16:38

trust more, and maybe which one looked

16:40

through more data, and that's how you

16:42

could maybe trust it more. So, let's

16:44

look at the cost. Fable here spent

16:46

$10.60 and took 10 minutes, almost 11

16:48

minutes, whereas Opus spent $8.34 and

16:52

took 20 minutes. So, a little bit less

16:54

efficient, once again, from Opus, but

16:56

ultimately ended up being cheaper. Okay.

16:58

Let's move on to the fifth experiment.

16:59

I'm sorry if I'm going fast, but I also

17:01

don't want to bore you guys just like

17:03

really, really diving into all these,

17:04

because there's a lot of things to go

17:05

through, but I want to show you a kind

17:06

of a wide range of stuff. So, this fifth

17:09

one. Your job is to create me a YouTube

17:11

video outline and a slideshow, an

17:12

Excalidraw an Excalidraw style

17:14

presentation for this YouTube video. I

17:16

want you to go through past LinkedIn

17:17

posts, school posts, YouTube videos, my

17:19

AI's plus Q&A. So, there's a lot of

17:20

things to dig through. And then I

17:22

basically told it, "You are a product

17:24

manager. You're in charge of agents. You

17:25

don't do anything. You just delegate

17:26

work and you review stuff." So, that's

17:28

what I wanted. Okay. So, it created the

17:30

presentation and the outline. Let's

17:32

first look at the outline. So, context

17:33

engineering for agents. We have a cold

17:36

open. We then move into what changed,

17:39

why the terms exist, the failure modes

17:41

of context, and we get into writing,

17:45

selecting, compressing, isolating. Okay.

17:49

So, it a pretty legit outline, as you

17:52

can see here. Let's open up the actual

17:55

Excalidraw slide deck it made for us.

17:57

Okay. So, context engineering for AI

17:58

agents. Your AI agent isn't dumb. Your

18:00

context is. Bigger windows didn't help.

18:03

So,

18:04

harder to see. I like that little touch.

18:06

We've got Excalidraw style boxes here.

18:08

Um Andrej Karpathy, Anthropic. We've got

18:11

a quote right here. Attention is a

18:14

budget. And honestly, this doesn't look

18:16

very branded the way my other Excalidraw

18:19

uh presentations look. So, I'm not sure

18:20

exactly what happened here, but this

18:21

doesn't feel exactly right. Poisoning,

18:24

distraction, confusion, clash, and rot.

18:27

We've got some other stats here. So, not

18:29

too bad, right? I would obviously make

18:30

some tweaks before I would get ready to

18:31

start, you know, thinking about how I'm

18:33

going to present this, but not too bad,

18:34

especially for one pass. Okay. Let's go

18:37

ahead and see what Fable did here, what

18:39

kind of topic it shows for us. So, if I

18:41

go to it it created outline, slides, and

18:44

notes. So, I'm going to go to the

18:44

outline. Context engineering for AI

18:46

agents. Wow, very similar. Okay. So,

18:49

we've got the hook. We've got the

18:50

section-by-section outline. What bad

18:52

context is costing you, the four moves,

18:54

write, select, compress, isolate. Okay.

18:56

So, these are finding similar things,

18:58

which is pretty interesting. I mean,

18:59

it's looking through them assuming

19:00

similar data sources. So, that's kind of

19:02

good to know, right? Like, the

19:03

consistency makes me feel good. This

19:06

looks more like what my YouTube video

19:09

ones typically do look like, though. So,

19:11

that means maybe Fable did a better job

19:13

navigating into my other project and

19:15

finding the right skills because I

19:17

forgot to mention this directory was a

19:19

completely fresh one. Both of these are

19:20

working in completely fresh

19:21

environments, so it's not inside of my

19:23

Herc 2 as all of my normal things are

19:25

running. This one had 29 slides, so

19:27

quite a bit. We've got um a big story

19:30

here, which is something that happened

19:31

to us. We have these different colors

19:33

here. This one looks way more like what

19:35

I typically am trying to build. We've

19:37

got this nice visual with context rot.

19:39

Gets lost in the middle. You know, we've

19:41

got these nice visuals here. I would say

19:42

this one is definitely a better

19:44

presentation. So, once again Fable is

19:46

kind of coming in on top when it comes

19:47

to like the actual visual elements. I'm

19:50

assuming they both did verification

19:51

loops of screenshotting and you know,

19:53

validating. I like these a lot. These

19:55

are nice slides. Yeah. I like these

19:58

slides better, so definitely I think

19:59

Fable takes the cake here. Let's look at

20:01

cost. So, Fable took about 8 minutes and

20:03

cost us 40 bucks. Wow, almost 41 bucks.

20:06

For that Opus let's see an hour and 15

20:09

minutes and 33 bucks. So

20:12

I don't know. I think that Fable wins

20:13

here even though it was a little bit

20:14

more expensive because it was still more

20:16

token efficient and it was faster. Okay.

20:18

Let's go to the next one, number six.

20:20

So, this one's interesting. This one's

20:21

very interesting. I wanted to

20:24

try to show you guys some computer use

20:25

stuff. I compared it a little bit with

20:27

Codex computer use and ultimately I

20:29

still like Codex computer use. I don't

20:30

know why. Um they're very similar now

20:33

but I think because I just have this

20:34

bias you know

20:36

already for Codex computer use I don't

20:38

know. Anyways, I don't use computer use

20:40

but I do use it a lot for verification

20:42

stuff. So here I told it to use

20:44

Playwright CLI and I told it this is

20:46

something interesting, right? I told it

20:47

to go to Google and play the snake game.

20:49

So, I don't know if you guys you

20:49

obviously know like what the snake game

20:51

is

20:52

but if I come here, you can just play

20:53

snake right here, which I used to do in

20:55

class all the time. So, I told

20:58

these two models to go to Google and to

21:00

do that. I wanted it to only play five

21:02

games and screenshot the score of each

21:05

game and give us the average, right?

21:08

Okay, you know what's really weird? I

21:09

actually just noticed something.

21:11

So, I made a big whoopsies here. I I

21:13

sent this off as Fable,

21:15

but I only but I actually used Opus. And

21:18

then for the Opus run, where I thought I

21:19

was using Opus, I was using Opus. So, I

21:22

did Opus twice here, but I'm not even

21:24

going to change that cuz I want to show

21:25

you guys this.

21:26

Same exact prompt to Opus, right? Same

21:29

exact prompt to the same exact model

21:32

twice.

21:33

But, we got drastically different

21:34

results. In this first version, the one

21:35

that I thought was Fable,

21:37

1 hour and 53 minutes, 18 bucks. But,

21:40

look what I had to do here. You were

21:42

told to only run five games and give me

21:43

the average score there. I'm not sure

21:45

why you went rogue. This thing started

21:46

running like 30 different games and it

21:49

just kept running and kept running

21:50

games. It was trying to like

21:52

maximize itself, right? Here, it got an

21:53

average of 63. So, it was running

21:55

batches of five games at a time and did

21:57

that so many times. I was like, why why

22:00

are you not following my instructions,

22:01

but Opus is so much better,

22:03

but I didn't realize that they were both

22:04

Opus. So, that's just a really good

22:06

reminder that like, at the end of the

22:07

day, these things are completely

22:09

non-deterministic. You don't know what

22:11

they're going to do.

22:12

So, anyways, this Opus run, 2 hours, 18

22:16

dollars. And then, the real Opus run

22:18

that I thought was Opus from the

22:19

beginning,

22:20

um

22:21

1 hour, 10 minutes, and $10. But, this

22:24

Opus run did much better, which is

22:26

really weird. 81 25 33 21 54, which is

22:30

like so so much better than the other

22:32

one did. So, I'm not sure exactly what

22:34

happened there, but

22:35

anyways, hopefully that was kind of

22:38

interesting to you guys. Okay, so let's

22:39

take a look at this last one that I have

22:41

to show off to you guys today. So, this

22:43

one was um a bit longer, right? You can

22:45

read this if you want, but basically

22:46

what I wanted was a simulator where we

22:49

could see different like buildings and

22:50

vehicles and we could stress test them

22:51

with different weather. We could add

22:53

weights and we could even build and

22:54

design our own um

22:57

structures and then test them. And this

22:59

one's pretty interesting, right? Cuz I

23:00

gave this same prompt obviously to Fable

23:02

and to Opus, but let's take a look at

23:05

the two differences first,

23:06

and then how long they ran. So, I'm not

23:08

going to tell you which one's which yet.

23:11

Let's just take a look. So, here's the

23:12

first one. As you can see, there's a lot

23:14

going on. Like, it's a little bit

23:15

overwhelming, right? So, if they wanted

23:17

to build something that was kind of

23:18

user-friendly and not intimidating, they

23:19

failed on that front, but that's not

23:21

exactly what we asked for. We can see

23:23

the different nodes, we can see the

23:24

different beams, we can add like weights

23:26

to them, we can see if they're passing

23:28

or failing. I don't even know, like,

23:29

this is one of my first times opening,

23:31

you know, this. I just opened them both

23:32

to look at them, but I didn't use them.

23:34

We can see I can add like snow, right?

23:36

So, if I come here, and if we want to

23:37

add snow,

23:39

we can see how things are changing. And

23:41

if I add some rain,

23:43

um this is adding it to the whole thing,

23:44

gust factor, gravity.

23:47

You can see it starts to change colors

23:48

right there because there was way too

23:50

much weight on here. And as you move

23:52

this stuff,

23:53

you know, I'm not an engineer, so I'm

23:55

not going to come in here and tell you

23:56

this is completely structurally

23:57

accurate, and you could use this for,

23:59

um

24:00

you know, making sure that your stuff

24:02

isn't going to fall and hurt people.

24:05

This is a good simulation, right?

24:06

Because you can see where things are

24:08

being

24:09

put under pressure and where you need to

24:11

increase,

24:12

um some stability and stuff like that.

24:14

Now, I will say this is pretty, um

24:15

overwhelming. Like, this UI, I don't

24:17

really understand what to do, but if

24:19

someone did understand how to get in

24:20

here and how to test out this different

24:22

stuff and, you know, build their own

24:24

custom things, I could see this being

24:25

very useful. And it just goes to show

24:27

that in less than 2 hours, I already

24:29

have this POC that I could go get

24:31

feedback on and iterate on and stuff

24:32

like that. We've even got these

24:33

skyscrapers in here that we can start to

24:35

add a bunch of different, you know,

24:36

things to. So, anyways, this was the

24:38

first version. We've got a bunch of

24:39

different machines, and then you can

24:40

also build your own. So, let's take a

24:43

look at the other version.

24:45

This one's a bit more user-friendly,

24:46

right? And this one, honestly, does look

24:47

a little bit more AI-made. This looks

24:49

very like legacy software. This one

24:51

clearly looks a little bit more AI, but,

24:53

you know, the UI I wasn't too concerned

24:54

with. But, this one's a lot simpler. I

24:57

can see the different things. I can

24:58

easily add wind or snow or an earthquake

25:01

and I can see how much pressure is being

25:03

put on these different points. You can

25:04

see I can change to a skyscraper or a

25:06

school or a vehicle and we can do the

25:08

same thing once again.

25:09

And this one also makes it way simpler

25:11

for me if I wanted to build and design

25:12

my own thing. So if I wanted to add a

25:14

few nodes here, I could add one there, I

25:16

could have one there, one there, one

25:17

there, one there and then I can start to

25:19

connect these, right?

25:21

So I can connect these here and here and

25:23

here and there and here and there and

25:25

you know, put some triangles in here if

25:26

we really want to, you know, start

25:28

getting fancy and building some nice

25:29

support. So anyways, this one just makes

25:32

a bit more sense and I didn't prompt it

25:33

to say, "Hey, you know, like this should

25:34

be easy to use and people should

25:35

understand it and the UI should not be

25:37

intimidating or whatever." Um

25:40

but anyways, like I can add weight to

25:41

these different things and I can start

25:42

to really put some pressure on this

25:44

stuff and, you know, see what it's going

25:46

to do. This is saying that this one is

25:48

unstable. So anyways, which one do you

25:50

think was which?

25:52

Because this one was actually Fable and

25:54

this one was actually Opus, which

25:57

I wouldn't have expected, honestly.

25:58

Because I think this one is honestly

26:01

much more well-designed when you think

26:03

about what you're actually seeing in the

26:05

data. So let's take a look at how much

26:06

these costed us. So Fable this only took

26:10

Fable 7 minutes

26:12

and it costed us 73 bucks. So it was

26:14

just that that just goes to show how

26:15

quick Fable can really run your session

26:18

if you're not being careful. And then

26:19

when we go to Opus here, this one took

26:21

us 2 hours and 26 minutes and 112 bucks.

26:25

So clearly Opus got stuck in a different

26:27

loop or had for some reason Opus

26:29

interpreted the verification criteria a

26:31

little bit differently and it ran way

26:33

more sub-agents and it did more stress

26:35

testing on it. And maybe that just goes

26:37

to show this line that we talked about

26:39

that Claude Opus is much stronger at

26:40

verifying its own work and iterating

26:42

carefully until it succeeds. So maybe

26:44

that just goes to show that in action

26:46

right there. Because the thing is I

26:48

think that Claude Fable could have

26:50

easily designed something and built

26:51

something this level of detail and much

26:53

better, like far, far better. I've seen

26:55

Fable do things that are way more

26:56

incredible than this, and it could have,

26:58

but it just decided for some reason,

27:00

based on the way I prompted it or

27:01

whatever it was, it just decided to be

27:03

done. It decided to be done here, which

27:05

as you guys know if you played with

27:06

Fable, it's capable of so much more. But

27:09

I think that that's a really good

27:10

reminder of the fact that once again,

27:11

these models are not deterministic, but

27:13

also that

27:14

Opus 5 interprets things different than

27:16

Opus 4.8 and interprets things different

27:18

than Fable 5, which means the first

27:21

thing that I did when I got Opus 5 is I

27:22

ran my skills. I ran my regular

27:24

workflows. I ran my regular things that

27:26

I do, you know, generating some YouTube

27:28

stuff and helping me out because I

27:29

wanted to see how it feels. That's why I

27:31

always say I take these benchmarks with

27:33

a grain of salt because it matters way

27:35

more about how you talk to it and how it

27:37

feels and how you prompt for

27:38

verification and all of that kind of

27:40

stuff. Okay, so here are the

27:42

consolidated results from those sessions

27:44

that I just showed you guys. Keep in

27:45

mind, I accidentally ran 10 with Opus

27:47

and eight with Fable. It should have

27:48

been nine and nine, but either way, the

27:51

numbers are still pretty telling. Look

27:53

at this. Opus 5 spent more, which once

27:55

again means that it's

27:57

more inefficient,

27:59

or I could have just said less

28:00

efficient, with spending tokens because

28:03

Opus 5, per token, is half the cost of

28:06

Fable 5. So for it to be more expensive

28:08

means that it's using way more tokens.

28:10

We can see the combined output tokens

28:12

right here, 2 million for Opus and

28:14

832,000 for Fable. We can see the active

28:17

time for Opus was 630 minutes, so an

28:19

average of an hour for all of these

28:20

sessions, whereas an average of about 25

28:23

minutes for the Fable sessions. Here's

28:24

the total. Now, we can see on the output

28:26

speed, we have some stats here with Opus

28:29

and with Fable, the cost per minute with

28:30

Opus and Fable, the cost per 1,000

28:33

output tokens, Opus and Fable, API calls

28:36

per session, and tool calls per session.

28:38

And just a breakdown of where the money

28:39

went. So, it's a pretty similar split,

28:41

right? A cache reads a lot of it, and

28:43

then we have cache write, and then we

28:44

have output and input. And on both of

28:46

these and on all of these runs, the

28:48

input was pretty minimal because it was

28:49

basically just, you know, um

28:52

creating output. And if you look at all

28:54

of the sessions from most expensive to

28:56

least expensive, it's honestly pretty

28:57

split besides the fact that Opus here

29:00

owns like the most expensive one. But

29:02

anyways, I will attach this exact

29:04

document in my free school community if

29:05

you guys really want to check out this,

29:07

you know, this data

29:08

session by session. But what I really

29:10

wanted to see was this headline up here.

29:12

This headline is super interesting to me

29:13

when it comes to the output tokens and

29:15

when it comes to the time for actually

29:18

running these models. So once again,

29:20

guys, I hope that you enjoyed this

29:21

comparison here where I showed you some

29:23

actual things that I have done. Like I

29:24

said, I rec- really recommend that you

29:26

just get in here with Opus 5 and you run

29:27

your skills and just see what feels good

29:29

and what doesn't. Have it do some

29:31

verification loops and see if you like

29:32

the way that Opus feels as an

29:34

orchestrator. I still personally like

29:36

the feeling of having Fable being an

29:38

orchestrator and I like to say something

29:39

like you shouldn't be writing any code

29:42

or executing anything. You should just

29:43

be telling Opus what to do and that

29:46

saves your session limit with Fable big

29:47

time because you can be running Fable

29:49

for multiple hours or even a day and it

29:51

won't even hit like 300k or 400k on the

29:54

context because all it's doing is

29:55

delegating. So try out that tip, but

29:57

also, you know, start using Opus for

29:59

that because I think a lot of us don't

30:00

realize how powerful like a an older

30:03

Sonnet model really is that you could

30:05

use that to drive most of your daily

30:07

knowledge work. And sometimes Opus 5 and

30:09

Fable 5 are both overkill. So really

30:11

think about what you're doing and

30:13

matching the intelligence of the model

30:14

with the intelligence needed for the

30:16

task. But anyways, I hope you guys

30:17

enjoyed this video. I hope you found it

30:18

insightful and if you did, please give

30:20

it a like. It helps me out a ton. And as

30:21

always, I appreciate you guys making it

30:22

to the end of the video and I'll see you

30:24

on the next one. Thanks, everyone.

30:52

>> Mhm.

Interactive Summary

The video provides a detailed, practical comparison between Claude Opus 5 and Fable 5 through a series of real-world experiments, including coding tasks, content generation, and data analysis. While benchmarks suggest Opus 5 is highly competitive and cost-effective, the user's hands-on tests reveal that performance is highly dependent on task-specific prompting, verification strategies, and the non-deterministic nature of the models. The video concludes that Fable 5 often excels in creative and design-oriented tasks, while Opus 5 demonstrates impressive rigor in verification and complex iterative processes, ultimately advising viewers to match model capabilities with specific task requirements.

Suggested questions

3 ready-made prompts