HomeVideos

Claude Opus 5 DESTROYS Fable 5

Now Playing

Claude Opus 5 DESTROYS Fable 5

Transcript

293 segments

0:00

Claude Opus 5 has dropped. It has

0:02

happened and it has totally blown my

0:05

mind. It is better than Claude Fable 5.

0:08

It is a fraction of the price. It is

0:11

significantly faster, but it has three

0:14

insane weaknesses I'm about to go over

0:17

that change absolutely everything. But

0:19

before we get into those three deal

0:20

breakers with Claude Opus 5, let's talk

0:22

about the incredible things that are

0:24

like legitimately revolutionary. First

0:26

of all, it beats Fable 5 on almost every

0:29

single benchmark. This is not one of

0:30

those channels where we look at charts

0:32

and benchmarks all day. Totally boring.

0:33

No, I'm just letting you know right now

0:35

it beats Fable 5 on every benchmark and

0:37

in a second I'm going to take you

0:39

through a world-famous new revolutionary

0:42

benchmark I've been working on for weeks

0:44

now that proves it's better than Fable 5

0:47

in almost every single way. It is half

0:50

the price. So this was the biggest issue

0:51

with Fable 5. It was totally

0:53

uneconomical for the average user, but

0:56

Opus 5 it's half the price. It comes in

0:59

at a good price and you get your full

1:01

limitations with your Claude plan.

1:03

Meaning you can use it to 100%. Claude

1:05

Fable 5 for some reason won't let you

1:06

use half your limitations, which is

1:08

really stupid and annoying. It's

1:09

significantly faster. It works lightning

1:11

quick, much better than Fable. Which

1:13

brings us to the question and I'll and

1:15

I'll show you my benchmarks in 1 second

1:17

here. Any reason to use Fable 5 anymore?

1:21

No, with one small exception which I'll

1:24

go through right after these benchmarks

1:26

I'm about to show you. But no, there

1:28

isn't a reason. And then as I said,

1:30

there's three major weaknesses which

1:33

we'll go over as well. Let's go into

1:34

these world-famous Finn benchmarks I

1:37

just created to show you why this is

1:38

better than Fable 5. All right, so this

1:40

is the new famous Finn benchmark. There

1:43

are five tests here we put both Fable 5

1:46

and Opus 5 through. In a second I'm

1:48

going to show you Opus 5 versus GPT 56

1:51

so we can know which one of those two is

1:53

better. And Opus 5 won basically all of

1:57

the benchmarks. So, starting with

1:59

benchmark number one, this was a 3D

2:02

roller coaster simulator. Both had to

2:04

build a 3D roller coaster simulator.

2:06

Let's show you Opus 5 first. This was

2:09

Opus 5, an absolutely beautiful

2:11

simulation. It built the entire track,

2:13

all the trees, the building route, the

2:14

sky, the clouds. And if I hit ride, you

2:17

can actually see the roller coaster in

2:20

real time. You can see the from the

2:23

user's perspective going around the

2:25

roller coaster. This is really, really

2:27

nice, detailed, beautiful, and was done

2:29

lightning quick. As you can also see, it

2:31

was done with 82 cents of credits and

2:34

about 32,000 tokens. We go over to Fable

2:37

5. It was actually done with less

2:40

tokens, but was significantly more

2:42

expensive, about 50% more expensive.

2:45

Let's see the output of it. And as you

2:47

can see, it's just not quite as

2:48

beautiful, not quite as detailed or

2:50

clear. It looks like just kind of like a

2:52

generation behind. If we go on the ride,

2:55

you can see the track doesn't look quite

2:57

right. The trees all kind of look

2:58

similar. There was no real detail to

3:01

anything, so it doesn't look nearly as

3:03

good. The next test is the pixel perfect

3:05

test where it has to clone an entire

3:07

website. It has to clone the Apple

3:10

website. If we go to the actual Apple

3:13

website, this is what it was tasked with

3:15

cloning. You can see college sorted. It

3:17

has a bunch of people holding Apple

3:19

devices, iPhones, MacBook Airs, MacBook

3:22

Pros. Now to Opus 5's recreation,

3:25

obviously, it doesn't look incredible,

3:27

but I'll show you Fable in a second. But

3:28

you can see it's pretty close. It had to

3:31

build every part of this website from

3:33

scratch. It's not copying and pasting

3:35

over images. It's building what a

3:37

MacBook looks like from scratch. What a

3:39

MacBook Pro, an iPad, a watch would look

3:41

like. As you can see, it has names,

3:44

recreates a TV. Let's see what this

3:46

looks like with Fable. Fable 5, not

3:48

quite the same. Remember this looked

3:50

like human beings on the real website

3:52

and in the Opus 5 website, the iPhones

3:55

cut off, the MacBook looks nothing like

3:57

a MacBook Air, the iPad looks nothing

3:59

like iPads, the watch doesn't look

4:01

anything like a watch. The cards kind of

4:04

lazily done. As you can see, nothing

4:06

looks quite the same. Opus 5, everything

4:09

looked much better. Does it look

4:10

perfect? No, but looks much better than

4:12

Fable. The next test was the gauntlet

4:15

test. And basically, this is a agentic

4:19

test where it tests the tool use of the

4:21

model. We give the model a whole bunch

4:23

of documents, PDFs, Excel spreadsheets,

4:26

a whole bunch of things, and have it

4:28

basically do a scavenger hunt where it

4:30

goes and has to find specific things in

4:33

all the documents. It tests its agentic

4:35

ability. Opus 5, it did a pretty good

4:38

job. It got five out of the eight

4:40

scavenger hunt items. Fable 5 cut off

4:43

halfway through because of content

4:46

blockage from Anthropic. It thought I

4:47

was doing something with cybersecurity.

4:49

It cut it off. This has nothing to do

4:51

with cybersecurity. It's about finding

4:53

specific things in different documents.

4:54

Fable 5 wouldn't allow me to do the

4:56

agentic test, so that didn't count.

4:59

There's a debug wall. So basically, what

5:01

this benchmark does is go online and

5:03

find like 15 different bugs from

5:06

open-source GitHub repos. And then it

5:08

hands it to each model and says, "Hey,

5:10

go through this and fix all the bugs in

5:12

all these open-source repos." Opus

5:14

actually took a bit longer than Fable 5.

5:18

Did it at about the same amount of

5:20

tokens, but did it at about a dollar

5:22

cheaper overall. So about 25% cheaper.

5:26

So this goes to Opus 5 as well. And then

5:29

the last test is breaking point. And

5:31

basically, the way this works is each

5:34

model is tasked with building a bridge.

5:36

It's basically a bridge simulator.

5:38

They're tasked with building a bridge,

5:40

and then the benchmark drives a car over

5:43

the bridge over and over and over and

5:45

over see how much weight the bridge can

5:47

hold. It's basically testing the

5:49

thinking ability. Okay, can you design a

5:51

bridge that holds tons and tons of

5:52

weight? Fable cost basically double Opus

5:55

to do this, but only was able to hold

5:57

slightly more weight than Opus. So, it

5:59

goes to Fable, but it was a lot more

6:01

expensive. Overall, Opus 5 beat Fable

6:04

beat him in almost every single

6:06

benchmark, did it for significantly

6:08

cheaper. Total cost was $6 for Opus,

6:11

$7.50 for Fable 5. And it's the winner.

6:15

Opus beats Fable for a fraction of the

6:17

price. So, let's talk about the

6:19

weaknesses of the model. There are a few

6:21

deal breakers for me here that are

6:23

stopping me from using this in my entire

6:26

stack. Number one, the personality

6:28

sucks. It absolutely sucks. I've never

6:31

been so annoyed talking to a Claude

6:33

model. This has actually been the

6:34

advantage of Claude models up to this

6:37

point. I've always loved talking to

6:39

Claude models. It's been their biggest

6:41

advantage against ChatGPT, but for the

6:43

first time it has flipped. I loathe the

6:45

personality of Opus 5. It is way too

6:48

verbose. It is not nearly concise

6:50

enough. It goes in a hundred different

6:52

directions when it's talking to you. You

6:54

ever have like that friend from high

6:56

school who thinks he's just like way

6:57

better than everyone else and way

6:59

smarter than everyone else? And when you

7:00

talk to them, they use like the biggest

7:02

words possible and go in a million

7:04

different directions to prove how smart

7:06

they are? That's what it feels like with

7:08

Opus 5. I've had to multiple times using

7:11

Opus 5 hit the stop button to get it to

7:13

shut up and I say "Please be way more

7:15

simple and concise and talk to me like

7:17

I'm 5 years old." I'm not kidding. For

7:19

the first time I've had to go to

7:21

Claude.md to edit its personality. I

7:23

just said, "Speak as simple as humanly

7:26

possible." And I highly recommend when

7:27

you use this model you do the same

7:29

thing. It's unfortunate cuz it also kind

7:31

of leaks into the way it works

7:32

sometimes, where I'll be like, "Fix this

7:34

bug." And it'll just do a hundred other

7:36

things before fixing the bug, which is

7:39

really, really annoying. It It appears

7:41

like it's just this like erratic, super

7:44

hyper intelligent being that can't stay

7:46

focused. For me, Fable 5 was actually

7:49

way more focused. And I actually enjoyed

7:52

talking to Fable 5 more. The issue is

7:55

Fable 5, you can only use 50% of your

7:57

budget on it, and it uses up all your

7:59

credits. So, I have to replace Fable 5

8:02

with Opus. So, highly recommend editing

8:05

your personality for Opus 5. It's just

8:06

It's just too much and it does too much.

8:09

The limits suck. Even though you can use

8:10

100% of your budget on Opus 5, the

8:13

Claude limits just absolutely suck

8:15

compared to ChatGPT. ChatGPT, that that

8:18

Tibo dude from Twitter is constantly

8:20

restarting the limits like every 5

8:23

minutes. You basically get unlimited

8:25

usage with ChatGPT. Claude, even though

8:28

you can use all your budget on Opus 5,

8:30

it still has lower budgets overall

8:34

Anthropic versus ChatGPT. I can still

8:36

see the meter going quicker, which gives

8:39

me like a level of anxiety as I'm giving

8:41

prompts. It makes me want to do less.

8:43

Because ChatGPT has unlimited usage

8:45

basically,

8:47

there's no anxiety when using I I'm more

8:49

free to be creative and explore and do

8:51

more things and do interesting things.

8:54

So, I you know, the limits still here

8:55

are a deal breaker for me. And then the

8:57

harness for Claude code is still just

9:00

not as good as Codex or I guess it's the

9:02

ChatGPT app now. The ChatGPT app is a

9:05

significantly better harness. Their new

9:07

voice mode, which video on that coming

9:09

in like the next 24 hours, maybe 48

9:12

hours, turn on notifications now and

9:14

subscribe, especially if this video has

9:16

been helpful for you. Video on that

9:18

coming very soon, but it is incredible.

9:20

It is excellent. Claude just added a

9:22

voice mode, but it's not nearly even

9:24

like a quarter of what the ChatGPT voice

9:26

mode is. That's coming soon again,

9:28

notifications on. Also, by the way, I'm

9:30

doing a boot camp on Opus 5 in an hour

9:33

from me filming this. It's going to be

9:35

recorded. It'll be in the Vibe Coding

9:37

Academy. Link for that down below.

9:39

Number one AI community on planet Earth.

9:41

Join that link down below. I promise

9:43

it'll be the best decision you ever

9:44

make. Now, here is my new stack. With

9:46

all that being said, here is my new

9:48

stack for super hard problems, Opus 5.

9:51

It's the smartest model out there. It

9:52

has the highest intelligence. It's

9:54

smarter than Fable 5. Fable 5 was

9:56

slightly smarter than 5.6. Opus 5 is

9:59

slightly smarter than Fable 5. If I'm

10:00

doing massive planning, I'm still

10:02

relying on Fable 5, mostly because I

10:04

just don't like the output of Opus from

10:06

like a talking perspective. So, if I

10:09

need talking, if I need a plan, if I

10:11

need to go back and forth, if I need a

10:12

brainstorm, I'd rather do with Fable 5.

10:15

I don't want to talk to Opus 5 to do

10:16

planning. I just want to shut up and

10:18

write code. Daily driver though, that's

10:20

ChatGPT 5.6. You get so much higher

10:22

limits. The voice mode is incredible.

10:24

Again, video coming soon. It is just

10:26

better to use overall out of the three.

10:30

So, daily driver, ChatGPT 5.6 it is.

10:33

Have you used Opus 5? How's it compare

10:35

to Fable 5 for you? Let me know down

10:37

below in the comments section. Hope this

10:39

was helpful. Way more videos coming out

10:41

on Opus 5, Claude Code, GPT voice mode,

10:44

Hermes agent, Opus 5 and Hermes agent,

10:46

all coming very soon. Make sure to

10:48

subscribe. Leave a like if you learned

10:50

anything at all. I'll see you in the

10:51

next video.

Interactive Summary

Claude Opus 5 is a newly released model that outperforms Claude Fable 5 across most benchmarks while being more cost-effective and faster. Despite these performance advantages, the creator highlights three significant weaknesses: an overly verbose and arrogant personality, restrictive usage limits that cause user anxiety compared to ChatGPT, and a less developed toolset/voice mode than its competitors. Consequently, while Opus 5 is recommended for high-complexity coding tasks, Fable 5 remains preferred for brainstorming and planning, and ChatGPT remains the recommended daily driver due to better limits and features.

Suggested questions

3 ready-made prompts