HomeVideos

Claude Opus 4.8 actually blew my mind...

Now Playing

Claude Opus 4.8 actually blew my mind...

Transcript

350 segments

0:00

Opus 4.8 is out and it's an absolute

0:02

smash home run. They implemented so many

0:04

changes and a lot of these were hidden.

0:06

They weren't even in the announcement.

0:08

In this video, we're going to go over

0:10

every single one of those changes. I'll

0:12

tell you exactly what you need to do to

0:15

take advantage of all these changes and

0:17

then we'll even run some fun tests with

0:18

it. If you stick with me until the end,

0:20

you are going to be a master of the most

0:22

powerful technology that's out there

0:23

right now, Claude Opus 4.8. Let's go.

0:26

Here we go. We are not a channel where

0:28

we sit here for 20 minutes reading blog

0:30

posts. Let's quickly go through all the

0:32

changes then we'll get straight into the

0:34

product. First things you need to know,

0:35

it smashes all the benchmarks. It beats

0:37

ChatGPT 5.5 and all the other models in

0:41

all the benchmarks. Google's not even

0:42

close. No one else is even close. This

0:44

destroyed all the benchmarks. I actually

0:47

believe this is a kind of watered-down

0:49

version of Mythos. I'll go into that a

0:52

little bit. I'll go in a little bit, but

0:53

I actually believe this is Mythos but

0:55

kind of a little bit weaker. It's the

0:57

same cost. This is mind-blowing and I

0:59

think this is a result of all the

1:01

compute Elon Musk just gave to Claude,

1:04

but it is the same cost as Opus 4.7,

1:07

which is this is the first release in

1:10

quite a bit of time where the price

1:12

didn't go up. For both OpenAI and

1:14

Anthropic, all their new releases, the

1:16

price ticked up slowly and slowly and

1:18

slowly. This is the first one in a while

1:20

where the cost did not go up, which

1:23

again, just a few weeks ago, Elon sold

1:26

to Anthropic tens of billions of dollars

1:28

of compute. I think that is a direct

1:30

result of that deal, which is really

1:32

amazing. Here's a big one and also a big

1:34

reason I think Elon's helped out

1:36

Anthropic a lot. Their fast mode is

1:39

cheaper. A big reason I have been using

1:43

ChatGPT 5.5 inside Codex more the last

1:46

couple weeks is their fast mode is dirt

1:49

cheap. You get way better performance

1:51

for not that much more money. Claude,

1:53

their fast mode was six times more

1:56

expensive, which is untenable. Their

1:58

limits were already low, so their fast

2:00

mode wasn't worth it. Now their fast

2:03

mode is three times cheaper than it was

2:04

before, which if we do the advanced

2:06

algebra here, comes out to just two

2:09

times more expensive than the regular

2:11

mode, if my math is correct there. So

2:13

their {slash} fast mode is actually

2:16

affordable if you're on the $200 plan,

2:18

and we'll go into exact recommendations

2:20

right after this, so stick around. Beats

2:22

Claude GPT 5.5. I thought 5.5 was the

2:25

first model ever that beat Opus at

2:28

coding, but this one came right back and

2:30

beat it. Four times less hallucinations.

2:34

So this is a big one, and this is a big

2:36

reason why I think this is actually just

2:37

Mythos, but watered down a little bit,

2:39

is this is one of the things they showed

2:42

off with Mythos was the reduction in

2:44

hallucinations, about four times

2:46

reductions. It's matching Mythos in a

2:47

lot of things they advertised before.

2:49

And for those who are new to the

2:50

channel, kind of new to the AI world,

2:52

Mythos is this model that Claude has

2:55

been advertising now for a couple

2:56

months. They've been advertising it some

2:58

sort of doomsday model that's

3:00

outrageously good that can hack any

3:01

website. They've been teasing us a bit.

3:04

It appears we're getting closer and

3:05

closer. One big thing to notice also, in

3:08

their announcement blog post, this

3:10

wasn't in the tweet as well, they expect

3:12

to bring Mythos class models to all the

3:15

customers in coming weeks, which means

3:17

they're doing it. They're actually going

3:19

to release Mythos. Again, I think Elon

3:22

must saved Anthropic from the dead here.

3:24

They were starting to lose in every

3:27

single facet because their lack of

3:29

compute, and now they're going to

3:30

release Mythos. I don't think it's any

3:32

coincidence that they're increasing

3:34

limits, making prices cheaper, and

3:36

releasing super powerful models within

3:38

weeks of buying tens of billions of

3:40

dollars of compute from Elon Musk. Now

3:42

from a new functionality perspective,

3:45

here's the two big things, and we're

3:47

going to demo this when we get into the

3:48

product. Dynamic workflows in Ultra

3:51

Code. What are these two things? Why are

3:53

they so big and important? Dynamic work

3:56

flows is now Claude code's ability to

4:00

tackle months of work in just a day.

4:04

What does that mean? If you give Opus 48

4:07

a very complex task, a big juicy meaty

4:10

girthy complex task, it will now spin up

4:14

between tens to thousands of sub agents

4:18

to tackle that task. Say you were trying

4:21

to implement a really big new feature or

4:24

one shot a big app. Before, if you gave

4:28

it to Opus, it would just have one agent

4:30

go there and add some code, take some

4:32

code away, add some code, do some

4:33

research, add some code. Now it's going

4:35

to take literally thousands of those

4:37

agents, send them out. They're all going

4:39

to be touching different pieces of your

4:41

code base, doing research, testing

4:43

things out. And simultaneously, these

4:46

tens of thousands of agents are going to

4:48

be writing code, testing, using the app,

4:51

doing regression tests, a whole bunch of

4:54

things. It's going to be really really

4:56

powerful, and that is now in Opus 48.

4:58

Again, this is going to allow you to do

5:00

months of work in just one afternoon.

5:03

And then you have ultra-code mode, which

5:05

is basically giving the keys of the

5:07

kingdom to Claude code, to Opus 48, and

5:10

saying, "Hey, use dynamic work flows

5:12

whenever you want." This is only You're

5:14

only using this if you got that $200 a

5:16

month plan. So this is Opus 48. I'm

5:18

going to go into my exact

5:19

recommendations of how to use it in a

5:21

second. Then we'll go into the product

5:23

and demo it out. But again, absolutely

5:25

massive changes here. Let's go into the

5:28

recommendations. Number one, switch all

5:29

tasks to Opus 48. There's no reason not

5:31

to. There's no reason not to go in to

5:34

Claude code right now, pull it open, and

5:37

choose Opus 48. Now you can choose the

5:39

million context if you want. I find the

5:42

million context you don't absolutely

5:44

have to use it. I find once you start to

5:46

fill up that million context window, the

5:49

performance actually degrades a good

5:51

amount. So, I'm actually an Opus 4 8,

5:53

which is the regular context type of

5:54

guy. From a effort perspective, I'm

5:57

recommending doing a high by default,

6:00

and then when you are building out much

6:02

bigger things, switching to extra or

6:05

max. But, I would by default stay in

6:07

high, and then only switch to extra and

6:09

max when you have to. Despite the fact

6:11

that Papa Elon allowed Anthropic to have

6:13

much more compute and capacity, it's

6:15

still not as high capacity as ChatGPT.

6:19

So, I'm sticking with high for default,

6:20

then doing extra and max if necessary.

6:22

When it comes to Hermes and open claw, I

6:25

wouldn't move it to Opus 4 8 just yet.

6:28

This is a big mistake a lot of people

6:30

make is they try to force their agents

6:32

into using latest version Opus the

6:34

moment it comes out. The issue is this

6:36

invariably leads to errors, leads to

6:38

crashes, and errors and crashes in open

6:41

claw and Hermes are not the most fun to

6:42

solve. I would wait until the official

6:45

releases, which typically come within 24

6:48

hours of the release of the model. So,

6:50

once it officially releases, then you

6:52

switch it over, and you'll have way less

6:53

crashes and and way less bad

6:56

reliability. As for the slash fast mode

6:58

and the new ultra code mode, I'm only

7:01

using those if you're on the $200 a

7:02

month plan. And even if you're on the

7:04

$200 a month plan, I don't know if I'm

7:07

using it for every single prompt. Like

7:09

ChatGPT Codex, I'm using fast mode for

7:12

literally everything cuz they give you

7:13

so much capacity. Claude, their limits

7:15

still aren't as high as ChatGPT. So, for

7:18

me, I actually have extra usage on,

7:20

which means once I get past limits, I

7:22

just pay through the API. So, I'm

7:24

actually going to be using fast and

7:25

ultra code for almost everything. But,

7:27

for you, if you don't have the extra

7:29

capacity or if you're not on $200 a

7:30

month plan, then I wouldn't use these

7:32

modes. Totally up to you. And here's the

7:34

last recommendation before we get into

7:35

products and we do some cool demos. You

7:38

need to lock the hell in. You need to

7:39

lock the hell in. I've been working with

7:41

a lot of people lately, watching how

7:43

they vibe code, seeing what they do. One

7:45

issue I'm seeing is AI is enabling a lot

7:48

of people to get wildly distracted. They

7:51

will send a prompt to their AI, and then

7:53

they will go and doom scroll for an

7:55

hour, despite the fact that their AI

7:57

finished the task like 50 minutes

7:59

earlier. You cannot get distracted. If

8:01

you can get into a flow state and lock

8:04

the F in, you are going to get so much

8:07

more done. I truly believe the number

8:10

one indicator of how successful someone

8:13

will be in 2026 is their level of focus.

8:17

Do not allow this extra power to mean

8:20

you can slack off more. Use this extra

8:22

power to get more done. So, really work

8:25

on your focus, put the phone away, close

8:27

social media, close Twitter, close

8:29

YouTube, and just lock in, and you'll

8:31

get so much more out of this tech. Now,

8:33

let's jump into the product and build

8:34

some cool things out. I'm using Claude

8:36

Code Desktop. You can use the CLI or the

8:39

extension to take advantage of Opus 4.8

8:42

right now. I'm going to run one of the

8:44

world famous Alex Finn benchmarks on

8:46

this model. This is a benchmark I've ran

8:48

on every single model. Up until now,

8:50

Opus 4.7's actually been king with by

8:52

far the best scores in all four of these

8:55

tests. We're going to run the 3D

8:57

first-person shooter test here, see how

8:59

it does, see how it compares to the

9:01

other models. If you want to run this

9:02

benchmark yourself, I'll put the uh

9:05

prompt for this down below, so you can

9:07

run your own world famous Alex Finn

9:09

benchmark. I'm going to hit enter on

9:11

this, and I'm going to send it off, and

9:12

we're going to see how it does.

9:13

Basically, what we're going to have it

9:14

do is we're giving it creative freedom.

9:16

We're saying build a 3D first-person

9:17

shooter using 3.js. Do whatever you

9:20

want. Make it as creative as humanly

9:22

possible. Add power-ups. Do whatever you

9:24

want. We'll see how good Opus does here.

9:27

Side note, remote control is active. I

9:29

There's actually a setting in Claude

9:31

Code not many people know about. You

9:33

should be using the setting. It turns

9:35

remote control on by default for every

9:38

single chat. What this allows you to do

9:40

is whenever you spin up a new chat in

9:42

Claude, you can actually go on your

9:44

phone, go into the code section in the

9:48

top left, and as you can see here, that

9:51

chat I just started is now on the

9:53

screen. Create the stylistic 3D

9:54

first-person shooter. So, I can now go

9:56

mobile whenever I want with every single

9:58

chat I start. So, make sure to turn that

10:00

on. A little tip for you there. A little

10:01

bonus tip. Go in the settings, turn on

10:03

remote control is active. The only thing

10:05

I ask for for that tip is you tip me

10:07

with a like down below, subscribe if you

10:10

learned anything so far, turn on

10:12

notifications, and I'm going to do a

10:14

full boot camp by an Opus 48 tomorrow in

10:17

the Vibe Coding Academy. Make sure to

10:19

join that number one AI community on

10:20

planet Earth. Link down below. Best

10:22

decision you'll ever make in your entire

10:24

life. All right, looks like it's done.

10:25

It even tested itself, which is sick.

10:27

Let's see how this is. Neon Assault.

10:29

It's always neon themed. I have no idea

10:31

why. First model that makes a non-neon

10:33

themed game, I'm going to give it a 10

10:35

out of 10. Here we go.

10:36

Let's engage. This is nice. This is

10:39

nice. These graphics are very, very

10:41

nice. Much better I mean, if you're this

10:43

is your first time watching my channel,

10:44

you might think, "What the hell is this

10:45

guy talking about? This sucks. This

10:46

isn't Cyberpunk 2027." But, if you

10:49

compare this to the default apps that

10:51

previous models have built, this is

10:53

pretty nice with from the walls to the

10:55

ground. These are the enemies. Oh, to

10:58

the way the gun shoots, to the way you

11:00

can see hit markers on the enemies. I

11:02

assume these are Even the power-ups look

11:03

nicer.

11:08

Wave two. So, they got combos, they got

11:11

waves.

11:12

This is for sure an upgrade and probably

11:15

the best version of this we've seen yet.

11:18

Oh, this is an enemy. Okay.

11:21

This is probably a step above what 47

11:24

gave to me. Probably just a small step.

11:26

So, I'm going to give it a 9.1. I'm

11:28

going to run the next three benchmarks

11:30

probably on a live stream in the next

11:32

week. If you want to see that, make sure

11:34

to turn on notifications down below for

11:36

that. Again, here's a reminder on my

11:38

recommendations. You want to be jumping

11:39

on this now. When they release new

11:42

technology, you have a distinct

11:45

advantage if you start using it right

11:47

away. Your competition probably isn't

11:49

using Opus 4. They're probably not using

11:52

the dynamic mode that sends out tens of

11:54

thousands of sub agents. They're

11:55

probably not using that. So, if you go

11:57

and you use this tech and you build out

12:00

really, really cool things, you are

12:02

going to have a distinct advantage over

12:04

the rest of the field. So, you want to

12:06

make sure today, carve off some time in

12:08

your calendar, go on do not disturb

12:10

mode, close out all the doom scrolls you

12:13

got, the tickety talks, the Twitters,

12:15

all of that, and lock in and use this

12:17

and build cool things cuz you have an

12:18

advantage right now over everyone else

12:21

if you take advantage of all these

12:23

different features and functionality

12:24

they just released. Let me know what you

12:26

want next about Claude. Do you want

12:29

tutorials and I build really complex

12:30

apps? Do you want deep dives in a

12:31

functionality? Do you want more

12:33

benchmarking to see if it's the best?

12:35

Let me know down in the comments. I'm

12:36

super curious what you want. All my

12:38

videos are based on your feedback. I

12:40

hope this is helpful. See you in the

12:41

next video.

Interactive Summary

This video covers the major release of Claude Opus 4.8, highlighting its impressive performance in benchmarks, the introduction of 'Dynamic Workflows' and 'Ultra Code' modes, and the overall improved cost-efficiency. The presenter also shares practical recommendations on how to integrate this model into your workflow, emphasizes the importance of focused 'vibe coding,' and demonstrates the model's capabilities with a 3D first-person shooter benchmark.

Suggested questions

4 ready-made prompts