HomeVideos

Grok 4.6 is Claude Fable 5, but dirt cheap

Now Playing

Grok 4.6 is Claude Fable 5, but dirt cheap

Transcript

352 segments

0:00

Grok 4.6 just dropped, and it's

0:02

official. SpaceX AI is now a frontier AI

0:06

company. Elon claims you can now get

0:08

above Chad GPT 5.6 and Opus 5 quality

0:13

for a fraction of the price and much,

0:16

much faster. Well, I put it to the test,

0:18

and today you are going to find out if

0:21

those claims are true. You're also going

0:23

to find out if you should be switching

0:24

to Grok 4.6, how to use Grok 4.6, and

0:28

which situations you should be using the

0:30

model. Now, let's lock in and get into

0:32

it. So, here are the benchmarks they're

0:34

giving. Unfortunately, I think all

0:35

benchmarks are fake, so we're not going

0:36

to spend much time here. Basically, what

0:38

they're claiming though is it's just as

0:40

good as Opus 5 and 5.6 sole for a

0:44

fraction of the price. They're not

0:45

claiming it's way better or it's the

0:47

best model ever made. They're saying

0:49

it's good as frontier, but you're

0:50

getting it for way better costs, which

0:53

is a good thing because Fable and Opus

0:56

and 5.6 sole are very, very, very

0:59

expensive models. The two places you can

1:01

be using it right now are Cursor and

1:03

Grok Build. I lean towards using Cursor.

1:06

It is a much more full-featured AI vibe

1:10

coding experience. It's basically kind

1:12

of like a stripped-down version of the

1:13

Chad GPT app, but it's very, very good.

1:16

Grok Build is solid, but it's just a

1:18

CLI. There's no interesting harness

1:21

features I think you'd come to expect

1:23

from harnesses in August 2026. So, I

1:27

lean much more toward using Cursor

1:29

instead of Grok Build if you want to use

1:30

this model. Now, let's talk about if

1:33

they came through on their claims. I ran

1:36

this model through the world-famous Alex

1:38

Finn benchmarking test. For those who

1:40

don't know, it is five different

1:42

benchmarks. I put it against Chad GPT

1:44

5.6 sole. I put it against Fable. Here

1:47

are the results. It beat 5.6 sole in my

1:50

benchmark. So, it went through these

1:52

five different tests. It did a whole

1:54

scavenger hunt for code on the internet.

1:56

It did a whole debugging exercise, it

1:59

built a bunch of simulators. Basically,

2:01

what it came down to is it scored better

2:03

than 5 6 soul. It did it in

2:06

significantly less time, did it in 18

2:08

minutes rather than 24 minutes with GPT,

2:11

and it did it for basically a third of

2:13

the price. Now, let's go through those

2:15

actual results here. Well, starting off

2:16

with the first one, which is a roller

2:18

coaster simulator test. Over on the

2:20

right is 5 6, over on the left is Grok.

2:23

Let's open this up and see what it looks

2:25

like. As you can see, it built a very

2:27

nice roller coaster simulator. That is

2:28

That is the I think the biggest roller

2:30

coaster out of all the simulators we've

2:31

done so far. Let's hit ride and see what

2:33

this is like. This looks great. The sky

2:35

looks amazing. I love the kind of

2:37

twilight colors in the sky. Clouds look

2:40

nice, and as you can see, the roller

2:41

coaster is looking good. If we go over

2:44

to chat GPT and I expand this, I'd say

2:47

the fidelity is not quite as good, it's

2:48

not quite as pleasant on the eyes. And

2:50

if we go ride here,

2:53

uh it looks it looks fine. It looks all

2:55

right. It did it for a fraction of the

2:57

price. It actually took a little bit

2:58

longer on this test, but I think the

3:00

results came out well, and you spent way

3:02

less money on it. Here are a couple of

3:04

the other tests against chat GPT, a

3:06

recreation of the Apple website. So, we

3:09

hand both models the Apple website, we

3:12

say, "Hey, recreate it as close to

3:13

graphical fidelity as you possibly can."

3:16

This is what the Apple website looked

3:18

like today. As you can see, there's the

3:19

phone, the air, uh all the other

3:21

devices, Ted Lasso, uh Sabrina

3:24

Carpenter. Let's see the recreation from

3:26

Grok. Not too bad, not too bad.

3:29

Obviously, it doesn't look nearly as

3:31

good, but it can't use graphic

3:32

generation, it has to recreate all the

3:34

graphics itself. Honestly, not too bad.

3:37

And then if we go over to GPT, GPT looks

3:41

all right. I like the header better, but

3:43

the phones look all kind of messed up

3:45

from there. That doesn't even look like

3:46

a laptop. Uh all the things down here

3:49

don't look too great. So, honestly, it

3:51

it it edged it out. Did it actually a

3:54

little bit longer, too, but much

3:55

cheaper. Other than that, we have an

3:57

agent test where we give it tons and

3:59

tons of files. Then we basically have it

4:00

go on a scavenger hunt, has to use

4:02

different tools to find things in CSVs

4:04

and PDFs and different files. It

4:07

actually dominated GPT. It did it in a

4:10

looks like a sixth of the time for about

4:14

a sixth of the cost, and it found and it

4:17

used the tools better. GPT couldn't even

4:20

find uh two of the scavenger hunts. Bug

4:22

fixing, we handed a open-source library

4:25

from GitHub, a a recent one. We see how

4:27

many bugs it can find. Both of them

4:29

found all 13 bugs in the code. Grok did

4:32

it faster, did it for about half the

4:34

cost. GPT took about a minute longer for

4:36

about a dollar more. And overall, it

4:39

came out that Grok beat GPT pretty

4:42

handily, which is pretty shocking. But

4:44

how did it handle against Fable? Let's

4:46

check that out. So, Grok actually ends

4:49

it up beating Fable 5. Now, I will say

4:51

this, the reason why Grok won is because

4:54

Fable 5 refused to do the scavenger hunt

4:57

because of its safety guardrails. If you

4:59

take out Grok's score from here, it is

5:02

about even when you take out Grok's

5:04

score from that one. But at the same

5:06

time, Fable's not letting you do

5:08

everything. Grok lets you do everything.

5:11

Grok beat it at about a tenth of the

5:13

cost, did it in a minute quicker. Let's

5:15

check out Fable beat it on pixel

5:17

perfect, the Apple recreation. Let's see

5:19

that real quick. Here is the Fable

5:21

recreation. Some of these actually look

5:23

really, really strong, some of the

5:25

devices. Fable actually did this one

5:27

faster, but again, four times the price.

5:30

Bug finding, Fable was four times the

5:32

cost at double the time. And so, Grok

5:35

ended up beating Fable 5, which is

5:37

really, really impressive. Now, I did

5:39

some other tests, too. I put it up

5:40

against Opus, Fable, and 5.6 on some of

5:43

these tests. So, here's a flight

5:45

simulator. I also put Grok against the

5:47

test with Opus, as well as Fable and

5:50

Soul and a few other tests here. So,

5:52

first we have kind of a Flappy Birds

5:54

test here. Here is the Grok one.

5:57

Actually looks really, really nice.

5:59

Ooh, next levels. You get the stars, you

6:01

get everything. I like that. It does It

6:03

runs the simulation a little bit quick

6:06

there. Oh, and things are just

6:07

collapsing. All right, so the physics

6:08

are a little bit off. Let's see here.

6:10

Graphics looks great though. Let's see

6:12

what Opus 5 looks like here. Opus 5

6:15

looks pretty good, too.

6:17

Physics actually work well. I like that

6:19

a lot. There doesn't appear to be second

6:21

levels though. You just do level Oh,

6:23

there we go. Next level. Here we go. All

6:25

right, so now they got new things out of

6:26

there. Physics are a little off, too, as

6:28

well. Just see it collapse before you'd

6:30

even do anything. I'd say graphically,

6:32

Grok probably outdoes it a little bit.

6:34

Gameplay-wise, Opus wins. And then let's

6:37

check out Fable 5. Fable 5 appears to be

6:39

completely broken. And then ChatGPT 5 6

6:43

Soul. This looks really good, too. I

6:45

like the graphics. Let's Let's fling

6:47

this thing. Uh the physics seem off here

6:50

and everything looks kind of messed up.

6:51

This is probably the weakest of them

6:53

all. So, let's see how much it cost, how

6:54

long it took. 9 minutes for Grok. It was

6:57

actually longer than 5 6. Yet it is half

7:00

the price of 5 6. By far the cheapest

7:03

out of all of them. And I think it got

7:04

the best results. Fable produced the

7:06

worst results and was the most expensive

7:08

and took the longest. So, let's revisit

7:11

those claims. Let's see what was good

7:13

and bad about Grok 4 6. I've been using

7:15

all day. I've tested it in Cursor, Grok

7:17

build. I showed you all my benchmarks.

7:19

Here's what it comes down to. The

7:21

positives. All the claims are true. It

7:23

is basically just as good as Opus and

7:27

ChatGPT 5 6 while being much cheaper and

7:31

much faster. Those are all 100% true.

7:36

Here's the issue though. The issue is

7:38

there's no kind of great general purpose

7:41

harness for this model. The strength

7:45

right now of ChatGPT is the ChatGPT

7:47

desktop harness is incredible. It has

7:51

everything. It does coding really well,

7:54

kind of like Cursor, but it also does

7:55

the general-purpose stuff. It can do

7:57

computer use really, really well. It can

7:59

do browser use really, really well. It

8:02

does everything you need to get all your

8:05

knowledge work done in one single place.

8:08

Grok doesn't really have that. Grok has

8:11

Cursor. Cursor is good. Cursor is a very

8:14

good app, but it is really built for

8:18

coding. Yes, it can do some more

8:19

general-purpose stuff, but it's not as

8:22

strong or as intuitive as a lot of stuff

8:24

in the ChatGPT app. If you are looking

8:27

to do pure coding, if you're all about

8:29

vibe coding and that's it, you don't do

8:31

kind of AI for all of your knowledge

8:33

work, then this actually is probably the

8:36

best choice for you. You're going to get

8:38

an excellent model at an excellent price

8:40

that's super fast, and you can use

8:41

Cursor for all your coding. And Cursor

8:43

is excellent at coding. The challenge is

8:46

it's just not a very great

8:48

general-purpose

8:49

AI agent harness. If you want to do

8:52

knowledge work with Grok, the best

8:54

option is Grok Bot. Grok Bot is a

8:56

fantastic app. I did a review on it

8:59

yesterday. You should check that out if

9:00

you haven't yet. But, it's more for your

9:03

kind of knowledge work. It's not as good

9:04

for coding, and it's a little bit

9:06

limited with local computer and browser

9:08

use. But, it is very, very good. So,

9:11

that is the challenge for Grok right

9:13

now. It is a great model. It is cheap,

9:15

and it is fast, but it doesn't have that

9:18

amazing, incredible harness just yet

9:20

like ChatGPT desktop app is. Claude, to

9:24

a lesser extent, the Claude desktop app

9:26

is really, really good. I wouldn't be

9:28

surprised if they completely rebrand

9:31

Cursor to like Grok agent or Grok

9:33

desktop very, very soon. The Cursor team

9:37

built Grok Bot, and they didn't call it

9:39

Cursorbot, they called it Grokbot. So, I

9:40

wouldn't be surprised if they repurpose

9:43

Cursor to be Grok desktop very soon to

9:46

be kind of the ChatGPT desktop app for

9:49

SpaceX AI. You know, that doesn't even

9:51

mention a lot of the other nice-to-have

9:53

tools that ChatGPT and Claude has like

9:56

ChatGPT voice, Claude voice, two

9:59

revolutionary technologies that's kind

10:01

of missing from Cursor. The mobile app,

10:03

right? Cursor just came out with their

10:04

mobile app. It's very good, but it's not

10:06

quite the ChatGPT and Claude mobile

10:09

apps. So, who is Grok for? Who should be

10:11

using it? Well, I think a lot of people

10:14

should be using it. I think if your

10:17

workflow is primarily vibe coding, this

10:21

is kind of the best pure vibe coding

10:23

model there is right now. It can build

10:25

just as good as the other models, but do

10:27

it for way faster, way less cost. If you

10:30

are an AI power user and you need voice

10:33

while you're on the go so you can talk

10:35

to your agent and build things on the go

10:37

and you want to be, you know, by the

10:39

pool and you want to send a command to

10:41

your agent to do and tinker on your

10:43

computer and do a bunch of things. Like

10:44

you're a power user like that, you need

10:46

kind of a power user harness, Cursor's

10:49

not quite there yet. ChatGPT is there,

10:51

Claude is basically there, but I will

10:54

say this, what SpaceX AI has pulled off

10:58

over the last several months is

10:59

miraculous. Grok was by far way behind

11:03

Claude and ChatGPT. Now they are neck

11:06

and neck, they are right there. So, that

11:08

wouldn't shock me if again Cursor turns

11:10

into Grok desktop and then it's just as

11:13

good as the other desktop tools and is

11:15

just as good of a harness. Grokbot

11:18

absolutely amazing. Everyone should be

11:21

using Grokbot. I've told you this is

11:24

like the best kind of for the normie AI

11:26

agent app I've ever used. It's

11:28

excellent. That should be used here as

11:30

well with Grok. I'd still use ChatGPT if

11:33

you do a ton of general purpose

11:35

knowledge work and you're a power user

11:37

that need AI using every device you

11:39

have. Claude, I'm still using Claude

11:42

basically just for Fable 5 business

11:45

planning. I think from a high-level

11:48

business strategy perspective, Fable 5

11:51

is still the best model. I don't use

11:53

Claude for literally anything else at

11:55

all. Claude has basically fallen behind

11:57

on all levels. I really feel like this

12:00

desktop app hasn't changed in a very

12:02

long time. It feels like it's been

12:04

absolutely months since anything has

12:06

changed in this app. I'm not sure what's

12:08

going on with the Claude team. It seems

12:11

like releases has slowed down

12:12

dramatically over the last few months. I

12:14

wonder if it's because of their compute

12:15

limitations. They just signed a deal

12:17

with SpaceX AI to get a whole bunch of

12:19

compute. Maybe they're getting that on

12:21

boarded and will be able to release

12:23

quicker in the future. I wouldn't be

12:25

surprised if they just have an explosion

12:27

of releases over the next few weeks. I

12:29

wouldn't be surprised we get Fable 5 1

12:31

very soon. There's a lot of rumors

12:33

around that and I'm sure this whole

12:35

competition flips on its head when that

12:37

comes out as well. But Grok 4 6 pound

12:39

for pound is probably the best vibe

12:41

coding model out there right now. If

12:43

you're tight on money or you're just

12:44

purely about vibe coding, there's no

12:46

better model to be using. I do it in

12:49

Cursor. Again, get Cursor, use it in

12:51

there. It's great inside Cursor. For me

12:54

personally, I'll still be using ChatGPT

12:55

and Claude as well because I have so

12:57

many of those general purpose use cases

12:58

I still do with AI agents. I'm going to

13:01

be doing a full Grok bot boot camp this

13:03

week in the vibe coding academy. Make

13:05

sure you sign up for that. Link for

13:06

that's down below. It's the number one

13:07

AI community on planet Earth. Best

13:09

decision you'll ever make joining that.

13:10

If you learn anything, make sure to

13:11

subscribe and turn on notifications.

13:13

Leave a like. So grateful you'd watch

13:15

this video. Seriously, thank you so much

13:17

and I will see you in the next one.

Interactive Summary

Grok 4.6 has been released as a strong contender in the frontier AI space, offering performance comparable to top models like GPT 5.6 and Claude Opus 5 at a fraction of the cost and speed. The reviewer puts Grok through various benchmarks, including code generation and simulation building, finding it highly effective for coding tasks, particularly within the Cursor environment. While Grok excels in coding and efficiency, it currently lacks the comprehensive general-purpose AI agent ecosystem found in ChatGPT. The video concludes that Grok is the top choice for 'vibe coding,' though power users requiring integrated voice and broad knowledge tools may still rely on other platforms.

Suggested questions

3 ready-made prompts