HomeVideos

xAI just caught up (Grok 4.6 is here)

Now Playing

xAI just caught up (Grok 4.6 is here)

Transcript

718 segments

0:00

Seems like Xi is on a bit of a tear

0:02

lately since they acquired Cursor.

0:03

They've been moving so much faster. We

0:05

went from a drought where they just kind

0:06

of didn't put out models that mattered

0:08

or really at all for months at a time to

0:10

getting a bunch back to back. We just

0:12

had Croc 45 not long ago and it really

0:15

impressed me. It's not the model I

0:16

default to for my everyday work, but

0:18

it's the first model that's good enough

0:20

from a lab that isn't anthropic or open

0:22

AI that I could see myself actually

0:23

daily driving it. And now Grock 46 is

0:26

here and they've went even further with

0:27

it. With everything I've seen and all

0:29

the playing I've been doing so far, it

0:31

seems incredibly promising. It's fast,

0:33

it's cheap, it's reliable, and it's good

0:35

at doing long running tasks with lots of

0:38

sub agents and not losing track of what

0:39

it's supposed to be getting done. This

0:41

is important because as the way we build

0:43

with agents changes, the ability for

0:45

agents to go off for a long time and

0:47

solve hard problems and come back with

0:49

working results is more important than

0:51

ever. Sadly, that doesn't mean this

0:53

model's perfect, though. There are rough

0:54

edges as there are with everything and

0:56

I've already encountered a handful of

0:57

those with my own testing. There's also

0:59

some big bright sides to this as well

1:01

that I've been experiencing the more

1:03

I've been using the official Grock Build

1:05

CLI which is awesome by the way. I

1:06

actually really like the CLI. There's

1:08

layers to this one. There are some

1:10

really crazy benchmarks which uh spoiler

1:13

Gro 46 is now neck andneck with GPT56

1:17

Soul according to artificial analysis.

1:20

We need to see if this holds up because

1:22

if it's actually getting that close at a

1:24

fifth the price, things are changing

1:26

quickly. While Intelligence seems to be

1:28

getting cheaper, my payroll isn't. So, I

1:30

hope you pardon me for a quick break for

1:32

today's sponsor. I have a real quick

1:34

question for you. If you're building an

1:35

app and you want to add video to it, how

1:37

are you going to do that? Whether it's a

1:38

video on the homepage that you want to

1:40

use to market your site or userfacing

1:42

uploads where they can upload whatever.

1:44

How do you know that these videos are

1:45

being served correctly, that they're

1:46

being processed properly, that they

1:48

don't have inappropriate content? And

1:50

what happens when you want to do things

1:51

like generate captions or get more info

1:53

about the video? Well, I hope you would

1:54

stumble on today's sponsor, MX, before

1:56

you get too far, because they solve all

1:58

of those problems and more. Whether

2:00

you're just adding one video to a small

2:01

homepage on a side project, or you're

2:03

trying to build a service that competes

2:05

with YouTube, there's a reason everyone

2:07

picks MX. Whether it's Fox, Band Camp,

2:09

or Robin Hood, or even companies like

2:10

Cursor or Dropbox, everyone leans on

2:13

these guys for a reason. It's cuz they

2:14

get video. I've built a lot of video

2:16

services, and if I learned anything in

2:18

that time, it's that processing and

2:20

managing user generated content is just

2:22

a huge rabbit hole you want to avoid to

2:24

the best of your ability. And that's why

2:26

I'm so thankful for the new MX Robots

2:28

product where you can do all of the

2:29

automation you would ever need to

2:31

trivially. You can set up Muk to answer

2:33

questions about videos as they're

2:34

uploaded using AI, which is super

2:36

powerful by itself. But when you combine

2:37

that with the captions and the

2:39

moderation, as well as a summary

2:40

feature, there is so much stuff you can

2:42

do without having to build all of the

2:44

details yourself. And once you're

2:46

processing that video, it gets really

2:48

expensive really fast. But when you use

2:50

a service like MX, you know you're

2:51

getting the best possible deal. I could

2:53

tell you all about how great the code is

2:55

and how easy it is to integrate, but

2:56

honestly, you should ask your agent. it

2:58

already knows about MX because these

2:59

guys have been the default for a long

3:01

long time. There's a reason I've been

3:02

building a MX for eight years and you

3:04

can figure out why at sidv.link/mucks.

3:06

Let's start from the official

3:07

announcement and then go into the real

3:09

world use cases, the benchmarks and my

3:11

own experience building with it.

3:12

Introducing Gro 46. Gro 46 builds on Gro

3:16

4.5 with a particular focus on

3:18

longrunning agents and more ambitious

3:20

interactive and visual work. translated

3:22

to actual English quickly. Grock 4 6

3:26

isn't a new pre-training. It is a new

3:29

post-training on Grock 45 because part

3:31

of acquiring cursor is getting all of

3:33

cursor's crazy RL and post-training

3:35

stuff. And now they have all of this

3:37

post- training and RL stuff. They're

3:39

able to improve the model after the

3:40

pre-training is done. And this is the

3:42

technique that makes these longunning

3:44

tasks and agentic work stuff much more

3:47

reliable. That's usually done in post-

3:49

training, not pre-training. As they

3:50

mentioned before, 46 is building on 45

3:52

with a particular focus on longrunning

3:54

agents and more ambitious interactive

3:56

and visual work. It stays with complex

3:57

tasks across many steps. Whether it's

3:59

researching topics, analyzing

4:00

information, working across code bases,

4:02

or turning ideas into polished apps or

4:04

work artifacts. As I mentioned before on

4:07

the artificial analysis intelligence

4:08

index, it is neck andneck with 56 soul

4:11

and right behind fable 5 and opus 5. I

4:14

am obligated to say that this shows that

4:16

these numbers aren't the best way to

4:17

measure how smart or useful a model

4:19

actually is because for my experience,

4:21

Opus 5 is significantly worse to use

4:24

than Fable and Saul. It looked and felt

4:26

really good day one, but the more I

4:28

started merging the code it wrote, the

4:29

more problems I ran into, the more

4:31

cleaning up after the Opus sloppus as I

4:34

now call it. Yeah, Opus 5 seemed like it

4:37

was good and by every measure seemed to

4:39

be solid and the more I use it, the more

4:40

I hate it and I'm just on Fable all the

4:42

time. So again, this bench isn't the

4:44

most trustworthy thing, but there are

4:46

others in here that are good, like deep

4:47

sui, where it is just behind soul and

4:50

fable and meaningfully ahead of Grock

4:52

45. Like this is a huge bump for a

4:54

post-training improvement to go from a

4:56

54% to a 66 is huge. Cursor bench saw a

5:00

smaller jump, but it did jump Grock 46

5:04

past 56 soul, but remember they

5:07

accidentally trained some of Cursor

5:09

Bench data into the models. I think they

5:11

might have cleaned it up with cursor

5:12

bench 32. Not sure, but you get the

5:14

idea. And then with Frontier Code, they

5:16

saw a meaningful jump as well from 56.6

5:19

to 61.3, putting it smack dab in the

5:21

middle between Soul and Fable. It's

5:23

available now in cursor and Grock build.

5:25

They've offered 2x usage inside of your

5:27

Grock build and cursor subs for the

5:29

first week, so you can start trying it

5:30

immediately as I have been doing. I am

5:32

just on my subscription on Twitter right

5:34

now because I already have the like blue

5:36

check on X because they pay me for it.

5:38

So, I'm on the the pro tier, the one

5:40

that kills ads as well. X premium

5:43

pricing. Let me see what the costs are

5:45

for the different tiers. Yeah, I have

5:47

the $40 tier because I don't like ads

5:49

and it pays for itself because I'm paid

5:51

to use Twitter, as stupid as that is.

5:53

So, I've just been using Gro 46 through

5:55

the included data I get with that

5:57

subscription. And I've used it to

5:59

generate a handful of games and also do

6:01

some real work in T3 code. Specifically

6:03

trying to get things to parody with both

6:06

cursor and gro in T3 code because we do

6:09

have the Grock bindings built in. Shout

6:11

out to the Grock build team for working

6:13

with us on that. Really proud that we

6:14

got that in as quick as we did. So I was

6:16

working on tidying it up using Grock and

6:18

so far pretty solid. Show you guys

6:21

everything I did with it after we get

6:22

through the official information. Grock

6:25

46 underwent a longer supplemental

6:27

training run than 45 did with curated

6:29

model generated data for reasoning and

6:31

advanced technical concepts, highquality

6:33

engineering data, and an improved

6:34

optimizer and training recipe produce a

6:37

stronger foundation for the SFT and RL

6:39

stages that followed. When they used Gro

6:41

45 to regenerate the SFT trajectories

6:43

across reasoning efforts, agent

6:45

harnesses in domains such as STEM,

6:46

software engineering, and knowledge work

6:48

as well as filtering out problematic

6:49

traces with modelbased checks. The

6:51

resulting SFT checkpoint shows strong

6:53

performance and improved behavior. 46 is

6:56

trained on a wide range of agentic RL

6:58

tasks including knowledge work, general

6:59

coding, and domain specific environments

7:01

for kernel optimization, webdev,

7:03

computerated design, and more. I'll show

7:05

you guys some Grock 46 designs, don't

7:06

worry. They tested Grock 46 on projects

7:09

designed to stretch its range and

7:10

ability to sustain work over many steps.

7:12

The model's especially strong at turning

7:14

a broad product idea into working first

7:16

versions. It can research unfamiliar

7:18

domains, structure the app, implement

7:20

the core interactions, and continue

7:21

refining the result through several

7:23

rounds of feedback. They've also seen it

7:25

doing more self tests and verifications

7:26

on longer trajectory runs with the model

7:28

checking its own work before moving on.

7:30

Good. That's been an issue I had with

7:32

grock models is they would just write

7:33

the whole thing and it would work

7:35

sometimes and then fail other times.

7:37

Sadly, that was not really my experience

7:40

with the game gen stuff I was doing,

7:42

which I will show you all momentarily.

7:45

They added more safeguards due to new

7:46

safety capabilities and potential to

7:48

hack and whatnot. Safeguard evaluation

7:51

work reflects 46 expanded capabilities

7:53

with our widest ever suite of

7:54

pre-eployment testing for capabilities

7:55

and safeguard calibration as well as

7:57

extensive post- deployment and

7:58

third-party testing. Let's see how this

8:00

goes. Do a thorough security audit of

8:04

this project. Are we ready to launch?

8:09

What should we fix first? We'll see how

8:12

it does with a thorough security audit

8:14

in a real codebase. I have it pointed at

8:16

my cloud that I'm still hopefully

8:18

getting out soon. Lake bed. They show a

8:20

chunk of eval here which we'll look at

8:21

while I wait for that to run.

8:23

Intelligence inex we already covered.

8:24

GDP val I don't care about. Cursor bench

8:27

it's still behind fable but it's now

8:28

ahead of 56 soul which is crazy. Deep

8:31

suite it's getting up there. It's neck

8:32

andneck with soul and fable. Actually

8:34

it's a little bit more behind there but

8:36

still huge improvement for 45. It got

8:38

bestin-class in AA briefcase and Harvey

8:41

Lab vals, which are interesting. Also,

8:44

the best in GDP val, but again, like I

8:45

don't care. It's GDP val. SpaceXI's Gro

8:48

4.6 scores a 61 on the artificial

8:50

analysis intelligence index, joining the

8:52

frontier in line with 56 soul with

8:54

standout agentic performance at lower

8:56

costs. 46 gains five points over 45 in

8:59

the intelligence index just over one

9:01

month after its release. That's a

9:02

five-point improvement in a few months

9:04

and 23 points compared to 43, which is

9:06

really not that long ago. The pace of

9:08

releases from these guys has been nuts

9:10

lately. And it looks like that pace will

9:12

maintain because Elon's already tweeting

9:14

saying that 4.7 is significantly better

9:16

than 4.6 and it should be ready in 3 to

9:18

4 weeks. The initial training is

9:20

complete and they're adding a massive

9:21

amount of SpaceX company data in

9:22

supplementary training and they're

9:24

expecting it to be quote something

9:26

special. Very interesting. This is the

9:29

jump where they are now in the frontier

9:31

range. comparable to the other frontier

9:33

models according to artificial analysis.

9:35

Much stronger agentic performance

9:36

especially with things like GDP val

9:38

frontier level intelligence at lower

9:39

costs because again it's way cheaper.

9:42

It's $2 per mill in and $6 per mill out

9:44

which is 60% below Opus 5's price and

9:48

even more than 56 souls price costs 84

9:51

cents per task which is similar to Kimmy

9:53

K3 with slightly higher intelligence

9:55

which gives it a really good place on

9:57

the Pareto Frontier. The cost versus how

10:00

smart it is. Cost per task is right

10:03

there with Kimmy K3. Slightly more

10:06

expensive than Gemini 36 Flash and

10:08

Terra. Slightly less expensive than 56

10:10

Soul. Comically cheaper than Opus 5 and

10:12

Fable 5 for realworld use. That is $3.14

10:16

for Fable and 84 cents for Gro 46. This

10:19

is more than double the price that Grock

10:21

45 was though, which is the important

10:23

detail they seem to be hiding. I'm

10:25

guessing this price gap is a token

10:26

difference between 46 and 45, which we

10:28

can confirm here. Yep. I was so pumped

10:32

that Grock 45 was as efficient as it

10:34

was. Seeing them lose that with 46 is a

10:37

little sad. They've increased the number

10:39

of tokens per run by over 30% here.

10:43

Previously, they were one of the few

10:44

major providers that had a model that

10:46

was more token efficient than 56 Soul.

10:48

Now, they're not. Now they are still

10:50

relatively efficient, like comparable or

10:53

better than Fable efficiency-wise, but

10:55

it's no longer that magic efficient

10:58

small fast cheap model that I wanted it

11:00

to be. This is something I've actually

11:03

felt in my own usage because I noticed

11:05

that the Gro 45 test when I did them ran

11:08

almost immediately, like creepily fast.

11:10

46 is back to the usual send off the

11:13

prompt and then go do something else

11:14

that I'm used to from Frontier. They

11:16

talk about their AA briefcase bench

11:18

where it is now Fable 5 tier. Context

11:21

window is the same at 500k tokens. The

11:23

price is the same as well. They did

11:26

increase the price of cash reads.

11:28

Previously it was only 3 cents per

11:30

million tokens red. Now it's 5 cents per

11:31

million tokens red for the cash. That

11:33

might also be affecting price here. Some

11:35

of these jumps are nuts. Like I know I

11:37

don't like GDP Val, but this is one of

11:40

the biggest jumps I've ever seen in that

11:42

benchmark. Like that's crazy.

11:45

It's clearly much better at these long

11:46

running things, but it also thinks a lot

11:48

more, too. And again, as I said, it is

11:50

now much more expensive, but it's also

11:52

slightly smarter. It's got a really good

11:53

score here in the cost versus

11:56

intelligence, but this does yank them

11:58

out of this golden area of like the

12:00

green corner where it is the cheapest

12:02

and the smartest and forces it out over

12:05

to be on the right side of the median

12:08

for cost on the log scale, which is sad.

12:10

They're clearly trying to rush their way

12:12

to the frontier. So now it's time to see

12:15

how it does with UI. I have other things

12:17

I've been working on with it in the

12:19

background, too. But I am very curious

12:21

how its design work is because it's a

12:23

different model, and it is nice to see

12:26

the different ways different models

12:27

handle design. As I mentioned in my Muse

12:29

Spark video, I was actually surprised at

12:31

how different the Muse models felt for

12:33

design work than the usual anthropic and

12:35

OpenAI slop. I've experienced less of

12:37

that with Grock models where they feel

12:39

more like a dumbed down version

12:40

design-wise. So, let's see what they

12:42

cooked. I'll start with none of the

12:44

skills on and just go through these

12:47

Grock 46 outputs. Shout out to Dra for

12:50

making the site. By the way, the

12:51

witchai.dev is super handy for this type

12:53

of video. Here's the first design that

12:55

it made. I don't like the noise pattern,

12:58

and I really don't like the way that the

13:00

text is so sharp with the blurryish

13:02

background. It just doesn't feel

13:04

cohesive. This feels like old era AI

13:07

Slavic. This is GPT 5.0 is how this

13:10

feels to me. The purple in particular is

13:14

This one has a nice little mock app in

13:15

the corner, which is fine. The coloring

13:17

is okay. Too many cards. Too old school

13:21

like Tailwind Templaty. This is the

13:24

usual brutalism that we see from a lot

13:25

of these. And this is a boring centered

13:28

page. Yeah, none of these impress me. I

13:31

did happen to notice that when you use

13:32

the design skill from Claude, it can and

13:35

do better. Some of them are sinful and

13:37

disgusting like whatever the this

13:38

is. Some of them have okay ideas on them

13:41

like this with the floating card things

13:43

when you hover. Yeah, could be tidied

13:45

up. This is sinful. This is sinful. Not

13:49

digging this. And just for reference,

13:51

here is Fable doing similar designs.

13:55

Much much better, like comically. So, I

14:00

hate that the terminal ones, but other

14:02

than that, or even like Opus does a much

14:04

better job here with these, or even 56

14:08

Soul, as cringe as it often is, is

14:10

meaningfully better at this type of

14:12

thing. Like, this is a much nicer

14:14

version of this style of page. So, not

14:16

great at design at all. I will show the

14:19

real work I was doing with it earlier,

14:21

but I first want to play with these fish

14:22

slop ports because they're fun and I'm

14:25

curious how it did. Those who aren't

14:26

familiar, I made a game at the end of

14:28

last year with Opus called Fish Slop

14:30

with Opus 45. Never got very far with

14:33

it. It was meant to be an insane

14:34

aquarium clone. So, I like throwing

14:36

models at the old codebase and say,

14:38

"Rebuild this. You can reuse assets or

14:39

whatever, but from scratch, make a new

14:41

version of this game using that codebase

14:43

as a reference." And it's a fun way to

14:45

see how well it like understands a lot

14:47

of end toend

14:49

work in a complex thing like the

14:51

relations between systems in a game. I'm

14:53

already noticing that like the controls

14:55

are uh bad. The sizing of elements is

14:58

all wrong. Like these fish pellets are

15:00

too big and the fish are too small. The

15:03

fact that the thing you had highlighted

15:05

on the bottom here stays highlighted

15:06

when you move off it is really bad.

15:10

Yeah, the movement just feels so not

15:13

great. The pace of the game seems wrong,

15:16

too, where things don't like grow fast

15:18

enough and you don't get to the next

15:19

stage quick enough. Yeah, this just

15:22

doesn't feel great. This is like Kimmy

15:26

K3 did a meaningfully better job with

15:28

this than Gro 46 seems to be. But this

15:32

is the easy version. The hard version is

15:35

3D. This is what I got when I first

15:38

opened the 3D version of the game. A

15:40

black screen. This is the first new

15:42

model I have pointed at the task of a

15:45

make the game 3D and had it just

15:48

outright fail. All the others might have

15:51

had bugs and issues with like the

15:53

movement or the mechanics or the models

15:54

they created in the 3D space. All of

15:56

them opened with a 3D environment. Even

15:59

like GBT4 was able to get that far. This

16:03

is the first one that just outright

16:05

failed to do the 3D port. I did tell the

16:07

model that it failed. I even gave it a

16:09

picture which showed me how nice the

16:10

rendering of pictures and things is

16:12

inside of the Gro CLI, which like the

16:14

Grock build CLI is actually genuinely

16:16

really nice. Well, I told it to fix it

16:18

with the screenshot. It appears that it

16:20

did. So, let's dive in and see. Oh god,

16:23

this is so cringe. They got the

16:25

directions wrong. So, up and down is

16:26

right, but left and right are inverted.

16:28

I don't know how you get right and left

16:30

inverted, but they did. The placement of

16:33

everything at the ground is entirely

16:35

wrong. Like, comically so. God, this is

16:40

very bad and broken. [clears throat]

16:44

The models for the fish are This

16:47

is the worst 3D pass I've seen on this

16:50

in a long time. Actually, this is like

16:53

last year I would have expected quality

16:55

like this. Just for reference, here's

16:57

the version that Muse12 made for way

17:01

cheaper. It got some directional issues

17:03

like up and down tilt it, but like it

17:07

functions. The movement's way better.

17:09

The core mechanics work a lot better. If

17:11

we want to go to real frontier, like I

17:13

don't know, uh, Kimmy, look at how much

17:17

further along this is. The modeling of

17:20

the things on the ground. The lighting's

17:23

a little broken, but like

17:26

this is a generational gap. This is an

17:29

open weight model, by the way. Like, you

17:30

can go download the weights online and

17:32

have a thing that is this much better at

17:34

3D. And the fish model is so much

17:37

better. It's actually the best 3D

17:38

modeling I've seen any of the models do

17:41

so far. So yeah, there's your

17:43

comparison. Grock is not even

17:46

unimpressive. It's like last generation.

17:50

It's so far ahead here. But then there

17:52

is realworld work. And for this I was

17:55

quite kind to Gro 45. I found it

17:58

surprisingly capable of doing like real

18:01

work where it had to touch different

18:02

things that were complex and stay on

18:05

task for longer running things. Well,

18:07

take a look at the security audit it

18:09

just did. It found a previous audit that

18:11

I did in the git history. So, that's

18:13

cool. So, it's already strong. The O

18:16

broker, the worker isolation storage

18:17

control plane will get us burned. Give

18:20

us advice for locking endpoints. Put it

18:22

on the public suffix list. I already was

18:24

planning on that. like bed chrome on

18:26

every capsule page. Finish the admin

18:29

trust and safety UI. Stop putting

18:30

identity tokens in URLs or a wider

18:32

launch. Refuse sandbox off flags in

18:35

production. Yeah, did a decent audit

18:37

here. Didn't get me any errors or

18:38

problems when it did it. There are other

18:40

deeper things it didn't find, but that's

18:43

acceptable. This is the more interesting

18:45

one I wanted to take a look at with

18:46

y'all. I asked Grock 46 to take a look

18:49

at how we have T3 code implementing

18:52

cursor because the current build is

18:54

using the outofdate ACP adapter for the

18:56

cursor CLI. They want us to move to the

18:58

SDK. I believe Julius is work on the new

19:00

orchestrator includes that. But I wanted

19:02

to see what it would do. So I asked it

19:04

to go through and audit things and see

19:06

how we would migrate from ACP to the

19:08

Cursor SDK. For long lived gooey hosts

19:11

that need a real model catalog, resume,

19:13

cancel images, and usage. The SDK is the

19:15

one that cursor is actually maintaining.

19:16

Yep. So why is ACP the wrong host

19:19

contract? Then built in fallbacks. Yeah,

19:21

these are all real problems that we've

19:23

had with the current cursor bindings. So

19:24

it was correct there. Official docs

19:27

license anywhere to not MIT. That is

19:30

annoying, but that's fine.

19:33

The SDK doesn't offer guey level approve

19:36

or deny. That's annoying, but we auto

19:38

run anyways for most things. Blocking

19:40

ask questions. Uh the SDK will not allow

19:43

that. and reusing agent login. That will

19:46

not work. So, we have to implement our

19:47

own login, which is annoying. You know

19:49

what I will do? I don't feel like

19:50

reading gro slop. We'll have 56 soul

19:53

share its thoughts momentarily. Plan

19:55

approval could not be completed because

19:56

the client disconnected. Plan mode

19:58

remains active. Okay, so it put itself

20:00

in plan mode. That is obnoxious. Did it

20:02

put the plan in here anywhere? It

20:05

didn't. That is very annoying. Yeah, the

20:06

Grock implementation needs a little bit

20:08

of work as well. I will put some time in

20:10

in the near future. I have to go turn on

20:12

the legacy plan mode for this.

20:14

While those plans are being audited,

20:16

I'll show off a little bit of it

20:18

behaving how I wanted it to. Here I

20:20

asked it, what gaps exist in T3 Code's

20:22

implementation of Grock build compared

20:23

to other harnesses. It went through it

20:26

made a plan. I don't know where it went

20:27

in the UI. There there's some weirdness

20:28

in the events that the Grock build like

20:30

ACP sends out. So, we have to put more

20:33

work into how we clean those events up.

20:35

But I asked at the end here, I wanted to

20:37

see it differently. So I just said HTML

20:39

which triggers my HTML skill and it

20:42

responded there. The UI collapsed it

20:44

again because their events are broken.

20:46

Uh I gave it some feedback on the plan.

20:49

Had it update then asked it to build the

20:52

whole plan file a PR and babysit which

20:55

it did. I can go open it on GitHub and

20:57

we can see what things thought here. We

21:00

got a bunch of feedback from Macroscope

21:02

on the changes. It tore things to shreds

21:04

here, but it seems like it followed my

21:06

instructions really well. You'll see

21:08

here that it used my account to reply to

21:09

things, but I have a little call out in

21:12

a skill I made for PR comments where I

21:14

tell every agent to open its comments

21:16

with note which model is responding on

21:18

behalf of Theo. This is not a small

21:21

change either. This is a thousandline PR

21:24

that it made to address a bunch of

21:26

different gaps in our coverage for Grock

21:28

Build. And I told it after it makes the

21:31

PR to babysit it once it was filed. Then

21:34

I told it to do one of my favorite

21:36

things, which is to tell the agent to

21:38

look through my actual history on my

21:40

actual computer and find gaps in what

21:42

events occurred and what we actually

21:44

process. Just told it to make a separate

21:46

PR stacked on the one that we currently

21:48

have to do all these additional changes.

21:50

This is a complex issue because it needs

21:52

to know how to like figure out and stack

21:55

PRs. It needs to keep track of the gap

21:58

between the old PR and the new one. All

22:00

the investigatory work it just did in

22:01

the context of how it applies with my

22:03

history on this machine as well as all

22:05

the other skills and things that it

22:07

pulled in. It's not easy for models to

22:10

deal with all of this like competing

22:12

context and stay on track. This is one

22:14

of the things again I thought Grock 45

22:16

did surprisingly well. It was able to

22:18

take multi-step unrelated work and do it

22:22

all cohesively and coherently in one

22:24

thread. So, we shall see how it handles

22:26

this here. I just put together a really

22:28

rough like how I would score Grock

22:31

compared to Fable and Soul for like

22:33

different categories. I think about

22:35

things like cost where Fable's way too

22:37

expensive. Soul's significantly better.

22:38

Probably do a little lower there, but

22:40

like reasonably priced. And then Grock

22:42

45 way better on cost. Intelligence,

22:45

Fable is the best right now. Soul is

22:47

surprisingly good, but not quite as

22:48

good. Grock is not even this. I would

22:51

put a little lower there. Then you have

22:53

speed where fable is very slow. Soul is

22:56

meaningfully faster simply because it's

22:57

so much more efficient. And then Grock

22:59

45, especially on the fast mode, flew.

23:01

It was super fast and really nice. Then

23:03

with thoroughess, Fable, I find, isn't

23:05

quite as thorough as it should be. It

23:07

misses things here and there, but it's

23:09

it's thoughtful, not thorough. Soul is

23:11

incredibly thorough. It checks every

23:13

single thing. It touches every single

23:14

edge, and it writes too much code as a

23:16

result. 45 just did not have any of

23:18

that. And then with orchestration

23:19

capabilities, Fable was the best. Soul

23:21

is surprisingly close in Grock 45.

23:24

Better than most, but still not quite

23:25

there. So, how does this all compare to

23:27

our new model Grock 46? Sadly, cost is a

23:31

regression. I'd say this is like a

23:34

sixish now. Intelligence seems like a

23:37

meaningful bump. I haven't seen too much

23:39

of it being way smarter, but I can

23:41

confidently say it's probably

23:42

meaningfully better. We'll give it a 6.5

23:44

there. Speed is where I'm seeing one of

23:46

the biggest regressions now where I

23:48

would put it at like a five and a half

23:50

at best right now. Thoroughess, it is

23:53

better, but it still misses things. I'll

23:55

bump it slightly there. In orchestration

23:56

capabilities, I did see it using sub

23:58

agents decently well. I'll give it a 6.5

24:00

here. Why not? The issue is the only

24:03

things I would consider Grock previously

24:05

to be a leader in were the speed and the

24:09

cost. And we saw regressions in both of

24:13

those categories with this release. The

24:15

cost went up, not per token, but since

24:18

the token efficiency went down, it's now

24:19

more expensive. And the speed went down

24:22

because it's less token efficient. So,

24:24

it's generating more tokens and it takes

24:26

longer. These two changes make Grock

24:29

much, much less interesting to me. But

24:32

the speed at which the intelligence is

24:34

growing suggests that Gro 47 will do a

24:37

lot more in all of these other

24:39

categories to catch up. The thing I

24:41

liked Grock 45 for wasn't that it was as

24:43

good as Frontier at specific things.

24:45

It's that the speed and the cost were

24:48

far enough away from Frontier that it

24:50

felt uniquely useful in those ways. And

24:52

I feel like we are losing that with

24:54

Grock 46 a little bit. I think that

24:56

summarizes my thoughts on this release.

24:59

Is it my favorite model? No. Am I going

25:01

to use it a whole lot after this video?

25:03

Probably not. But does it have me

25:04

excited for Grock 4.7? Absolutely yes.

25:08

It makes a lot of sense why Elon's

25:09

talking more about 47 than 46 right now.

25:13

The future seems very clear and it seems

25:15

like a future where Grock catches up

25:17

very fast and I'm honestly pretty

25:19

excited for that. We need more

25:20

competition. We need more good models

25:22

and we need more people fighting to make

25:23

them cheaper. Let me know how y'all

25:25

feel. Is this an exciting release for

25:26

you or are you going to just ignore it?

25:27

Let me know in the comments. And until

25:29

next time, peace nerds.

Interactive Summary

The video provides a detailed analysis of the newly released Grok 4.6 model. While the model shows improvements in intelligence, agentic performance, and thoroughness, the reviewer highlights significant regressions in speed and cost-efficiency, which were previously Grok's key competitive advantages. The video includes practical testing through game generation, UI design evaluation, and real-world software engineering tasks, concluding that while 4.6 is a step forward, the anticipation for Grok 4.7 is much higher due to the rapid pace of development.

Suggested questions

4 ready-made prompts