xAI just caught up (Grok 4.6 is here)
718 segments
Seems like Xi is on a bit of a tear
lately since they acquired Cursor.
They've been moving so much faster. We
went from a drought where they just kind
of didn't put out models that mattered
or really at all for months at a time to
getting a bunch back to back. We just
had Croc 45 not long ago and it really
impressed me. It's not the model I
default to for my everyday work, but
it's the first model that's good enough
from a lab that isn't anthropic or open
AI that I could see myself actually
daily driving it. And now Grock 46 is
here and they've went even further with
it. With everything I've seen and all
the playing I've been doing so far, it
seems incredibly promising. It's fast,
it's cheap, it's reliable, and it's good
at doing long running tasks with lots of
sub agents and not losing track of what
it's supposed to be getting done. This
is important because as the way we build
with agents changes, the ability for
agents to go off for a long time and
solve hard problems and come back with
working results is more important than
ever. Sadly, that doesn't mean this
model's perfect, though. There are rough
edges as there are with everything and
I've already encountered a handful of
those with my own testing. There's also
some big bright sides to this as well
that I've been experiencing the more
I've been using the official Grock Build
CLI which is awesome by the way. I
actually really like the CLI. There's
layers to this one. There are some
really crazy benchmarks which uh spoiler
Gro 46 is now neck andneck with GPT56
Soul according to artificial analysis.
We need to see if this holds up because
if it's actually getting that close at a
fifth the price, things are changing
quickly. While Intelligence seems to be
getting cheaper, my payroll isn't. So, I
hope you pardon me for a quick break for
today's sponsor. I have a real quick
question for you. If you're building an
app and you want to add video to it, how
are you going to do that? Whether it's a
video on the homepage that you want to
use to market your site or userfacing
uploads where they can upload whatever.
How do you know that these videos are
being served correctly, that they're
being processed properly, that they
don't have inappropriate content? And
what happens when you want to do things
like generate captions or get more info
about the video? Well, I hope you would
stumble on today's sponsor, MX, before
you get too far, because they solve all
of those problems and more. Whether
you're just adding one video to a small
homepage on a side project, or you're
trying to build a service that competes
with YouTube, there's a reason everyone
picks MX. Whether it's Fox, Band Camp,
or Robin Hood, or even companies like
Cursor or Dropbox, everyone leans on
these guys for a reason. It's cuz they
get video. I've built a lot of video
services, and if I learned anything in
that time, it's that processing and
managing user generated content is just
a huge rabbit hole you want to avoid to
the best of your ability. And that's why
I'm so thankful for the new MX Robots
product where you can do all of the
automation you would ever need to
trivially. You can set up Muk to answer
questions about videos as they're
uploaded using AI, which is super
powerful by itself. But when you combine
that with the captions and the
moderation, as well as a summary
feature, there is so much stuff you can
do without having to build all of the
details yourself. And once you're
processing that video, it gets really
expensive really fast. But when you use
a service like MX, you know you're
getting the best possible deal. I could
tell you all about how great the code is
and how easy it is to integrate, but
honestly, you should ask your agent. it
already knows about MX because these
guys have been the default for a long
long time. There's a reason I've been
building a MX for eight years and you
can figure out why at sidv.link/mucks.
Let's start from the official
announcement and then go into the real
world use cases, the benchmarks and my
own experience building with it.
Introducing Gro 46. Gro 46 builds on Gro
4.5 with a particular focus on
longrunning agents and more ambitious
interactive and visual work. translated
to actual English quickly. Grock 4 6
isn't a new pre-training. It is a new
post-training on Grock 45 because part
of acquiring cursor is getting all of
cursor's crazy RL and post-training
stuff. And now they have all of this
post- training and RL stuff. They're
able to improve the model after the
pre-training is done. And this is the
technique that makes these longunning
tasks and agentic work stuff much more
reliable. That's usually done in post-
training, not pre-training. As they
mentioned before, 46 is building on 45
with a particular focus on longrunning
agents and more ambitious interactive
and visual work. It stays with complex
tasks across many steps. Whether it's
researching topics, analyzing
information, working across code bases,
or turning ideas into polished apps or
work artifacts. As I mentioned before on
the artificial analysis intelligence
index, it is neck andneck with 56 soul
and right behind fable 5 and opus 5. I
am obligated to say that this shows that
these numbers aren't the best way to
measure how smart or useful a model
actually is because for my experience,
Opus 5 is significantly worse to use
than Fable and Saul. It looked and felt
really good day one, but the more I
started merging the code it wrote, the
more problems I ran into, the more
cleaning up after the Opus sloppus as I
now call it. Yeah, Opus 5 seemed like it
was good and by every measure seemed to
be solid and the more I use it, the more
I hate it and I'm just on Fable all the
time. So again, this bench isn't the
most trustworthy thing, but there are
others in here that are good, like deep
sui, where it is just behind soul and
fable and meaningfully ahead of Grock
45. Like this is a huge bump for a
post-training improvement to go from a
54% to a 66 is huge. Cursor bench saw a
smaller jump, but it did jump Grock 46
past 56 soul, but remember they
accidentally trained some of Cursor
Bench data into the models. I think they
might have cleaned it up with cursor
bench 32. Not sure, but you get the
idea. And then with Frontier Code, they
saw a meaningful jump as well from 56.6
to 61.3, putting it smack dab in the
middle between Soul and Fable. It's
available now in cursor and Grock build.
They've offered 2x usage inside of your
Grock build and cursor subs for the
first week, so you can start trying it
immediately as I have been doing. I am
just on my subscription on Twitter right
now because I already have the like blue
check on X because they pay me for it.
So, I'm on the the pro tier, the one
that kills ads as well. X premium
pricing. Let me see what the costs are
for the different tiers. Yeah, I have
the $40 tier because I don't like ads
and it pays for itself because I'm paid
to use Twitter, as stupid as that is.
So, I've just been using Gro 46 through
the included data I get with that
subscription. And I've used it to
generate a handful of games and also do
some real work in T3 code. Specifically
trying to get things to parody with both
cursor and gro in T3 code because we do
have the Grock bindings built in. Shout
out to the Grock build team for working
with us on that. Really proud that we
got that in as quick as we did. So I was
working on tidying it up using Grock and
so far pretty solid. Show you guys
everything I did with it after we get
through the official information. Grock
46 underwent a longer supplemental
training run than 45 did with curated
model generated data for reasoning and
advanced technical concepts, highquality
engineering data, and an improved
optimizer and training recipe produce a
stronger foundation for the SFT and RL
stages that followed. When they used Gro
45 to regenerate the SFT trajectories
across reasoning efforts, agent
harnesses in domains such as STEM,
software engineering, and knowledge work
as well as filtering out problematic
traces with modelbased checks. The
resulting SFT checkpoint shows strong
performance and improved behavior. 46 is
trained on a wide range of agentic RL
tasks including knowledge work, general
coding, and domain specific environments
for kernel optimization, webdev,
computerated design, and more. I'll show
you guys some Grock 46 designs, don't
worry. They tested Grock 46 on projects
designed to stretch its range and
ability to sustain work over many steps.
The model's especially strong at turning
a broad product idea into working first
versions. It can research unfamiliar
domains, structure the app, implement
the core interactions, and continue
refining the result through several
rounds of feedback. They've also seen it
doing more self tests and verifications
on longer trajectory runs with the model
checking its own work before moving on.
Good. That's been an issue I had with
grock models is they would just write
the whole thing and it would work
sometimes and then fail other times.
Sadly, that was not really my experience
with the game gen stuff I was doing,
which I will show you all momentarily.
They added more safeguards due to new
safety capabilities and potential to
hack and whatnot. Safeguard evaluation
work reflects 46 expanded capabilities
with our widest ever suite of
pre-eployment testing for capabilities
and safeguard calibration as well as
extensive post- deployment and
third-party testing. Let's see how this
goes. Do a thorough security audit of
this project. Are we ready to launch?
What should we fix first? We'll see how
it does with a thorough security audit
in a real codebase. I have it pointed at
my cloud that I'm still hopefully
getting out soon. Lake bed. They show a
chunk of eval here which we'll look at
while I wait for that to run.
Intelligence inex we already covered.
GDP val I don't care about. Cursor bench
it's still behind fable but it's now
ahead of 56 soul which is crazy. Deep
suite it's getting up there. It's neck
andneck with soul and fable. Actually
it's a little bit more behind there but
still huge improvement for 45. It got
bestin-class in AA briefcase and Harvey
Lab vals, which are interesting. Also,
the best in GDP val, but again, like I
don't care. It's GDP val. SpaceXI's Gro
4.6 scores a 61 on the artificial
analysis intelligence index, joining the
frontier in line with 56 soul with
standout agentic performance at lower
costs. 46 gains five points over 45 in
the intelligence index just over one
month after its release. That's a
five-point improvement in a few months
and 23 points compared to 43, which is
really not that long ago. The pace of
releases from these guys has been nuts
lately. And it looks like that pace will
maintain because Elon's already tweeting
saying that 4.7 is significantly better
than 4.6 and it should be ready in 3 to
4 weeks. The initial training is
complete and they're adding a massive
amount of SpaceX company data in
supplementary training and they're
expecting it to be quote something
special. Very interesting. This is the
jump where they are now in the frontier
range. comparable to the other frontier
models according to artificial analysis.
Much stronger agentic performance
especially with things like GDP val
frontier level intelligence at lower
costs because again it's way cheaper.
It's $2 per mill in and $6 per mill out
which is 60% below Opus 5's price and
even more than 56 souls price costs 84
cents per task which is similar to Kimmy
K3 with slightly higher intelligence
which gives it a really good place on
the Pareto Frontier. The cost versus how
smart it is. Cost per task is right
there with Kimmy K3. Slightly more
expensive than Gemini 36 Flash and
Terra. Slightly less expensive than 56
Soul. Comically cheaper than Opus 5 and
Fable 5 for realworld use. That is $3.14
for Fable and 84 cents for Gro 46. This
is more than double the price that Grock
45 was though, which is the important
detail they seem to be hiding. I'm
guessing this price gap is a token
difference between 46 and 45, which we
can confirm here. Yep. I was so pumped
that Grock 45 was as efficient as it
was. Seeing them lose that with 46 is a
little sad. They've increased the number
of tokens per run by over 30% here.
Previously, they were one of the few
major providers that had a model that
was more token efficient than 56 Soul.
Now, they're not. Now they are still
relatively efficient, like comparable or
better than Fable efficiency-wise, but
it's no longer that magic efficient
small fast cheap model that I wanted it
to be. This is something I've actually
felt in my own usage because I noticed
that the Gro 45 test when I did them ran
almost immediately, like creepily fast.
46 is back to the usual send off the
prompt and then go do something else
that I'm used to from Frontier. They
talk about their AA briefcase bench
where it is now Fable 5 tier. Context
window is the same at 500k tokens. The
price is the same as well. They did
increase the price of cash reads.
Previously it was only 3 cents per
million tokens red. Now it's 5 cents per
million tokens red for the cash. That
might also be affecting price here. Some
of these jumps are nuts. Like I know I
don't like GDP Val, but this is one of
the biggest jumps I've ever seen in that
benchmark. Like that's crazy.
It's clearly much better at these long
running things, but it also thinks a lot
more, too. And again, as I said, it is
now much more expensive, but it's also
slightly smarter. It's got a really good
score here in the cost versus
intelligence, but this does yank them
out of this golden area of like the
green corner where it is the cheapest
and the smartest and forces it out over
to be on the right side of the median
for cost on the log scale, which is sad.
They're clearly trying to rush their way
to the frontier. So now it's time to see
how it does with UI. I have other things
I've been working on with it in the
background, too. But I am very curious
how its design work is because it's a
different model, and it is nice to see
the different ways different models
handle design. As I mentioned in my Muse
Spark video, I was actually surprised at
how different the Muse models felt for
design work than the usual anthropic and
OpenAI slop. I've experienced less of
that with Grock models where they feel
more like a dumbed down version
design-wise. So, let's see what they
cooked. I'll start with none of the
skills on and just go through these
Grock 46 outputs. Shout out to Dra for
making the site. By the way, the
witchai.dev is super handy for this type
of video. Here's the first design that
it made. I don't like the noise pattern,
and I really don't like the way that the
text is so sharp with the blurryish
background. It just doesn't feel
cohesive. This feels like old era AI
Slavic. This is GPT 5.0 is how this
feels to me. The purple in particular is
This one has a nice little mock app in
the corner, which is fine. The coloring
is okay. Too many cards. Too old school
like Tailwind Templaty. This is the
usual brutalism that we see from a lot
of these. And this is a boring centered
page. Yeah, none of these impress me. I
did happen to notice that when you use
the design skill from Claude, it can and
do better. Some of them are sinful and
disgusting like whatever the this
is. Some of them have okay ideas on them
like this with the floating card things
when you hover. Yeah, could be tidied
up. This is sinful. This is sinful. Not
digging this. And just for reference,
here is Fable doing similar designs.
Much much better, like comically. So, I
hate that the terminal ones, but other
than that, or even like Opus does a much
better job here with these, or even 56
Soul, as cringe as it often is, is
meaningfully better at this type of
thing. Like, this is a much nicer
version of this style of page. So, not
great at design at all. I will show the
real work I was doing with it earlier,
but I first want to play with these fish
slop ports because they're fun and I'm
curious how it did. Those who aren't
familiar, I made a game at the end of
last year with Opus called Fish Slop
with Opus 45. Never got very far with
it. It was meant to be an insane
aquarium clone. So, I like throwing
models at the old codebase and say,
"Rebuild this. You can reuse assets or
whatever, but from scratch, make a new
version of this game using that codebase
as a reference." And it's a fun way to
see how well it like understands a lot
of end toend
work in a complex thing like the
relations between systems in a game. I'm
already noticing that like the controls
are uh bad. The sizing of elements is
all wrong. Like these fish pellets are
too big and the fish are too small. The
fact that the thing you had highlighted
on the bottom here stays highlighted
when you move off it is really bad.
Yeah, the movement just feels so not
great. The pace of the game seems wrong,
too, where things don't like grow fast
enough and you don't get to the next
stage quick enough. Yeah, this just
doesn't feel great. This is like Kimmy
K3 did a meaningfully better job with
this than Gro 46 seems to be. But this
is the easy version. The hard version is
3D. This is what I got when I first
opened the 3D version of the game. A
black screen. This is the first new
model I have pointed at the task of a
make the game 3D and had it just
outright fail. All the others might have
had bugs and issues with like the
movement or the mechanics or the models
they created in the 3D space. All of
them opened with a 3D environment. Even
like GBT4 was able to get that far. This
is the first one that just outright
failed to do the 3D port. I did tell the
model that it failed. I even gave it a
picture which showed me how nice the
rendering of pictures and things is
inside of the Gro CLI, which like the
Grock build CLI is actually genuinely
really nice. Well, I told it to fix it
with the screenshot. It appears that it
did. So, let's dive in and see. Oh god,
this is so cringe. They got the
directions wrong. So, up and down is
right, but left and right are inverted.
I don't know how you get right and left
inverted, but they did. The placement of
everything at the ground is entirely
wrong. Like, comically so. God, this is
very bad and broken. [clears throat]
The models for the fish are This
is the worst 3D pass I've seen on this
in a long time. Actually, this is like
last year I would have expected quality
like this. Just for reference, here's
the version that Muse12 made for way
cheaper. It got some directional issues
like up and down tilt it, but like it
functions. The movement's way better.
The core mechanics work a lot better. If
we want to go to real frontier, like I
don't know, uh, Kimmy, look at how much
further along this is. The modeling of
the things on the ground. The lighting's
a little broken, but like
this is a generational gap. This is an
open weight model, by the way. Like, you
can go download the weights online and
have a thing that is this much better at
3D. And the fish model is so much
better. It's actually the best 3D
modeling I've seen any of the models do
so far. So yeah, there's your
comparison. Grock is not even
unimpressive. It's like last generation.
It's so far ahead here. But then there
is realworld work. And for this I was
quite kind to Gro 45. I found it
surprisingly capable of doing like real
work where it had to touch different
things that were complex and stay on
task for longer running things. Well,
take a look at the security audit it
just did. It found a previous audit that
I did in the git history. So, that's
cool. So, it's already strong. The O
broker, the worker isolation storage
control plane will get us burned. Give
us advice for locking endpoints. Put it
on the public suffix list. I already was
planning on that. like bed chrome on
every capsule page. Finish the admin
trust and safety UI. Stop putting
identity tokens in URLs or a wider
launch. Refuse sandbox off flags in
production. Yeah, did a decent audit
here. Didn't get me any errors or
problems when it did it. There are other
deeper things it didn't find, but that's
acceptable. This is the more interesting
one I wanted to take a look at with
y'all. I asked Grock 46 to take a look
at how we have T3 code implementing
cursor because the current build is
using the outofdate ACP adapter for the
cursor CLI. They want us to move to the
SDK. I believe Julius is work on the new
orchestrator includes that. But I wanted
to see what it would do. So I asked it
to go through and audit things and see
how we would migrate from ACP to the
Cursor SDK. For long lived gooey hosts
that need a real model catalog, resume,
cancel images, and usage. The SDK is the
one that cursor is actually maintaining.
Yep. So why is ACP the wrong host
contract? Then built in fallbacks. Yeah,
these are all real problems that we've
had with the current cursor bindings. So
it was correct there. Official docs
license anywhere to not MIT. That is
annoying, but that's fine.
The SDK doesn't offer guey level approve
or deny. That's annoying, but we auto
run anyways for most things. Blocking
ask questions. Uh the SDK will not allow
that. and reusing agent login. That will
not work. So, we have to implement our
own login, which is annoying. You know
what I will do? I don't feel like
reading gro slop. We'll have 56 soul
share its thoughts momentarily. Plan
approval could not be completed because
the client disconnected. Plan mode
remains active. Okay, so it put itself
in plan mode. That is obnoxious. Did it
put the plan in here anywhere? It
didn't. That is very annoying. Yeah, the
Grock implementation needs a little bit
of work as well. I will put some time in
in the near future. I have to go turn on
the legacy plan mode for this.
While those plans are being audited,
I'll show off a little bit of it
behaving how I wanted it to. Here I
asked it, what gaps exist in T3 Code's
implementation of Grock build compared
to other harnesses. It went through it
made a plan. I don't know where it went
in the UI. There there's some weirdness
in the events that the Grock build like
ACP sends out. So, we have to put more
work into how we clean those events up.
But I asked at the end here, I wanted to
see it differently. So I just said HTML
which triggers my HTML skill and it
responded there. The UI collapsed it
again because their events are broken.
Uh I gave it some feedback on the plan.
Had it update then asked it to build the
whole plan file a PR and babysit which
it did. I can go open it on GitHub and
we can see what things thought here. We
got a bunch of feedback from Macroscope
on the changes. It tore things to shreds
here, but it seems like it followed my
instructions really well. You'll see
here that it used my account to reply to
things, but I have a little call out in
a skill I made for PR comments where I
tell every agent to open its comments
with note which model is responding on
behalf of Theo. This is not a small
change either. This is a thousandline PR
that it made to address a bunch of
different gaps in our coverage for Grock
Build. And I told it after it makes the
PR to babysit it once it was filed. Then
I told it to do one of my favorite
things, which is to tell the agent to
look through my actual history on my
actual computer and find gaps in what
events occurred and what we actually
process. Just told it to make a separate
PR stacked on the one that we currently
have to do all these additional changes.
This is a complex issue because it needs
to know how to like figure out and stack
PRs. It needs to keep track of the gap
between the old PR and the new one. All
the investigatory work it just did in
the context of how it applies with my
history on this machine as well as all
the other skills and things that it
pulled in. It's not easy for models to
deal with all of this like competing
context and stay on track. This is one
of the things again I thought Grock 45
did surprisingly well. It was able to
take multi-step unrelated work and do it
all cohesively and coherently in one
thread. So, we shall see how it handles
this here. I just put together a really
rough like how I would score Grock
compared to Fable and Soul for like
different categories. I think about
things like cost where Fable's way too
expensive. Soul's significantly better.
Probably do a little lower there, but
like reasonably priced. And then Grock
45 way better on cost. Intelligence,
Fable is the best right now. Soul is
surprisingly good, but not quite as
good. Grock is not even this. I would
put a little lower there. Then you have
speed where fable is very slow. Soul is
meaningfully faster simply because it's
so much more efficient. And then Grock
45, especially on the fast mode, flew.
It was super fast and really nice. Then
with thoroughess, Fable, I find, isn't
quite as thorough as it should be. It
misses things here and there, but it's
it's thoughtful, not thorough. Soul is
incredibly thorough. It checks every
single thing. It touches every single
edge, and it writes too much code as a
result. 45 just did not have any of
that. And then with orchestration
capabilities, Fable was the best. Soul
is surprisingly close in Grock 45.
Better than most, but still not quite
there. So, how does this all compare to
our new model Grock 46? Sadly, cost is a
regression. I'd say this is like a
sixish now. Intelligence seems like a
meaningful bump. I haven't seen too much
of it being way smarter, but I can
confidently say it's probably
meaningfully better. We'll give it a 6.5
there. Speed is where I'm seeing one of
the biggest regressions now where I
would put it at like a five and a half
at best right now. Thoroughess, it is
better, but it still misses things. I'll
bump it slightly there. In orchestration
capabilities, I did see it using sub
agents decently well. I'll give it a 6.5
here. Why not? The issue is the only
things I would consider Grock previously
to be a leader in were the speed and the
cost. And we saw regressions in both of
those categories with this release. The
cost went up, not per token, but since
the token efficiency went down, it's now
more expensive. And the speed went down
because it's less token efficient. So,
it's generating more tokens and it takes
longer. These two changes make Grock
much, much less interesting to me. But
the speed at which the intelligence is
growing suggests that Gro 47 will do a
lot more in all of these other
categories to catch up. The thing I
liked Grock 45 for wasn't that it was as
good as Frontier at specific things.
It's that the speed and the cost were
far enough away from Frontier that it
felt uniquely useful in those ways. And
I feel like we are losing that with
Grock 46 a little bit. I think that
summarizes my thoughts on this release.
Is it my favorite model? No. Am I going
to use it a whole lot after this video?
Probably not. But does it have me
excited for Grock 4.7? Absolutely yes.
It makes a lot of sense why Elon's
talking more about 47 than 46 right now.
The future seems very clear and it seems
like a future where Grock catches up
very fast and I'm honestly pretty
excited for that. We need more
competition. We need more good models
and we need more people fighting to make
them cheaper. Let me know how y'all
feel. Is this an exciting release for
you or are you going to just ignore it?
Let me know in the comments. And until
next time, peace nerds.
Ask follow-up questions or revisit key timestamps.
The video provides a detailed analysis of the newly released Grok 4.6 model. While the model shows improvements in intelligence, agentic performance, and thoroughness, the reviewer highlights significant regressions in speed and cost-efficiency, which were previously Grok's key competitive advantages. The video includes practical testing through game generation, UI design evaluation, and real-world software engineering tasks, concluding that while 4.6 is a step forward, the anticipation for Grok 4.7 is much higher due to the rapid pace of development.
Videos recently processed by our community