Grok 4.6 is Claude Fable 5, but dirt cheap
352 segments
Grok 4.6 just dropped, and it's
official. SpaceX AI is now a frontier AI
company. Elon claims you can now get
above Chad GPT 5.6 and Opus 5 quality
for a fraction of the price and much,
much faster. Well, I put it to the test,
and today you are going to find out if
those claims are true. You're also going
to find out if you should be switching
to Grok 4.6, how to use Grok 4.6, and
which situations you should be using the
model. Now, let's lock in and get into
it. So, here are the benchmarks they're
giving. Unfortunately, I think all
benchmarks are fake, so we're not going
to spend much time here. Basically, what
they're claiming though is it's just as
good as Opus 5 and 5.6 sole for a
fraction of the price. They're not
claiming it's way better or it's the
best model ever made. They're saying
it's good as frontier, but you're
getting it for way better costs, which
is a good thing because Fable and Opus
and 5.6 sole are very, very, very
expensive models. The two places you can
be using it right now are Cursor and
Grok Build. I lean towards using Cursor.
It is a much more full-featured AI vibe
coding experience. It's basically kind
of like a stripped-down version of the
Chad GPT app, but it's very, very good.
Grok Build is solid, but it's just a
CLI. There's no interesting harness
features I think you'd come to expect
from harnesses in August 2026. So, I
lean much more toward using Cursor
instead of Grok Build if you want to use
this model. Now, let's talk about if
they came through on their claims. I ran
this model through the world-famous Alex
Finn benchmarking test. For those who
don't know, it is five different
benchmarks. I put it against Chad GPT
5.6 sole. I put it against Fable. Here
are the results. It beat 5.6 sole in my
benchmark. So, it went through these
five different tests. It did a whole
scavenger hunt for code on the internet.
It did a whole debugging exercise, it
built a bunch of simulators. Basically,
what it came down to is it scored better
than 5 6 soul. It did it in
significantly less time, did it in 18
minutes rather than 24 minutes with GPT,
and it did it for basically a third of
the price. Now, let's go through those
actual results here. Well, starting off
with the first one, which is a roller
coaster simulator test. Over on the
right is 5 6, over on the left is Grok.
Let's open this up and see what it looks
like. As you can see, it built a very
nice roller coaster simulator. That is
That is the I think the biggest roller
coaster out of all the simulators we've
done so far. Let's hit ride and see what
this is like. This looks great. The sky
looks amazing. I love the kind of
twilight colors in the sky. Clouds look
nice, and as you can see, the roller
coaster is looking good. If we go over
to chat GPT and I expand this, I'd say
the fidelity is not quite as good, it's
not quite as pleasant on the eyes. And
if we go ride here,
uh it looks it looks fine. It looks all
right. It did it for a fraction of the
price. It actually took a little bit
longer on this test, but I think the
results came out well, and you spent way
less money on it. Here are a couple of
the other tests against chat GPT, a
recreation of the Apple website. So, we
hand both models the Apple website, we
say, "Hey, recreate it as close to
graphical fidelity as you possibly can."
This is what the Apple website looked
like today. As you can see, there's the
phone, the air, uh all the other
devices, Ted Lasso, uh Sabrina
Carpenter. Let's see the recreation from
Grok. Not too bad, not too bad.
Obviously, it doesn't look nearly as
good, but it can't use graphic
generation, it has to recreate all the
graphics itself. Honestly, not too bad.
And then if we go over to GPT, GPT looks
all right. I like the header better, but
the phones look all kind of messed up
from there. That doesn't even look like
a laptop. Uh all the things down here
don't look too great. So, honestly, it
it it edged it out. Did it actually a
little bit longer, too, but much
cheaper. Other than that, we have an
agent test where we give it tons and
tons of files. Then we basically have it
go on a scavenger hunt, has to use
different tools to find things in CSVs
and PDFs and different files. It
actually dominated GPT. It did it in a
looks like a sixth of the time for about
a sixth of the cost, and it found and it
used the tools better. GPT couldn't even
find uh two of the scavenger hunts. Bug
fixing, we handed a open-source library
from GitHub, a a recent one. We see how
many bugs it can find. Both of them
found all 13 bugs in the code. Grok did
it faster, did it for about half the
cost. GPT took about a minute longer for
about a dollar more. And overall, it
came out that Grok beat GPT pretty
handily, which is pretty shocking. But
how did it handle against Fable? Let's
check that out. So, Grok actually ends
it up beating Fable 5. Now, I will say
this, the reason why Grok won is because
Fable 5 refused to do the scavenger hunt
because of its safety guardrails. If you
take out Grok's score from here, it is
about even when you take out Grok's
score from that one. But at the same
time, Fable's not letting you do
everything. Grok lets you do everything.
Grok beat it at about a tenth of the
cost, did it in a minute quicker. Let's
check out Fable beat it on pixel
perfect, the Apple recreation. Let's see
that real quick. Here is the Fable
recreation. Some of these actually look
really, really strong, some of the
devices. Fable actually did this one
faster, but again, four times the price.
Bug finding, Fable was four times the
cost at double the time. And so, Grok
ended up beating Fable 5, which is
really, really impressive. Now, I did
some other tests, too. I put it up
against Opus, Fable, and 5.6 on some of
these tests. So, here's a flight
simulator. I also put Grok against the
test with Opus, as well as Fable and
Soul and a few other tests here. So,
first we have kind of a Flappy Birds
test here. Here is the Grok one.
Actually looks really, really nice.
Ooh, next levels. You get the stars, you
get everything. I like that. It does It
runs the simulation a little bit quick
there. Oh, and things are just
collapsing. All right, so the physics
are a little bit off. Let's see here.
Graphics looks great though. Let's see
what Opus 5 looks like here. Opus 5
looks pretty good, too.
Physics actually work well. I like that
a lot. There doesn't appear to be second
levels though. You just do level Oh,
there we go. Next level. Here we go. All
right, so now they got new things out of
there. Physics are a little off, too, as
well. Just see it collapse before you'd
even do anything. I'd say graphically,
Grok probably outdoes it a little bit.
Gameplay-wise, Opus wins. And then let's
check out Fable 5. Fable 5 appears to be
completely broken. And then ChatGPT 5 6
Soul. This looks really good, too. I
like the graphics. Let's Let's fling
this thing. Uh the physics seem off here
and everything looks kind of messed up.
This is probably the weakest of them
all. So, let's see how much it cost, how
long it took. 9 minutes for Grok. It was
actually longer than 5 6. Yet it is half
the price of 5 6. By far the cheapest
out of all of them. And I think it got
the best results. Fable produced the
worst results and was the most expensive
and took the longest. So, let's revisit
those claims. Let's see what was good
and bad about Grok 4 6. I've been using
all day. I've tested it in Cursor, Grok
build. I showed you all my benchmarks.
Here's what it comes down to. The
positives. All the claims are true. It
is basically just as good as Opus and
ChatGPT 5 6 while being much cheaper and
much faster. Those are all 100% true.
Here's the issue though. The issue is
there's no kind of great general purpose
harness for this model. The strength
right now of ChatGPT is the ChatGPT
desktop harness is incredible. It has
everything. It does coding really well,
kind of like Cursor, but it also does
the general-purpose stuff. It can do
computer use really, really well. It can
do browser use really, really well. It
does everything you need to get all your
knowledge work done in one single place.
Grok doesn't really have that. Grok has
Cursor. Cursor is good. Cursor is a very
good app, but it is really built for
coding. Yes, it can do some more
general-purpose stuff, but it's not as
strong or as intuitive as a lot of stuff
in the ChatGPT app. If you are looking
to do pure coding, if you're all about
vibe coding and that's it, you don't do
kind of AI for all of your knowledge
work, then this actually is probably the
best choice for you. You're going to get
an excellent model at an excellent price
that's super fast, and you can use
Cursor for all your coding. And Cursor
is excellent at coding. The challenge is
it's just not a very great
general-purpose
AI agent harness. If you want to do
knowledge work with Grok, the best
option is Grok Bot. Grok Bot is a
fantastic app. I did a review on it
yesterday. You should check that out if
you haven't yet. But, it's more for your
kind of knowledge work. It's not as good
for coding, and it's a little bit
limited with local computer and browser
use. But, it is very, very good. So,
that is the challenge for Grok right
now. It is a great model. It is cheap,
and it is fast, but it doesn't have that
amazing, incredible harness just yet
like ChatGPT desktop app is. Claude, to
a lesser extent, the Claude desktop app
is really, really good. I wouldn't be
surprised if they completely rebrand
Cursor to like Grok agent or Grok
desktop very, very soon. The Cursor team
built Grok Bot, and they didn't call it
Cursorbot, they called it Grokbot. So, I
wouldn't be surprised if they repurpose
Cursor to be Grok desktop very soon to
be kind of the ChatGPT desktop app for
SpaceX AI. You know, that doesn't even
mention a lot of the other nice-to-have
tools that ChatGPT and Claude has like
ChatGPT voice, Claude voice, two
revolutionary technologies that's kind
of missing from Cursor. The mobile app,
right? Cursor just came out with their
mobile app. It's very good, but it's not
quite the ChatGPT and Claude mobile
apps. So, who is Grok for? Who should be
using it? Well, I think a lot of people
should be using it. I think if your
workflow is primarily vibe coding, this
is kind of the best pure vibe coding
model there is right now. It can build
just as good as the other models, but do
it for way faster, way less cost. If you
are an AI power user and you need voice
while you're on the go so you can talk
to your agent and build things on the go
and you want to be, you know, by the
pool and you want to send a command to
your agent to do and tinker on your
computer and do a bunch of things. Like
you're a power user like that, you need
kind of a power user harness, Cursor's
not quite there yet. ChatGPT is there,
Claude is basically there, but I will
say this, what SpaceX AI has pulled off
over the last several months is
miraculous. Grok was by far way behind
Claude and ChatGPT. Now they are neck
and neck, they are right there. So, that
wouldn't shock me if again Cursor turns
into Grok desktop and then it's just as
good as the other desktop tools and is
just as good of a harness. Grokbot
absolutely amazing. Everyone should be
using Grokbot. I've told you this is
like the best kind of for the normie AI
agent app I've ever used. It's
excellent. That should be used here as
well with Grok. I'd still use ChatGPT if
you do a ton of general purpose
knowledge work and you're a power user
that need AI using every device you
have. Claude, I'm still using Claude
basically just for Fable 5 business
planning. I think from a high-level
business strategy perspective, Fable 5
is still the best model. I don't use
Claude for literally anything else at
all. Claude has basically fallen behind
on all levels. I really feel like this
desktop app hasn't changed in a very
long time. It feels like it's been
absolutely months since anything has
changed in this app. I'm not sure what's
going on with the Claude team. It seems
like releases has slowed down
dramatically over the last few months. I
wonder if it's because of their compute
limitations. They just signed a deal
with SpaceX AI to get a whole bunch of
compute. Maybe they're getting that on
boarded and will be able to release
quicker in the future. I wouldn't be
surprised if they just have an explosion
of releases over the next few weeks. I
wouldn't be surprised we get Fable 5 1
very soon. There's a lot of rumors
around that and I'm sure this whole
competition flips on its head when that
comes out as well. But Grok 4 6 pound
for pound is probably the best vibe
coding model out there right now. If
you're tight on money or you're just
purely about vibe coding, there's no
better model to be using. I do it in
Cursor. Again, get Cursor, use it in
there. It's great inside Cursor. For me
personally, I'll still be using ChatGPT
and Claude as well because I have so
many of those general purpose use cases
I still do with AI agents. I'm going to
be doing a full Grok bot boot camp this
week in the vibe coding academy. Make
sure you sign up for that. Link for
that's down below. It's the number one
AI community on planet Earth. Best
decision you'll ever make joining that.
If you learn anything, make sure to
subscribe and turn on notifications.
Leave a like. So grateful you'd watch
this video. Seriously, thank you so much
and I will see you in the next one.
Ask follow-up questions or revisit key timestamps.
Grok 4.6 has been released as a strong contender in the frontier AI space, offering performance comparable to top models like GPT 5.6 and Claude Opus 5 at a fraction of the cost and speed. The reviewer puts Grok through various benchmarks, including code generation and simulation building, finding it highly effective for coding tasks, particularly within the Cursor environment. While Grok excels in coding and efficiency, it currently lacks the comprehensive general-purpose AI agent ecosystem found in ChatGPT. The video concludes that Grok is the top choice for 'vibe coding,' though power users requiring integrated voice and broad knowledge tools may still rely on other platforms.
Videos recently processed by our community