Grok 4.6 is Fable now
542 segments
So, huge news. Grok 4.6 just got
released as I'm recording this. But,
here's like the big headline. xAI caught
up with Anthropic. Now, the benchmarks
for Grok 4.6 High are showing that
they're equivalent or or close to Fable
5 Max. So, this is the first Grok
release that is at the time of release
is just neck to neck with the best
OpenAI model and the best Anthropic
model. 4.5 closed the distance and 4.6
but I'm right there neck to neck. Now,
of course, this was foretold, so to
speak. This was kind of predictable.
xAI's or Spec Space xAI's purchase of
Cursor probably had a lot to do with
this. In fact, as they say here, the
biggest difference between 4.5 and 4.6
was a much longer supplemental training
run. So, we believe that 4.5 and 4.6,
they are kind of the the V9 model or
base as it's referred to. It's a 1.5
trillion parameter base model and 4.6
had a much longer supplemental training
run versus 4.5. So, they used a lot of
the data from Cursor. And as they say
here, curated model generated data for
reasoning and advanced technical
concepts, high-quality engineering data,
and an improved optimizer and training
recipe. This produced a stronger
foundation for the supervised
fine-tuning and reinforcement learning
stages that followed. Now, of course,
most of us at this point view benchmark
numbers with some suspicion cuz we've
seen some places before kind of like
just trying to game the system, get high
scores on and those benchmarks, but the
models themselves kind of failed basic
tasks when people actually got their
hands on them. Is that the case here? I
don't think so. I've only played around
this model for a few hours at this point
and I got to say from my first initial
experiences just right out of the gate,
it's a beast. One of the first prompts
that I did was I asked to kind of do a
clone of Portal 2. So, kind of being
able to fire different portals, blue and
whatever, orange, right? You walk into
one, you walk out the other. Something
that a lot of the previous models up to
very recently just weren't able to do. I
used Grok Build to do this. It worked
probably for two to three hours or
thereabouts, and I don't want to say one
shot cuz I had a few suggestions and I I
steered it a little bit during the
process, including asking it to add some
sound effects and voice overs with 11
Labs, but pretty much yeah, it one shot
one sort of level from Portal 2. Created
all the mechanics, created all the
models, which you're going to see in
this demo in just a second. You know,
all the three objects it's created. The
conservation of momentum, the ability to
see through the portals, the ability
ability to see the character on the
other side of the portal. It's important
to understand that like this is not a
low-end model. This feels like a
high-end smart model. It created the
entire room in one shot, including the
puzzle that you're supposed to solve for
multiple elements. So, take a look. So,
this is chamber 07. Grok 4.6 build,
basically a replica of Portal 2.
>> Chamber 07, dual gate certification.
Please proceed.
>> And it's looking pretty good. We've got
reflections, we got lighting. Test
chamber 007. Blue gate.
>> Blue gate established.
>> Select the blue gate, and here is the
>> Both gates are live.
>> Let's see here. Oh, wow. All right. So,
we have a character model with
reflections. We're carrying that gun.
Uh that is pretty nifty. Okay, let's go
through here. Oh, wow. Okay, that's
pretty smooth, I got to say.
All right. So, if we wanted to go
through here and have my character here.
Wow, okay. So, I can see myself there.
So, what happens if I walk through it?
Oh, hey. That is
>> one, emerge from the other. Momentum is
conserved.
>> is conserved as we go through it.
>> Orange gate established.
>> Oh, look at that. Okay, so that thing
should
I did not expect that to happen.
I mean, it's supposed to happen. I just
didn't expect it to.
Okay, so this is our
>> Wait, you mad at Cuba quiet. Do not
introduce it to the acid.
>> Do not introduce it to the acid. The the
green bubbly stuff is the acid. Okay, so
basically I guess what we want to do is
slide this kind of like what is that
game where you slide the thing across
the ice? Loogie? Loogie? Something like
that. Okay, we go through here and as it
slides
>> Yes.
>> Hey, that is pretty good. I got to say.
I'm statistically a success apparently.
Well, thank you. So, there's a lot of
really incredible and amazing things
there. It's very impressive. Of course,
keep in mind this was just the first I'm
pretty sure this is the first prompt I I
threw its way. I was trying to see if it
works in Grok bot, which is the other
big release from yesterday. But, I
wasn't able to confirm that it was yet
available in Grok bot. So, I did Grok
build instead. That's where it built all
of this. And it was built with Grok 4.6
on the high settings. Later I realized
there's an extra high setting. So, I I
do want to test around it with that, you
know, turning up to 11. But, this was
Grok 4.6 high and this is very similar
to what I would expect Fable 5 to come
up with on on the first attempt. So,
again, Fable 5 level model from xAI.
But, wait, there's more because pricing
starts at $2 per million input tokens
and $6 per million output tokens.
Additionally, there's a fast variant,
which is twice the price. My next test
that I'm running within I'll post it
more about this later. So, I'm trying to
get it to build from the ground up fresh
without reusing any of the existing code
cuz I tried this before for a different
model. So, this is a starting fresh.
Here I'm using cursor by the way to to
do this. Important point on cursor and
Grok build because right now during
launch week, so that's right now, you
have double the usage for these new
models if you're using cursor or Grok
build. So, if you have those, do take
advantage. And so, here in cursor I used
a Grok 4.6 extra high fast. So, again,
fast is that model that it's twice the
speed, but it uses up twice the usage.
Right now, you get, you know, twice the
the usage, so it kind of evens out. But,
here I'm trying to get it to do a fully
AI automated run kind of streamer. So,
in this case, it's going to be playing
Pokémon Red and playing the game using
voice to narrate what it's doing.
Probably a HeyGen avatar kind of a video
avatar that's going to explain
everything while it's playing the game.
It's also going to be able to talk to
chat. And so, for that, I'm making sure
that it adds some sort of a censorship
so that people don't jailbreak it, which
could be bad. So, we'll see. We'll see
if this model survives contact with, you
know, you all people. So far, the
results are good. So, here's kind of
what we have so far. Just keep in mind,
this is a bare-bones prototype. If I
click start game, so notice it loads up
Pokémon Red. It goes quickly through the
initial screen, and I know it's got past
this right now. And so, right now, it
doesn't have all the logic and all that
stuff quite yet. Oh, the last time it
did get out of the house. Right now, it
seems to be stuck on something. But, the
point is, on its first attempt, it was
able to take a Pokémon Red emulator,
hook it up to make sure that it's able
to have enough of a harness to control
the game. It created a fake chat. It
created text-to-speech. So, that little
model that's over there, as it's
talking, that gets gets a real
voiceover. And so, all this was built
within the first hour or two of Grok 4.6
going live. So, again, I'm not claiming
anything about this model yet. I'm
saying just straight out of the gate,
this thing is feeling pretty beefy. So,
looking at these numbers, I'm not
doubting them. I think this is kind of
on the level that it's going to operate
at. This is Cognition posting about
Frontier Code 1.1. So, notice Grok 4.6
is right behind the Fable 5 and Opus 5
series of models. It's it's up there.
What caught my eyes, Elon responded to a
few of these messages saying that Grok
4.7 will exceed all current models. So,
the next release, 4.7, which we're
expecting in 3 to 4 weeks, will be
number one on the [clears throat] chart
assuming no one else releases new models
which is not going to happen. But as of
right now if everybody else was stuck in
time for three to four weeks, it would
be number one on the leader boards. And
for that model they're adding a massive
amount of SpaceX company data in
supplemental training. Okay, so so far
the only thing that's kind of like is
confirmed so far is what we know about
Grok 4.6. Some of the things that are
coming in the future, these are
speculations of certain announcements by
Elon, some of the kind of rumors that
are floating around. But we're expecting
4.7 to be a 2.1 trillion parameter
model. Again, that's not verified by
official sources as far as I know, but
here is the important thing to
understand. So, 4.5 wasn't that long
ago. Now we got hit with 4.6. We have
4.7 coming out in say three to four
weeks, and Grok 5 is targeted to be
released before the end of the year. So,
as you can imagine, that's a kind of a
crazy release schedule. Why is this
happening? xAI reorganized engineering
to ship major updates every two to three
weeks, it seems like. So, kind of keep
all this in mind, right? So, they caught
up, they got to Feat of Strength level,
and they are just gunning to release,
release, release. The pricing is $2 per
million input, $6 per million output, so
it's a half the price of comparable
OpenAI and Anthropic frontier models.
And we also have this SpaceX company
data that's going to be used in
supplemental training, but there's more,
and this is kind of the the other side
of the coin. This is the other big
thing. So, this is Grok Bot. And if I
understand correctly, 4.6 isn't running
this quite yet. I I did a few tests
here. I I don't think 4.6 went live on
this thing yet. Now, as you recall from
before, we talked about the different
benchmarks. One of them was GPT Val. So,
GPT Val takes real jobs, real projects
that's done by real people out there
that work in specific industries, and
these models are asked to complete those
projects. And then real humans sort of
like verify. So, these humans have a lot
of background working in those
industries, and the industries are from
engineering to like AutoCAD to some
hospitality and travel industry like a
travel agent. There's finance, video
editing, like it's a broad range. It
really represents a large part of the
job market. And the 4.6 model beats out,
you know, Fable 5 and everything else.
Why is that important? Because it goes
hand in hand with this with Grok Bot.
So, Grok Bot is xAI's always on, always
online, always available agent. So, it's
on desktop apps. I believe it's on
Linux, Mac OS, it's on Windows with
Android versions coming soon. And the
whole pitch is these are your AI
teammates. So, this is not like a
chatbot. These This is a team of agents
that work for you. One of the first
things that it kind of suggested you do
is you create, you know, let's say a
chief of staff. That's the role that it
kind of recommends. And so, I I jump in
here and I give it this this task.
Initially, I was want to see how far I
could take that AI streamer. So, I told
it, "Do this. Create whatever agents you
need." It does. So, everything here,
most of these agents except one, it
created automatically by itself. Then it
begins running the various tasks and
handing off the sort of the delegates
work to those sub-agents. So, you're not
talking to the million of AI agents
working in the background. You're
talking to your chief of staff or
whatever you want to call it. It
delegates, it thinks about it, it asks
you questions, but those questions are
might be the ones that are surfaced by
the agents. You're You have one point of
contact. Or you can have multiple points
of contact, but the point is that you're
not talking to every independent agent
on here. If you're wondering, "Well, how
could this be always on? Isn't this on
your desktop?" Well, no. Each of these
agents, Grok Bot, it has its own virtual
machine in the cloud that it can use 24
hours a day. You can actually see the
screen here. So, here it is. This is
what's available for it to use. So, you
can open up Chrome, terminal, whatever
you want. What's interesting about this
is there's actually a button to teach it
a task. This is kind of brilliant
because there's going to be a lot of
things that these AI agents have to
navigate that they're going to struggle
with. I had one of my AI agents not
through Grok, it was I think it was
OpenAI or Claude, I I forget which, but
I had it go online and go through an
online messaging app like WhatsApp or
Slack or what whatever it was to try to
see if somebody had mentioned something.
While I was doing that, it accidentally
sent a message. Now, this wasn't a big
deal, it caught the error immediately,
deleted it, and the message was just the
command it was run to to try to search
through all the messages by keyword. So,
it was something like search whatever
keyword, right? So, it wasn't a big
deal, but it could be bad in in a
different circumstance, right? Because I
didn't ask it to message people, but
with these teach a task, I can actually
show it, it gets recorded, and then
Grok, these bots, they're able to
execute those skills. Now, of course,
this is great if you're trying to work
with it, and it's kind of great for XAI
because it's basically a lot of people
kind of contributing data to these
models. And by the way, of course, you
have the choice to to opt out or not, so
that's entirely up to you, but for the
people that choose to opt in, of course,
their data can be used to further
improve these models. One other agent
that I built was my X news real-time
person that went through and basically
checks real-time news on X. So, X slash
Twitter, it goes out there and it starts
searching for me. It's super fast at
finding those news, so I just ask a
question, you know, like, "Did So, I
just ask it questions like, "Did anybody
talk about this? Did Elon say something
about this? Did somebody blah blah what
whatever you want?" It goes through, it
searches it, and it just explains it to
you. So, one of the things I asked it
is, "Grok 4.6, is it live on Grok bot?
Is this where we're using it?" And I
asked the other model, too, but it
wasn't able to say which model it's
running on, so I don't think as of that
time Grok 4.6 was live here yet as far
as I know. So, it found the official
posts, it found some comments by Elon,
and it said that so we're not 100% sure,
it's likely coming very very soon. Do
you want me to watch X for when they
turn it on?" I said, "Yes." It
automatically created routines and it's
going to keep checking, you know, three
times a day and for the next 10 days,
after which he'll notify me one way or
another. So, this is kind of awesome cuz
now I have this bot whose job it is to
keep searching X and when we get new
updates, it notifies me. Since this is
always-on, on 24/7, it's not like when I
shut down my computer, it goes to sleep
and forgets it. This will notify me for
sure. So, the big point here is that
these are meant to be more like your
colleagues and coworkers that you
delegate work to. So, this is kind of
digital labor, right? So, this has been
positioned by XAI as digital labor as
actual workers. So, a company that has
access to this basically is able to get
a lot of the cognitive labor handed off
to these agents that are online 24/7,
that are able to do research, you know,
when you're sleeping, put together
projects, hand them back to you. And
this is kind of I would say how they
design it is innovative. Now, not like
where they're breaking new frontiers or
whatever, but there's a lot of little
things here and there that just seem
like when you start using you're like,
"Oh, yeah, this feels right." One is
kind of they pop up with these boxes
more and more where you get to choose
your own answer. Your initial impression
might be like, "Oh, that's for, you
know, people that are just starting out.
Maybe they don't quite know exactly what
they're trying to do." But working with
this, I realized that there's a lot when
you try to go fast, this is very
helpful. Cuz when you give it a project,
it might have a bunch of questions for
you. Each time it asks you a question,
you have to kind of sit there and and
read it and think about it and then type
out an answer or just speak an answer.
There's been a number of times when it
shows me something like this and just at
a glance, I kind of know, "Oh, yeah, I
know exactly which one of those options
I want." Sometimes you don't even have
to really consciously read the question.
You get what it's asking you. So, at
some point it asked me this question,
right? And the answers were like, "Add
Heygen or just use text-to-speech." Like
I as soon as I glanced at it, I knew
exactly what it was asking without
needing to read the question. It was
going, "Do you want to build out the
video avatar or just leave that for
later?" Like I knew and I was like,
"Yeah, go ahead and build that out."
Things like this might seem tiny, but it
makes things go a lot faster. It's less
of a cognitive load, right? You're not
wasting cycles having to read the I mean
in your own brain, so to speak, reading
through everything. You're kind of like
cuz you know what you want. The model
doesn't. And so, any little polish we
can add to make that transfer of
knowledge faster and smoother and takes
less effort, it's a big deal, especially
as more and more work begins to get done
like this. So, love this so far. The X
News AI agent that adds the routine and
then just says, "Yep, I'll go ahead and
I'll start checking it three times a day
for the next 10 days." I didn't ask it
for that. It says here, "Do you want me
to watch X for when they flip it?"
Right? When it goes on. Yes. It's like,
"All right, got it." And it adds the
routine. This is beautiful. This is one
of the things where I like this so much
better than Fable 5, as much as I love
Fable 5, but the way that it talks half
of the time, half the answers, I'm like,
"What are you saying?" I mean, at this
point there's like entire memes by it.
It's got this weird way of speaking, you
know, and yeah, it's kind of funny, but
after a while it kind of gets old. It's
just like just tell me what I need to
know. What are we talking about? So far,
going through these and working with
Grok, it's just so much faster. You just
kind of fly through it. So, definitely
encourage everybody to download this.
It's super simple to set up. When you do
so, create your kind of chief of staff
or whatever name you want to use for it,
then pin it to the top, and then try to
use this as much as possible, and have
it create its own structure underneath
it so that you're sort of talking to
just this thing. And then and maybe
there's other sort of side agents like
this X News one that you have off to the
side that that it's its own thing. I
don't need it going through the chief of
staff, but for a lot of this work, have
one agent in charge of it. And of
course, this is easily cross-device,
right? So, I can go from here to using
it on my phone to wherever. I'm not
relying on the computer being on at any
given moment cuz it has its own
computer. So, between Grok bot and Grok
4.6 coming out and the new models that
are coming out, 4.7 and Grok 5, I mean,
those are not out yet, but so far it
looks like XAI is just in a great great
position to start winning. Check it out
within Grok bot if you have access to
Cursor. You get some double the usage
during launch week or if you're using it
with Grok build, same thing. But, I got
to say this is
this is kind of exciting. It seems like
Google is lagging slightly behind, but
we still have, you know, multiple
frontier models because here comes xAI
to fill in Google's spot. So, congrats
to xAI, congrats to Cursor. I was always
very impressed with Cursor and all the
stuff that they've been able to do. So,
them joining forces with xAI, which has
just incredible amounts of compute and
and talent and capital in order to
really put these two things together. I
feel like they're going to have a lot of
success in the future. But, let me know
what you think. If you've had a chance
to do more testing of Grok 4.6 more than
I have so far, let me know. What do you
think of it? Is it, in your opinion, as
good as Fable 5 and GPT 5.6? Will you be
using Grok bot? Let me know in the
comments. If you made this far, thank
you so much for watching. My name is Wes
Roth. I will see you in the next one.
Ask follow-up questions or revisit key timestamps.
This video details the launch of xAI's new model, Grok 4.6, highlighting that it has reached parity with top frontier models like Fable 5. The presenter demonstrates Grok 4.6's capabilities by building a playable Portal 2 replica and discusses the introduction of 'Grok Bot', an agent-based platform designed for continuous, autonomous task management. The video also covers the aggressive release schedule xAI has adopted, including upcoming versions 4.7 and 5, and the advantages of integrating with Cursor for development workflows.
Videos recently processed by our community