Grok 4.6 is Actually Good… And Claude Keeps Getting Better
653 segments
This was the biggest week of the year
for Elon Musk and his AI efforts at
SpaceX.
>> [music]
>> They released Grok 4.6, which shows that
they are finally catching up to OpenAI
and Anthropic. They also released Grok
Bot, which is their new super app, which
will rival GPT work and Claude co-work.
And I have a lot of thoughts about this
platform and what makes it so
interesting, which I'll talk about
today. But we have way more to cover. We
will also discuss the latest updates
inside Claude Code and Codex, as well as
the latest DeepSeek V4 model that they
launched today. And we even have some
news from Gemini. This is an
agent-native update where we put the
latest advancements on the frontier of
AI agent platforms and models into
context so that we can actually use them
to improve our business.
Okay, so we have a lot to cover today.
Let's dive straight into the SpaceX
update. So SpaceX said this yesterday,
"Introducing Grok 4.6. It delivers
frontier intelligence and is a
significant improvement over Grok 4.5 at
the same price." And so you'll notice
here that the three areas that they led
were economically valuable work, long
professional tasks, and legal work. And
so if you also notice that it isn't
coding, right? They didn't lead in any
of the coding benchmarks. And I believe
that this is because they are focused on
general agent tasks because this is
simply their priority, which explains
their brand new platform that they
worked on with Cursor, which is called
Grok Bot. Grok Bot is their brand new
super app, which we'll talk about in
just a second, that is focused on
non-coding work. They are trying to get
everyone within a company to interact
with AI agents to help them get work
done. They're focused on knowledge work.
And I still believe that the best and
easiest way to use Grok 4.6, this brand
new model, is directly inside Cursor. As
soon as you update Cursor, it will
default to Grok 4.6 fast, and you can
use it directly inside Cursor. So, I've
not yet had enough time to actually do a
deep test of Grok 4.6, but I do think
there's some interesting use cases to
discuss that people have posted on
Twitter. So, here's DHH. So, Fable, I
think this was last week, one-shotted a
Rust rewrite of the terminal text
effects Python library in 11 million
tokens.
If you don't know what that means,
that's perfectly fine. Fable did a
really, really hard thing. And then,
today, he tweeted that he used SpaceX's
new model, Grok 4.6, with just a couple
of nudges, it was able to repeat this
feat in in about an hour and a half, and
the key takeaway is that it was only
$55.
That was about 1/10 of the cost of Fable
implementation for the same work. So,
take a look at the pricing for these
models. If we look at Grok 4.6 compared
to Opus, Soul, and Fable,
right? If we combine the input and
output prices per 1 million token, we
get $8 for Grok 4.6, $30 for Opus 5, $35
for 5.6 Soul, and $60 for Claude Fable.
And so, that means that Claude Fable 5
is 7.5 times more expensive than Grok
4.6, 5.6 Soul, 4.4 times, and Opus 5,
and Grok 4.6 is better than Opus. It is
straight-up better, and Opus is still
3.75 times more expensive than Grok 4.6.
This is a very good model, and it is a
reasonable price, and it is legitimately
on the frontier. And beneath this
Cognition post, Elon commented, "Grok
4.7 will exceed all current models,
which includes Fable." He said, "That
said, Anthropic is a great company and
will probably release improved models
soon. However, the SpaceX training
corpus is so awesome and unique that I
would be shocked if any model is better
at real-world engineering than 4.7. And
so that was their model release. And if
you watch my channel, you know that I
actually don't dive too deep into model
releases. I care mostly about practical
use cases of AI. Like how do we take
these advancements and actually turn
them into like real business outcomes or
how do we actually improve our
productivity or make our lives better.
And so now we're going to the next
update by SpaceX, which is their new
Grok bot platform. This is Grok bot and
Grok bot is a desktop app and an iOS
app. This right here is the desktop app.
So this app right here was being worked
on by the Cursor team. So Cursor was
working on this platform for many
months. I think like four or five months
and this was going to be their general
knowledge worker platform. Cursor's the
coding tool and this platform Grok bot,
which internally they were calling sand
and I believe the name that they were
going to use was actually dot. They
bought dot.com for like $7 million and
instead they went with Grok bot. And
this platform was going to be Cursor's
version of Claude co-work or GPT work,
but it has some key differences that I
think makes it pretty unique. And so one
of those things is that instead of
creating new sessions all the time like
you do in GPT work or Claude co-work
where on the left side panel, right? If
we were to go to Claude and if you were
using co-work inside Claude, you just
see all these different like chats that
get lost while you use it. What Grok bot
did is they just said each session, each
one of these sessions is like its own
agent. And so instead of like creating a
bunch of new sessions, if we just create
a new bot here, what it does, instead of
it being like a new session where you
just go and type in your request, it'll
immediately try and figure out what the
purpose of this session is and then it
will actually name the agent. And so
it's like, "Hi Riley, I'm here. What do
you want me around for?" Could be email,
content, code, a specific workflow, or
something else entirely. Weekly agent
updates, look at my YouTube and Notion
to get context. Your job is to help me
with these every week. And so, you can
honestly think of this as each one of
these is like your own little bot. And
you can set up plugins, just like any of
the other platforms. And you can also
set up skills, like I have my scrape
creator's skill right here. And all of
these agents share plugins and skills.
It's just that these new bots have their
own name, title, and little description.
And then when you create automations,
they get added right here. And so, each
session or agent has its own
automations, or they call them routines.
And you can see here, it just updated
help Riley with weekly agent updates by
pulling context from YouTube. And I'm
going to say, "Name yourself weekly
update." And then I can change the
title, and that will just change this
little tag right here. And so, I can put
title as like, "Help with updates." I
don't know. And it just shows up right
here. And so, this agent has a very
specific role for me. It just helps me
with these weekly updates. Helps me do
research, and that is going to be the
purpose. And so, whenever I want to work
with this, I would just come to the
weekly update agent. And I could say,
"Hey, every weekday,
um present me with the AI news for the
day at 10:00 a.m." And so, I can just
ask the weekly update bot, or yeah, Grok
bot, to create a routine. And so, it
created this routine, and you can see it
right here. If we open up this side
panel, you can see that we have this
weekly agent updates, and then we also
have weekly AI news. It's 10:00 a.m. on
a deliver Riley's daily AI news briefing
in chat. So, it'll give it to me every
day at 10:00 a.m. And so, this is
fundamentally different than Claude
Co-work, right? We could go to Co-work
and we could say every morning at 9:00
a.m.
do a task. You can see here it created
this morning, hello.
And notice here that like we have all
these different chats. Most of these
chats I'll never return to. It's hard to
return to them because like these just
kind of get lost. And so, it's hard to
like pick up pick back up on previous
work that we were working on. And so,
you'll notice here that Claude Co-work
when I create this chat, it adds
scheduled as like a global setting. So,
it doesn't really have much to do with
this chat session anymore. The scheduled
tasks, and if you have like a ton of
scheduled tasks, they live up here in
this scheduled section, whereas in Grok
Bot, they actually live inside of the
agent itself or inside the session. So,
I created this new session and now I
have a weekly update bot and the
routines live within here. For example,
my partnership bot has its own It has
its own routines. My content bot has its
own routines. Like it scrapes from all
my favorite creators every morning at
9:16 a.m. But, these bots have different
cron jobs or routines that live inside
the agent itself. And I If I press
command K, I can get a kind of a a
zoomed out view of all the different
routines that I have and here it will
actually list the name of the agent and
the name of the routine. So, I can see
all the weekly update agent routines,
the developer uh routine. We also have a
partnership bot routine and then I have
a to-do list bot, which prints my
reminder to do my most important task
every single day. I forgot to mention
that every single agent that you create
comes with its own little computer. So,
your it runs in the cloud. So, this is a
full computer in the cloud that you can
use and your agent, most importantly,
your agent can use this browser and you
can sign in to your stuff on this
browser and you can even
teach a task. And when you teach a task,
you can record yourself. This is very
similar to record and replay on Codex.
For those of you who watch my content,
you can do the same thing, but you use
your own computer. Here, I'm teaching
the agent to do a task in its computer,
right? This is a virtual computer
running in the cloud. You can actually
go in and see all of its files on the
computer. It's its own
thing that you have full visibility
into, which is a very new and
interesting thing, especially in a
platform like this. Real quick, before
the next update, I want to talk about an
update by the sponsor of this video,
GenSpark. One of my biggest inspirations
for getting into agents in the first
place was to keep track of everything I
do to get things done. I talk for a
living. Podcast, calls, meetings, random
ideas in the car. Normally, 90% of that
just evaporates. So, a few months back,
I started clipping this to the back of
my phone, GenSpark's second brain note.
I hit record when something's worth
keeping. There's a physical light, so
it's never a guessing game whether it's
on. It's SOC 2 and ISO 27001 certified
and it works in over 100 languages. So,
I use it everywhere, not just at my
desk. Here's the part that actually got
me. It's not just a recorder. Twice a
day, it views what I said and figures
out what to do with it. Told someone I'd
send them an email, it has the draft
ready. Agreed on a time on a call, it's
sitting on my calendar ready for
approval. Rift on a video idea out loud,
there's a script waiting inside Notion.
That's second brain. It remembers. The
part that does the work is GenSpark's
super agent. Second brain retrieves,
super agent executes.
All I do is say yes. They just opened
the first to the public GenSpark second
brain note, 10% off through the link in
the description. Speaking of agents that
actually get things done for you. And
so, the last thing that I'll say on this
is it's important to note the evolution
of these general agent platforms, and I
want to take a quick look into this.
You know, in the end of like 2025, so if
this was like end of 2025 and this was
kind of the first half of 2026,
we got these platforms that kind of
looked very similar. They were kind of
like the evolution of Claude code
running in your terminal, and they've
evolved into these like super apps. So,
like even the Hermes agent looks a lot
like Claude, co-work, um Codex has GPT
work, which looks similar. Um the open
Claude desktop app looks a lot like this
where it's this kind of agent platform
where you have a bunch of sessions, you
have skills, plugins, you have
artifacts, automations, um yeah, which
are like scheduled, and they're starting
to look relatively similar. It is very
interesting to note that the the two
previous general agent platforms that
have gone viral, which are Buzz and
Grokbot, have a new shape to them.
Right here Grokbot, you can kind of like
see your team of AI agents. I showed you
that you can create new agents really
quickly, right? You can just create
agents, you can give it a purpose, and
each one has its own routines. And then
the platform that went viral before this
was Buzz. And so, this platform Buzz
allows you to create
channels with a bunch of different
agents, and you can even add people to
your Buzz. And so, Buzz looks exactly
like Slack, and it's kind of like this
agent native Slack or Slack meant to be
used with humans and agents, which I
find to be very interesting. And
publicly, Anthropic hasn't really talked
about co-work that much. They talked
more about Claude tag, which is their
new platform that allows you to
basically create agents inside your
company Slack. To my knowledge, you can
only use it with Teams, but this is kind
of their focus right now is creating an
agent for groups of people or
enterprises or small businesses. I feel
like that's kind of the shift that we're
moving into. So, maybe the first half of
2026 or the first you know, first 2/3 of
the year was about the personal agent
and the rest of this year is kind of
about how do you create your own
personal team of agents? And then in in
regards to Buzz and Claude Tag, how do
you add an agent so that your entire
team can get access to the same agent so
you can collaborate using AI agents.
Okay, so now let's discuss updates
coming out of Anthropic. Yesterday,
Claude announced that your Claude Chrome
sessions now carry over to desktop web
and mobile. Conversations are saved and
your skills and connectors work inside
the browser. Basically, what Claude and
OpenAI are doing is they are inserting
GPT work in the case of OpenAI and
Claude co-work in the case of Anthropic
directly into your browser. I actually
have both of them set up. I can use
Claude here. And what they announced is
I can say, "Please tell me more about
this." And whatever I type in here
is basically the same as using Claude
co-work. It has access to the same
connectors, the same skills, and
everything. And all of the chats, right?
I can view the history. All of these
chats will actually sync into my Claude
app. I can very easily move this
conversation over to the Claude desktop
app.
And you can see here, it's named this
chat more information request. And I can
go back to Claude and I can see that
more information request is right here.
So, I can very easily switch from the
chat that I had in any browser, right?
It's just a Chrome extension, back to
Claude. So, it's equivalent to coming
here and using Claude, except you can do
it directly from your
uh Chrome extension in Chrome. The next
update to Claude is pretty interesting.
Sonnet 5 had an introductory price, and
it was scheduled to increase back up in
price at a certain date, but Claude has
decided that it would actually stay at
this cheaper price. Open AI is doing
something similar with their Terra
and uh Luna
models. So, basically, their
non-frontier models are getting much
cheaper, and I believe, and many
believe, that this is due to the
pressure from Chinese models out of Deep
Seek Kimmy, uh z.ai, and then now also
Grok, right? Because if their middle
models are way more expensive than these
other alternatives that are actually
better, then people have no reason to
use them. So, this is causing them to
lower their prices, and I expect this to
continue. The non-frontier models from
Anthropic and Open AI will continue to
get cheaper if the competition from
China and in the US continues such that
their frontier models are cheaper than
their middle models, everyone's just
going to use these models, and they'll
never use the middle models like Sonnet
and Opus and Terra and Luna. And for the
final update regarding Claude, uh this
guy released phone harness. So, this
isn't actually coming directly from
Anthropic, but if you look look up
GitHub phone-harness,
you will find a repo which will allow
you to fully control your phone from any
agent, not just Claude code, but Claude
code, CodeX, etc. You can fully control
your phone with an AI agent. I haven't
tested this out. I just thought this was
really cool. Thought I'd share it really
quickly. Okay, so we were about to talk
about the brand new Deep Seek V4 model,
which was just released this morning
officially, and Deep Seek said that
we're uh launching DeepSeek V4 today,
and they claimed that it was nearly as
good as Fable. And here's a very quick
summary after summarizing everything and
everyone's takes on Twitter. I'm trying
to figure out this story, but here's a
concise summary here. So, apparently
DeepSeek released a new pro model last
night and talked like it was almost as
good as the top expensive models like
Fable.
And it was only a tiny fraction of the
price. However, many people started
testing this, and they came back and
they said, "No, it's only a little bit
better than DeepSeek's previous model,
which was DeepSeek
V4 Flash." And so, this was
embarrassing, and they have basically
taken the model down. In fact, if you go
to Cursor right now and you switch to
DeepSeek V4 Pro, I'm using this via Open
Router.
It's not an official model inside Cursor
yet, and if I say, "Hi," it actually
will immediately fail. And so, this
model just doesn't work right now, and
so I can't fully cover it because there
is some backlash. Apparently, it's not
that much better than DeepSeek V4 Flash,
and they're also getting some flak for
their pricing. So, apparently the
pricing has come up, and they went to
usage-based pricing. But luckily for us,
58 minutes ago Gemini officially
released 3.7 Flash. And apparently, this
is a very fast model, and it's brand
new. And I do notice here, though, a red
flag is that the models they're
comparing against, right? You can see
Gemini Flash is being compared to Claude
Sonnet and GPT Terra. So, this is the
third best model by Claude and the
second best model by OpenAI. That's what
they're comparing it to. And so, they
haven't released a new pro model in a
while, and so this 3.7 Flash is a fast
mid-tier model. And if you want to test
it out, again, I like to do it inside
Cursor. I have it running because I'm
using open router. And if you were to
sign up for open router and just ask
cursor how to set it up, you get it set
up in like 2 minutes inside cursor. But
I can say, "Hey." Or you could use it
inside the anti-gravity. I'm sure they
have it inside anti-gravity. You can use
the new Gemini 3.7 model. Since it's
only a mid model,
and it's not that good, it's just like
pretty fast. I don't think I'll be
testing it too much. But if you want to
test it out, you can. Okay, so for the
final big thing that I want to talk
about today is I want to talk about what
I call the chat verse work problem. Many
people are confused about how to use
AI agents, specifically the Claude
co-work and the GPT work features inside
these platforms.
Signal said, "It's not clear to me
that OpenAI realizes how strange the
distinction between chat and work feels
in practice and how fragmented the
entire chat GPT experience has become.
If you leave it on work, every simple
question or search becomes an
expedition. It starts thinking,
planning, and using tools when you just
wanted a quick answer." What he's
talking about is in chat, you now have
chat and work. And these are two very
different products. Obviously, you know
what chat GPT is. But GPT work is a
bigger thing, right? Here, you can
actually do things on Slack. You can
have it control your Gmail. You can have
it fully control your notion. You can
literally get, if you set up the right
scheduled tasks, you can have it fully
reply and send out emails, and you can
get it to run like it is a full agent
platform. And what he's talking about
here is just hard to understand the
distinction between chat and work. I
made a full video on GPT work. It's
incredibly powerful, but the distinction
between chat, work, and then the Codex
app is pretty confusing right now.
He goes on to say that there's no good
default. Chat is often too limited for
real tasks. Work is too slow and
cumbersome for normal queries.
Constantly switching between them makes
the entire product infinitely more
complex. Also, the lack of sync between
mobile and desktop plus what's a local
chat versus a cloud chat is a mess. And
again, these are things that I talk
about in my last video. It is relatively
confusing the difference between a local
chat and a cloud chat. Chat GPT went
from the most usable, simple consumer
experience to confusing AF in such a
short period of time. Pretty nuts. I do
think it's incredibly powerful, but I do
think he's talking about a very
difficult problem, which I call the
chat versus work problem. And I tweeted
this about the brand new Grokbot
yesterday.
I showed you the Grokbot platform
earlier. And then I tweeted this
specifically around how the platform is
set up. I said that I think Grokbot is
behind Codex and GPT work for a lot of
reasons. There's a lot of things that
GPT work can do that Grokbot cannot do.
But I will say their biggest innovation
is personifying the chat sessions. An
agent or a Grokbot is basically a named
session or a named chat session with a
mini system prompt. For example, on
Grokbot, if you just click on the name,
right, this description is just a mini
system prompt. And all of the plugins
that you can create, right, I can add
Gmail, that is a plugin. I can also use
skills. And all of these skills are
global to all of my agents. They
basically just made each chat session
have its own system prompt and your its
own name. So that when I want to create
a content script, I'll just go to my
content agent. Or if I want to scrape
from social media, I'll go to my content
agent. If I want to do my weekly update
agent, I'll go to my weekly update
agent. If I need something handled in my
partnership bot, um I will go to my
partnership bot. Not only that, but like
if something happens inside Slack in
this Slack channel, it will
automatically ping me. And it stays
organized by these chat sessions. And
so, I think the biggest innovation of
Grokbot was kind of making this analogy
easier to understand. And so, I think
we're going to see a lot of innovation
over the next 3 months as the frontier
labs try and figure out what is the best
setup for people to use AI agents in
their business. And I think there's just
going to be a lot of innovation. Because
this is something I would have never
thought I wanted, but once I used it, I
was like, "Okay, this actually makes
more sense." The automations should not
live at a global level, they should just
live within the bot. I have a feeling
that the other labs are going to copy
this design because I do find it's a lot
easier to get started, and it just
intuitively makes sense that you have a
bot. Each bot has its own computer, and
it has its own routines. I find it to be
an easy interface to pick up and
understand. Another thing is that
workspace agents are coming soon to chat
GPT. And I know this because I quote
tweeted their workspace agents. And so,
workspace agents are only on the GPT
team plans. It allows you to create
these like agents, right? You can create
agents. And these are as powerful as GPT
work, except they do have their own
system prompts. You can message these
custom agents, and you can even add them
to Slack and message them through there.
And it's only available to teams. And
so, I tweeted this. I said, "Wish this
wasn't only a Teams plan and only on
web. This should be on desktop as well."
And the Someone from the OpenAI team,
Andrew, said, "Yes." So, this indicates
that this is coming very soon to the
chat GPT desktop app, and it won't just
be a Teams plan, which I think is really
cool. This is This is the most slept on
OpenAI product yet, in my opinion.
And then finally, one update is that
codex you can get and download on Linux.
So codex the codex app or the chat GPT
app where I use codex
you can get this on Linux now. So that
is a new update on the codex side. And
yeah, that's basically everything. So
this was an agent native update. I'm
Riley Brown. Thank you guys so much for
watching and please like, please
subscribe. It helps you out a ton. I'll
see you here for the next video.
Ask follow-up questions or revisit key timestamps.
This video covers a major week in AI, focusing on the release of SpaceX's Grok 4.6 model and the introduction of their new 'Grok Bot' platform. The host discusses how Grok 4.6 is challenging current frontier models in terms of performance and cost-efficiency. A significant portion of the video is dedicated to the 'chat versus work' problem, where the host highlights how Grok Bot improves upon the user interface of existing tools like Claude Co-work and GPT Work by personifying chat sessions into specialized agents. The video also touches on updates from Anthropic, Gemini, and DeepSeek, alongside future developments expected in AI agent platforms.
Videos recently processed by our community