Codex Is Winning… So Why Am I Still Paying for Claude?
607 segments
The last two weeks in the world of AI
agents have been very interesting. Mark
Zuckerberg and Meta released their new
coding agent, Musecode, to rival Claude
Code and Codeex. And you can use Muse
Code in the terminal to build apps or as
a general agent. And it only takes 2
minutes to set up. I'll show you how.
But we have way more to talk about like
Deep Seek, Kimmy, and Quen and more
models coming out of China. We also need
to talk about Codec's latest updates. We
also need to talk about Cursor and how
they're shifting from a developer tool
to a super app just like Claude and
Codeex. We will also discuss Google and
how they might actually be out of the AI
race. And we need to talk about Buzz for
creating teams of agents and all of the
new updates there. This is an agent
native update where we cover the most
important news and updates on Frontier
AI platforms and models so you can use
AI agents to be more productive. I'm
Riley Brown. Let's go.
The first agent native update comes from
Meta, which is very rare. They don't
usually do that many things on the
frontier, but here we go. Mark
Zuckerberg on Twitter said, "Releasing
Muse Code in beta today. It's a terminal
coding agent that takes on complete
software engineering tasks across large
repos, planning changes, writing code,
validating the results powered by Muse
Spark 1.2, the model, a coding focused
model update. And based on the
benchmarks that he posted, it's right in
between Opus and GPT 5.6 Terra, which is
their middle model between Soul and
Luna. You can see Muse Spark 1.2. So,
here are the prices of the meta model
compared to 5.6 Soul and Claude Opus.
Uh, and if you add up the input cost and
the output cost, you get $35 input plus
output for GBT 5.6. And for Opus, you
get $30. And for Meta Muse Spark, with
the top model on there, it's only $5.50.
So that's a total difference of like
five to 6x. So it's significantly
cheaper than using the top models at
OpenAI. Now, when you start comparing it
to OpenAI's GPT 5.6 Terra, it gets
pretty comparable, which we'll talk
about later because OpenAI is lowering
the prices of their other model. But
always remember, every single update
that I talk about is best understood by
actually testing it out, not just
looking at some charts. So now, how do
we actually use Muse code? It's very
simple. What you should do is you should
go to either Claude Code or Codeex. If
you watch my videos, you know that I use
codec a lot more. And I'm just going to
go to Codeex and say, "Hello there,
Codex. Can you please download the
latest Muse code by Meta? Mark
Zuckerberg announced it yesterday. The
terminal coding agent. Please download
it now so I can use it in my terminal.
So all we have to do is run this and we
will wait a few seconds and it's going
to download it to our computer. Okay. So
now it's done and remember this is a
terminal coding agent. So you can run
this in the terminal. I can either go to
the terminal app on my computer or if
I'm using the codeex app or chatgbt on
the codeex version. I can press commandJ
and this is the terminal. And now I can
type in muse. And now I need to select
trust and continue or quit. I'm going to
click trust and continue. And here we
can either log in with browser or set an
API key. And so we're logging into Meta
Platform. And so if I click login with
browser, it'll automatically take me to
the browser. And here I can actually
just sign in with meta. Once you sign
in, it'll look a lot like this. And now
we can just use Musecode. So I can say
hello, what model are you? And here it
says I am Musecode powered by Metam Muse
Spark. And in order to use this going
forward, I normally just do it inside
the terminal, the normal terminal. And
we can zoom in. And remember, all you do
is type muse. And now you're using it
inside the terminal. And you can full
screen it if you want. I'm going to say,
"Hey, I want you to create a bowling
simulator game." Like it should be like
we bowling. I want you to make a wee
bowling and then I want you to run it
locally so I can play it.
Now we're creating Wii bowling. And I
highly recommend just testing this out
for yourself. Do it for knowledge work
tasks. Get it to create documents,
spreadsheets, that type of thing. And
this model is really fast. And in my
opinion, it's somewhere between obus and
sonnet. By default, it will tell you
something like this sandbox
blocks persistent servers. So, we
actually need this to be in yolo mode to
get it to automatically be able to run
terminal commands so that it can run
whatever app that you create. And we can
do this by going to codeex. And we can
open codeex and we can say something
like this. I need you to uh set it so
that um muse code always defaults to
yolo mode. Please just do that right now
for every session.
And this will allow your Muse code to do
anything you want or do anything on your
computer. So you won't have to like give
weird permissions. And then after you do
that, uh, if you were to open up a new
terminal session by pressing command N
and you were to type in Musecode
here, you can see you are on YOLO mode.
So now Musecode has full control over
everything and you can get it to do
everything. And so now I want to check
on that bowling game. And okay, so this
is the game that it created. We have
bowling and we can up the spin. So, this
will spin it left, which means I need to
aim it like here.
Oh,
okay. Maybe I won't do spin.
There you go. So, we're using Muse Code.
This isn't the greatest game ever
created, but that's probably a problem
of prompting. I just wanted to empower
you with the ability to test it. It's
really easy to test out. I highly
recommend testing out the new meta
terminal coding agent. Okay, for the
second agent native update, I want to
talk about the changes to codeex over
the past 2 weeks. We're going to come to
the mobile app in just a second. Let's
start off with the desktop app. In the
desktop app, we have a few new changes.
In the side panel over here, you'll see
that we have the normal view right here
where I can see all of my chats, but
there's also this notifications bar. And
so this is kind of the activity pane. So
instead of them showing in this like
fixed project view right here, it'll
show you the most recent ones. And it's
really intuitive to use and it'll show
you the exact folder that they're
running in as well. So you can see where
the directory is on your computer. So,
this is just kind of feels a lot more
like notifications
and it shows the most recent one that
completed, which I personally use most
of the time. Okay, for the second update
to the desktop app, which is my favorite
design update that they've made in a
very long time, involves the inapp
browser. And so, I in this chat created
a website. And so, this is a website.
It's kind of ugly, but that's okay. And
if we were to full screen it so that the
agent chat disappears. Now they updated
this bottom part right here. So I can
fully I can click on this and it opens
up and I can very easily say like
the line that goes beneath episodes. I
don't really like that very much.
Uh can you also please make the like
dark purple behind the logo at the top
better? The whole top bar is pretty
ugly. Can you please fix that?
And so I can just run it right here. And
it's really easy to see what the agent
is doing. And then I can pin it to the
bottom. So you can see here if I scroll
down to the bottom, I see this little
chat GBT logo. And now I can open it up
and I can get rid of it. So it allows me
to fully immerse myself in the website.
Right? This is basically full screen.
And if I ever want to make changes, I
just come down to the bottom, press this
open AI thing right here. Now, what I do
want to let you know is if you use
Whisper Flow, it will get in the way of
this. So, as you can see here, I moved
Whisper Flow to the right side of my
screen. And if you have Whisper Flow,
see how it gets in the way. So, what you
need to do is you need to come over
Whisper Flow and just drag it to the
left or to the right. And I choose to
put it on the right. So, now I can use
Whisper Flow. Whisper flow is right
here. And the OpenAI edit is here. And
you can notice here that browser use is
coming into play and it's actually
controlling the browser. So that's just
a fun new design of the inapp browser
inside codeex. Well, staying on the
topic of browsers, they also made
updates to their Chrome extension and
you can find this in plugins. And if you
just type in Chrome and you come here
and you download this when you go to
your Chrome browser, you will see this
chat GPT icon and you can move it to the
front slot if you really like it. And
now what we can do is we can select this
and we can just say please prepare a
tweet based on the recent um memories
that you have of me based on what we did
today. come up with the best possible
tweet and I can fire this off,
control this browser and make the tweet,
but don't post it, just cue it up.
And so all of the chats that you do
inside this Chrome extension sync with
the codeex app. So you can see this is
called draft tweet from today. I can
very easily just go to chat GPT. If we
go to chatgpt, you can see that it says
draft tweet from today. And so we are
connected to Chrome. And so here it
actually queued it up. Okay. So the
biggest AI shift isn't better answers,
it's agency. Okay. Well, we need to work
on the tweet, but it queued it up. And
if you want to open it straight inside
codeex from here, all you need to do is
go here. You're going to hit these three
dots and click open in app. And you're
going to open chat GBT. And that will
automatically open it directly inside
the app. So you can basically just use
the codeex app inside Chrome. And this
is just a nice new update. Okay, for the
last update involving chat GBT, we have
a new toggle here at the top. So now you
can toggle between chat and work. I just
released or I might be just about to
release an hourong video on chat GBT
work. It is the most in-depth guide
ever. I was wrong about Codeex. I said
that I wish they didn't split up Codeex
and GPT work. I was wrong. GBT work is
insanely useful. It's completely
cloud-based and it is I I use it every
day. So, I highly recommend starting to
use this from your phone. You can
basically control your email calendar,
all of the plugins that you would find
inside the chat GPT app right up here.
You can basically fully control with GPT
work. So, all of the apps you add, you
can fully control from your phone, which
is just incredibly useful. And then
after this update, it's just more
accessible here at the top between chat
and work. I hardly ever use chat
anymore. I just use work. For the next
agent native update, I want to talk
about this tweet chatbt desktop app or
formerly known as codeex. Their biggest
competitor right now is cursor, not claw
desktop. And the reason I tweeted this
is I have it on good authority. I've
talked to some people and cursor is in a
middle of a revamp. They are about to
make major changes to their platform and
it is going to become a full super app
just like codeex. And the first signs
that we're seeing of this is you can now
connect cursor to Google Workspace. This
is their CTO at Cursor. Everyone thinks
of Cursor as a tool for coding. We
thought so too. But inside the company,
many of our use cases aren't coding at
all. Research, data analysts, bug
triage, and project management, just to
name a few. As it turns out, coding
agents are pretty good foundation for
all kinds of work. Could be a sign of
what's to come. So very clearly they're
going to make changes to their platform.
Inside cursor uh you will see this
customize tab. This is equivalent to the
plugins tab inside codeex. And if you
come to the top and you just type in
Google drive, you will see this Google
drive setup here. I've already set it
up. So you will need to authenticate it.
But once you authenticate it, which just
means sign into your Google Drive files,
you can then try it in chat. Here I have
cursor set up to the new DeepSseek V4
model. Um, and so now this model's
basically free to use. I'm not even
kidding. And so I'm going to go ahead
and stop this and I'm going to say,
"Hey, can you please create a quick
spreadsheet on um just the steps on how
to um train an AI model along with the
description on how to do it? Um, make it
look professional and then send me the
link to the Google Sheets for this."
And so now I'm using DeepSeek V4 Flash,
the cheapest model in the world by
Deepseek. It's like basically free to do
this. Maybe a scent or two. And it will
create a Google Sheets link. And now
it's working inside cursor. And there
you go. It says done. Your Google sheet
is created and ready. It worked for 44
seconds. So it's really fast. I can open
this up. And boom. Here we go. The point
is is it can control uh Google Docs,
Google Sheets, Google Slides, anything.
And you can use a better model than Deep
Seek V4 if you want, but that's just the
one that I've been using for fun. And
for the fourth Agent Native update, we
have three new models coming out of
China. We have Kimmy K3, we have
Deepseek V4 Flash, and we also have Quen
3.8 Max. And these models are all good
in their own way. Kimmy K3 is the best
out of all three of the new models, but
it's also the most expensive. I think
many people were really excited by how
good the model was, but they were
underwhelmed by the price. It is quite
expensive. It's like almost as expensive
as Sonnet for certain tasks, but I
highly recommend trying it out. Then
Deepseek V4 Flash is probably the most
interesting because for certain tasks,
it is 105 times more cheap than Fable 5.
And don't worry, I'll show you how to
get this model set up inside Cursor in
just a second, as well as all the other
models. But I do want you to keep in
mind that regarding this Deepseek V4
flash model, which is really good for
how expensive it is or how cheap it is.
They said this, Deepseek said this. They
said, "We plan to raise the overall
pricing for Deepseek API services in the
near future with a significant increase
expected. Please plan your usage
accordingly." So that's not good. Many
people are like, "Oh no, um, how
expensive is it going to be?" However,
DAX from Open Code said, "On the
upcoming DeepSk price increase, we've
been able to reproduce their current
prices even on rented GPUs." So, this
likely isn't because they're losing
money. It's traffic shaping because
they're overloaded. So, he's saying that
we're going to be able to host this
DeepSeek model in the US for a similar
price that it is now. They're saying
that the price increases just because so
many people around the globe are trying
to use it. Klein posted this about
DeepSeek V4 Flash. While Deepseek V4
Flash is significantly cheaper on price
per token, this could be misleading if
the overall cost per task ends up being
higher due to more turns being made.
However, Artificial Analysis reports
DeepSeek completing the same benchmark
tasks as Fable at 105 times lower cost,
which is absolutely insane. So, my take
on this is as follows. I believe that
DeepS V4 Flash may get more expensive in
the future. Maybe it'll get three times
as expensive. Maybe it'll get five times
as expensive in the short run because
there's so much traffic and so much
demand for this model. But I think over
the long run, like over the next two
months, this model will get
significantly cheaper or a model just as
good as it will get as cheap as the
model is now. And so for the next week
or two, I highly recommend trying these
models. The truth of the matter is every
few weeks new open models are released
and the costs are simply going down. And
so you can do so much with a model uh
just as simple as V4 Flash. A lot of
tasks you can do and this model is 20 uh
to 100 times cheaper depending on the
task than Fable. So I highly recommend
testing out these models and seeing what
you can do with them. For the next agent
native update, I want to talk about
Anthropic. And there's a vibe going
around Twitter and even on YouTube that
people are getting a little fed up with
Anthropic. Not only is Fable only on API
usage now, so it's incredibly expensive
to use Fable, many people are showing
disappointment surrounding their new
models, Opus 5 and Sonnet 5, and that
they completely wasted time building
them. And I'm not kidding. A lot of
people, you know, AI researchers are
literally saying that Opus 5 sucks. John
Enis said that I have decided Opus 5
extra high is basically trash. It
doesn't use its thinking budget to do
anything more productive. It just uses
it to thrash around to do pointless and
sometimes harmful things. Fable is a
good model and I will miss it. But I
think I'm finally ready to just let
Claude go or at least until they release
a new model. And I think finally more
people are realizing that the vibe is
shifting from anthropic to codeex again
at least for now. Humal Hussein said uh
it's crazy how the consensus shifted
away from claude being the favorite to
codeex. It's not just vibes things
favoring codecs. It's a better harness.
Uh the codeex desktop is significantly
better. Better pricing especially with
the latest pricing updates to the new
codeex models. less refusals and you can
use your subscription freely wherever
you want which is really cool. Open uh
Anthropic has a lot more guards against
using your subscription on other tools.
However, I just tweeted this. However, I
just tweeted this prediction. Anthropic
is pulling back the slingshot and will
try to go on another run soon. I think
they got high on their momentum from Q1
and Q2 and tried to launch too many
things and their products got confusing
and their momentum wore off and they
just launched so many different products
that people couldn't even keep track of
everything. Like when you tried to use
Claude Co-work and connect it to your
phone, it was called Dispatch, but if
you did the same thing with Claude Code,
it was called Claude Remote. Um, they
released a law product, they released
Claude Design, which is a good product.
just got buried under all the different
product announcements. They had an
extension and then they had the whole
mythos fiasco, Fable, Sonnet, Opus. And
honestly, their model releases besides
Fable just haven't been that good since
like Claude 4.6. All of them have felt
relatively similar. However, there's one
thing that just keeps me using
Anthropic's products, and that is the
fact that it is just so much better at
front end. I gave Codeex and Enthropic
the same prompt. And this is what Codex
created and it's just so much worse than
Anthropic. It is just so much better at
front end. And this is not just like
frontend for landing pages and other
things like that. It's also for
spreadsheets, docs, and presentations.
Opus and Fable blows the OpenAI models
out of the water for these like
knowledgework type documents and
front-end design. For everything else, I
think GPT 5.6 Soul is basically as good
as Fable. It's just so much better at
front-end design and it created this in
one prompt and I just think it looks so
good. Like this is I'm working on my
website for agent native and it's just
so much better. So, this is the one
thing that keeps me using the claw
desktop app. It's just better at design.
And don't forget that the claw desktop
app, if you go to home, there is design
mode. So, you can use claw design
directly inside the desktop app. It's
just they've launched so many things
that so many people forget about claw
design and it's still really good. And
the final thing that I want to discuss
is I believe that we are entering a new
era in the world of agents. I think the
first half of the year was kind of the
openclaw personal agent movement and I
believe that we're moving into the team
of AI agent movement. Now I've already
made a video talking about Buzz and so
this is Jack Dorsey's new platform where
it is a Slack
clone basically except it's made to be
used with AI agents. These are my
existing codecs and claude code running
in this slack-like interface. And I can
at@mention claude and I can at@mention
codeex and say hey work together to
build an app that lets me use deepseek
and you can at mention them and you'll
notice here that they both reacted to
it. You can see Codeex and Claude Code
reacted to it. And now you can see that
they're both working on a response.
Claude Code and Codeex are going to
respond. And I can message them as if
they were just humans and they will
actually work together to get it done.
And if you come to the agents tab, you
can see your full team of AI agents. And
you can add cursor, you can add Devon,
and you can make any model the default
model. So I could come here and create
an agent and I could customize the agent
and I could use cursor for example and I
could choose any model from cursor and
for this I'm actually going to go ahead
and make the default model Kimmy K3. So
now we're using Kimmy K3 with cursor and
I can give it instructions. you are a
content agent and I can very easily name
this agent cursor with Kimmy and I can
create an agent and now this agent is
running and so I could very easily go
into my content channel and I could add
or I could actually just go like this.
Hey codeex,
uh please add cursor with Kimmy to this
channel and all other channels and
codeex or any other agent can fully
control this version of Slack. And so it
can add any agent to any channel and it
can even create agents. And here at the
bottom you can see who's working. So I
can see that codeex is working and you
can see that cursor with Kimmy is
requesting approval and this just kind
of a team of agents that can do things
together. And you can see here curs uh
CW which is cursor with Kimmy just
responded and it said I handled it
myself. I joined all seven channels
content visual coding general research
management and notes. And the reason I'm
showing you this is I believe we're in
the very early innings of working with
the team of AI agents. It's not super
easy yet. The same way that OpenClaw
wasn't very easy early on. And now GBT
work is basically OpenClaw running
inside chat GBT and it can basically do
anything for you. And so I think we've
advanced really far on the personal
agent side which is advancing these AI
agents that we can access through our
phones. I think we're about to see the
AI agent for teams really take shape
over the next four to six months. So
definitely be on the lookout and as this
progresses, I highly recommend if you
work at a large company to be the person
who can build a team of agents for your
team or create an agent that everyone on
your team can access to get things done.
And I think we're going to see so many
platforms like this. we're going to see
a lot more Slack agents that make agents
really easy to configure. And so that's
something I would keep a close eye on.
Anyway, that's an update. Those are the
things that have really interested me
over the past two weeks. I really hope
you like this video. And if you could
like and subscribe, it would help me out
a ton. I really appreciate you guys. I
will see you here for the next
Ask follow-up questions or revisit key timestamps.
This video provides an overview of the latest developments in AI agents over the past two weeks. Key highlights include the release of Meta's Musecode terminal coding agent, significant UI and functionality updates to the Codeex/ChatGPT platform (including the new 'Work' mode), the evolution of Cursor into a broader productivity tool, the emergence of new Chinese AI models (Kimmy K3, DeepSeek V4 Flash, Qwen 3.8 Max), and a critical look at Anthropic's current standing. Additionally, the video explores the transition from individual personal agents to collaborative 'teams of agents' using platforms like Buzz.
Videos recently processed by our community