HomeVideos

Grok 4.6 is Actually Good… And Claude Keeps Getting Better

Now Playing

Grok 4.6 is Actually Good… And Claude Keeps Getting Better

Transcript

653 segments

0:00

This was the biggest week of the year

0:02

for Elon Musk and his AI efforts at

0:04

SpaceX.

0:05

>> [music]

0:05

>> They released Grok 4.6, which shows that

0:08

they are finally catching up to OpenAI

0:10

and Anthropic. They also released Grok

0:11

Bot, which is their new super app, which

0:13

will rival GPT work and Claude co-work.

0:16

And I have a lot of thoughts about this

0:18

platform and what makes it so

0:19

interesting, which I'll talk about

0:21

today. But we have way more to cover. We

0:23

will also discuss the latest updates

0:25

inside Claude Code and Codex, as well as

0:28

the latest DeepSeek V4 model that they

0:30

launched today. And we even have some

0:32

news from Gemini. This is an

0:35

agent-native update where we put the

0:36

latest advancements on the frontier of

0:39

AI agent platforms and models into

0:41

context so that we can actually use them

0:43

to improve our business.

0:47

Okay, so we have a lot to cover today.

0:49

Let's dive straight into the SpaceX

0:51

update. So SpaceX said this yesterday,

0:53

"Introducing Grok 4.6. It delivers

0:56

frontier intelligence and is a

0:57

significant improvement over Grok 4.5 at

1:00

the same price." And so you'll notice

1:02

here that the three areas that they led

1:05

were economically valuable work, long

1:08

professional tasks, and legal work. And

1:11

so if you also notice that it isn't

1:13

coding, right? They didn't lead in any

1:15

of the coding benchmarks. And I believe

1:17

that this is because they are focused on

1:19

general agent tasks because this is

1:21

simply their priority, which explains

1:24

their brand new platform that they

1:26

worked on with Cursor, which is called

1:28

Grok Bot. Grok Bot is their brand new

1:31

super app, which we'll talk about in

1:32

just a second, that is focused on

1:35

non-coding work. They are trying to get

1:37

everyone within a company to interact

1:39

with AI agents to help them get work

1:42

done. They're focused on knowledge work.

1:44

And I still believe that the best and

1:46

easiest way to use Grok 4.6, this brand

1:49

new model, is directly inside Cursor. As

1:52

soon as you update Cursor, it will

1:53

default to Grok 4.6 fast, and you can

1:57

use it directly inside Cursor. So, I've

1:59

not yet had enough time to actually do a

2:01

deep test of Grok 4.6, but I do think

2:03

there's some interesting use cases to

2:05

discuss that people have posted on

2:06

Twitter. So, here's DHH. So, Fable, I

2:09

think this was last week, one-shotted a

2:12

Rust rewrite of the terminal text

2:14

effects Python library in 11 million

2:17

tokens.

2:18

If you don't know what that means,

2:19

that's perfectly fine. Fable did a

2:21

really, really hard thing. And then,

2:25

today, he tweeted that he used SpaceX's

2:28

new model, Grok 4.6, with just a couple

2:31

of nudges, it was able to repeat this

2:34

feat in in about an hour and a half, and

2:36

the key takeaway is that it was only

2:38

$55.

2:40

That was about 1/10 of the cost of Fable

2:43

implementation for the same work. So,

2:45

take a look at the pricing for these

2:48

models. If we look at Grok 4.6 compared

2:51

to Opus, Soul, and Fable,

2:54

right? If we combine the input and

2:56

output prices per 1 million token, we

2:59

get $8 for Grok 4.6, $30 for Opus 5, $35

3:06

for 5.6 Soul, and $60 for Claude Fable.

3:09

And so, that means that Claude Fable 5

3:11

is 7.5 times more expensive than Grok

3:14

4.6, 5.6 Soul, 4.4 times, and Opus 5,

3:19

and Grok 4.6 is better than Opus. It is

3:22

straight-up better, and Opus is still

3:25

3.75 times more expensive than Grok 4.6.

3:29

This is a very good model, and it is a

3:31

reasonable price, and it is legitimately

3:33

on the frontier. And beneath this

3:35

Cognition post, Elon commented, "Grok

3:37

4.7 will exceed all current models,

3:41

which includes Fable." He said, "That

3:43

said, Anthropic is a great company and

3:45

will probably release improved models

3:47

soon. However, the SpaceX training

3:49

corpus is so awesome and unique that I

3:52

would be shocked if any model is better

3:54

at real-world engineering than 4.7. And

3:57

so that was their model release. And if

4:01

you watch my channel, you know that I

4:02

actually don't dive too deep into model

4:04

releases. I care mostly about practical

4:06

use cases of AI. Like how do we take

4:09

these advancements and actually turn

4:10

them into like real business outcomes or

4:13

how do we actually improve our

4:14

productivity or make our lives better.

4:16

And so now we're going to the next

4:18

update by SpaceX, which is their new

4:20

Grok bot platform. This is Grok bot and

4:22

Grok bot is a desktop app and an iOS

4:25

app. This right here is the desktop app.

4:27

So this app right here was being worked

4:29

on by the Cursor team. So Cursor was

4:32

working on this platform for many

4:33

months. I think like four or five months

4:35

and this was going to be their general

4:37

knowledge worker platform. Cursor's the

4:40

coding tool and this platform Grok bot,

4:43

which internally they were calling sand

4:45

and I believe the name that they were

4:47

going to use was actually dot. They

4:48

bought dot.com for like $7 million and

4:52

instead they went with Grok bot. And

4:54

this platform was going to be Cursor's

4:56

version of Claude co-work or GPT work,

4:59

but it has some key differences that I

5:01

think makes it pretty unique. And so one

5:03

of those things is that instead of

5:06

creating new sessions all the time like

5:09

you do in GPT work or Claude co-work

5:11

where on the left side panel, right? If

5:13

we were to go to Claude and if you were

5:15

using co-work inside Claude, you just

5:17

see all these different like chats that

5:19

get lost while you use it. What Grok bot

5:23

did is they just said each session, each

5:26

one of these sessions is like its own

5:29

agent. And so instead of like creating a

5:31

bunch of new sessions, if we just create

5:33

a new bot here, what it does, instead of

5:35

it being like a new session where you

5:38

just go and type in your request, it'll

5:40

immediately try and figure out what the

5:41

purpose of this session is and then it

5:43

will actually name the agent. And so

5:46

it's like, "Hi Riley, I'm here. What do

5:48

you want me around for?" Could be email,

5:50

content, code, a specific workflow, or

5:52

something else entirely. Weekly agent

5:55

updates, look at my YouTube and Notion

5:59

to get context. Your job is to help me

6:04

with these every week. And so, you can

6:07

honestly think of this as each one of

6:09

these is like your own little bot. And

6:11

you can set up plugins, just like any of

6:14

the other platforms. And you can also

6:16

set up skills, like I have my scrape

6:17

creator's skill right here. And all of

6:21

these agents share plugins and skills.

6:23

It's just that these new bots have their

6:26

own name, title, and little description.

6:29

And then when you create automations,

6:31

they get added right here. And so, each

6:34

session or agent has its own

6:36

automations, or they call them routines.

6:39

And you can see here, it just updated

6:41

help Riley with weekly agent updates by

6:43

pulling context from YouTube. And I'm

6:45

going to say, "Name yourself weekly

6:49

update." And then I can change the

6:50

title, and that will just change this

6:52

little tag right here. And so, I can put

6:54

title as like, "Help with updates." I

6:57

don't know. And it just shows up right

6:59

here. And so, this agent has a very

7:01

specific role for me. It just helps me

7:04

with these weekly updates. Helps me do

7:06

research, and that is going to be the

7:08

purpose. And so, whenever I want to work

7:10

with this, I would just come to the

7:11

weekly update agent. And I could say,

7:13

"Hey, every weekday,

7:16

um present me with the AI news for the

7:21

day at 10:00 a.m." And so, I can just

7:24

ask the weekly update bot, or yeah, Grok

7:28

bot, to create a routine. And so, it

7:30

created this routine, and you can see it

7:32

right here. If we open up this side

7:34

panel, you can see that we have this

7:36

weekly agent updates, and then we also

7:39

have weekly AI news. It's 10:00 a.m. on

7:41

a deliver Riley's daily AI news briefing

7:44

in chat. So, it'll give it to me every

7:46

day at 10:00 a.m. And so, this is

7:48

fundamentally different than Claude

7:50

Co-work, right? We could go to Co-work

7:51

and we could say every morning at 9:00

7:53

a.m.

7:54

do a task. You can see here it created

7:56

this morning, hello.

7:57

And notice here that like we have all

7:59

these different chats. Most of these

8:00

chats I'll never return to. It's hard to

8:02

return to them because like these just

8:05

kind of get lost. And so, it's hard to

8:07

like pick up pick back up on previous

8:10

work that we were working on. And so,

8:12

you'll notice here that Claude Co-work

8:15

when I create this chat, it adds

8:17

scheduled as like a global setting. So,

8:20

it doesn't really have much to do with

8:22

this chat session anymore. The scheduled

8:25

tasks, and if you have like a ton of

8:27

scheduled tasks, they live up here in

8:29

this scheduled section, whereas in Grok

8:31

Bot, they actually live inside of the

8:36

agent itself or inside the session. So,

8:39

I created this new session and now I

8:40

have a weekly update bot and the

8:43

routines live within here. For example,

8:45

my partnership bot has its own It has

8:49

its own routines. My content bot has its

8:52

own routines. Like it scrapes from all

8:54

my favorite creators every morning at

8:56

9:16 a.m. But, these bots have different

9:00

cron jobs or routines that live inside

9:03

the agent itself. And I If I press

9:05

command K, I can get a kind of a a

9:07

zoomed out view of all the different

9:09

routines that I have and here it will

9:11

actually list the name of the agent and

9:15

the name of the routine. So, I can see

9:17

all the weekly update agent routines,

9:20

the developer uh routine. We also have a

9:23

partnership bot routine and then I have

9:25

a to-do list bot, which prints my

9:29

reminder to do my most important task

9:31

every single day. I forgot to mention

9:33

that every single agent that you create

9:35

comes with its own little computer. So,

9:38

your it runs in the cloud. So, this is a

9:40

full computer in the cloud that you can

9:42

use and your agent, most importantly,

9:45

your agent can use this browser and you

9:48

can sign in to your stuff on this

9:50

browser and you can even

9:52

teach a task. And when you teach a task,

9:55

you can record yourself. This is very

9:57

similar to record and replay on Codex.

10:01

For those of you who watch my content,

10:03

you can do the same thing, but you use

10:04

your own computer. Here, I'm teaching

10:07

the agent to do a task in its computer,

10:10

right? This is a virtual computer

10:11

running in the cloud. You can actually

10:12

go in and see all of its files on the

10:15

computer. It's its own

10:17

thing that you have full visibility

10:18

into, which is a very new and

10:21

interesting thing, especially in a

10:23

platform like this. Real quick, before

10:25

the next update, I want to talk about an

10:27

update by the sponsor of this video,

10:29

GenSpark. One of my biggest inspirations

10:30

for getting into agents in the first

10:31

place was to keep track of everything I

10:33

do to get things done. I talk for a

10:35

living. Podcast, calls, meetings, random

10:37

ideas in the car. Normally, 90% of that

10:40

just evaporates. So, a few months back,

10:41

I started clipping this to the back of

10:44

my phone, GenSpark's second brain note.

10:47

I hit record when something's worth

10:48

keeping. There's a physical light, so

10:50

it's never a guessing game whether it's

10:51

on. It's SOC 2 and ISO 27001 certified

10:55

and it works in over 100 languages. So,

10:58

I use it everywhere, not just at my

11:00

desk. Here's the part that actually got

11:02

me. It's not just a recorder. Twice a

11:04

day, it views what I said and figures

11:06

out what to do with it. Told someone I'd

11:08

send them an email, it has the draft

11:10

ready. Agreed on a time on a call, it's

11:13

sitting on my calendar ready for

11:14

approval. Rift on a video idea out loud,

11:17

there's a script waiting inside Notion.

11:20

That's second brain. It remembers. The

11:22

part that does the work is GenSpark's

11:24

super agent. Second brain retrieves,

11:27

super agent executes.

11:29

All I do is say yes. They just opened

11:31

the first to the public GenSpark second

11:34

brain note, 10% off through the link in

11:36

the description. Speaking of agents that

11:39

actually get things done for you. And

11:40

so, the last thing that I'll say on this

11:42

is it's important to note the evolution

11:44

of these general agent platforms, and I

11:47

want to take a quick look into this.

11:49

You know, in the end of like 2025, so if

11:52

this was like end of 2025 and this was

11:54

kind of the first half of 2026,

11:58

we got these platforms that kind of

11:59

looked very similar. They were kind of

12:02

like the evolution of Claude code

12:04

running in your terminal, and they've

12:06

evolved into these like super apps. So,

12:09

like even the Hermes agent looks a lot

12:11

like Claude, co-work, um Codex has GPT

12:15

work, which looks similar. Um the open

12:18

Claude desktop app looks a lot like this

12:21

where it's this kind of agent platform

12:23

where you have a bunch of sessions, you

12:25

have skills, plugins, you have

12:26

artifacts, automations, um yeah, which

12:30

are like scheduled, and they're starting

12:32

to look relatively similar. It is very

12:34

interesting to note that the the two

12:36

previous general agent platforms that

12:39

have gone viral, which are Buzz and

12:42

Grokbot, have a new shape to them.

12:45

Right here Grokbot, you can kind of like

12:47

see your team of AI agents. I showed you

12:49

that you can create new agents really

12:52

quickly, right? You can just create

12:54

agents, you can give it a purpose, and

12:55

each one has its own routines. And then

12:58

the platform that went viral before this

13:00

was Buzz. And so, this platform Buzz

13:03

allows you to create

13:05

channels with a bunch of different

13:07

agents, and you can even add people to

13:10

your Buzz. And so, Buzz looks exactly

13:13

like Slack, and it's kind of like this

13:15

agent native Slack or Slack meant to be

13:17

used with humans and agents, which I

13:20

find to be very interesting. And

13:21

publicly, Anthropic hasn't really talked

13:24

about co-work that much. They talked

13:26

more about Claude tag, which is their

13:28

new platform that allows you to

13:29

basically create agents inside your

13:31

company Slack. To my knowledge, you can

13:33

only use it with Teams, but this is kind

13:36

of their focus right now is creating an

13:38

agent for groups of people or

13:40

enterprises or small businesses. I feel

13:43

like that's kind of the shift that we're

13:45

moving into. So, maybe the first half of

13:47

2026 or the first you know, first 2/3 of

13:51

the year was about the personal agent

13:53

and the rest of this year is kind of

13:55

about how do you create your own

13:57

personal team of agents? And then in in

13:59

regards to Buzz and Claude Tag, how do

14:03

you add an agent so that your entire

14:06

team can get access to the same agent so

14:08

you can collaborate using AI agents.

14:11

Okay, so now let's discuss updates

14:13

coming out of Anthropic. Yesterday,

14:16

Claude announced that your Claude Chrome

14:17

sessions now carry over to desktop web

14:20

and mobile. Conversations are saved and

14:22

your skills and connectors work inside

14:25

the browser. Basically, what Claude and

14:27

OpenAI are doing is they are inserting

14:30

GPT work in the case of OpenAI and

14:32

Claude co-work in the case of Anthropic

14:35

directly into your browser. I actually

14:38

have both of them set up. I can use

14:42

Claude here. And what they announced is

14:44

I can say, "Please tell me more about

14:48

this." And whatever I type in here

14:52

is basically the same as using Claude

14:54

co-work. It has access to the same

14:56

connectors, the same skills, and

14:58

everything. And all of the chats, right?

15:02

I can view the history. All of these

15:04

chats will actually sync into my Claude

15:07

app. I can very easily move this

15:08

conversation over to the Claude desktop

15:10

app.

15:11

And you can see here, it's named this

15:13

chat more information request. And I can

15:16

go back to Claude and I can see that

15:18

more information request is right here.

15:20

So, I can very easily switch from the

15:23

chat that I had in any browser, right?

15:25

It's just a Chrome extension, back to

15:27

Claude. So, it's equivalent to coming

15:29

here and using Claude, except you can do

15:31

it directly from your

15:33

uh Chrome extension in Chrome. The next

15:36

update to Claude is pretty interesting.

15:38

Sonnet 5 had an introductory price, and

15:42

it was scheduled to increase back up in

15:45

price at a certain date, but Claude has

15:47

decided that it would actually stay at

15:49

this cheaper price. Open AI is doing

15:52

something similar with their Terra

15:55

and uh Luna

15:57

models. So, basically, their

15:59

non-frontier models are getting much

16:00

cheaper, and I believe, and many

16:02

believe, that this is due to the

16:04

pressure from Chinese models out of Deep

16:07

Seek Kimmy, uh z.ai, and then now also

16:10

Grok, right? Because if their middle

16:13

models are way more expensive than these

16:15

other alternatives that are actually

16:17

better, then people have no reason to

16:19

use them. So, this is causing them to

16:21

lower their prices, and I expect this to

16:24

continue. The non-frontier models from

16:26

Anthropic and Open AI will continue to

16:28

get cheaper if the competition from

16:31

China and in the US continues such that

16:34

their frontier models are cheaper than

16:36

their middle models, everyone's just

16:38

going to use these models, and they'll

16:39

never use the middle models like Sonnet

16:42

and Opus and Terra and Luna. And for the

16:44

final update regarding Claude, uh this

16:46

guy released phone harness. So, this

16:47

isn't actually coming directly from

16:49

Anthropic, but if you look look up

16:51

GitHub phone-harness,

16:53

you will find a repo which will allow

16:56

you to fully control your phone from any

16:59

agent, not just Claude code, but Claude

17:00

code, CodeX, etc. You can fully control

17:03

your phone with an AI agent. I haven't

17:05

tested this out. I just thought this was

17:06

really cool. Thought I'd share it really

17:08

quickly. Okay, so we were about to talk

17:10

about the brand new Deep Seek V4 model,

17:13

which was just released this morning

17:15

officially, and Deep Seek said that

17:17

we're uh launching DeepSeek V4 today,

17:20

and they claimed that it was nearly as

17:23

good as Fable. And here's a very quick

17:25

summary after summarizing everything and

17:27

everyone's takes on Twitter. I'm trying

17:29

to figure out this story, but here's a

17:31

concise summary here. So, apparently

17:34

DeepSeek released a new pro model last

17:36

night and talked like it was almost as

17:38

good as the top expensive models like

17:40

Fable.

17:41

And it was only a tiny fraction of the

17:43

price. However, many people started

17:45

testing this, and they came back and

17:46

they said, "No, it's only a little bit

17:48

better than DeepSeek's previous model,

17:49

which was DeepSeek

17:51

V4 Flash." And so, this was

17:54

embarrassing, and they have basically

17:57

taken the model down. In fact, if you go

18:00

to Cursor right now and you switch to

18:03

DeepSeek V4 Pro, I'm using this via Open

18:06

Router.

18:07

It's not an official model inside Cursor

18:09

yet, and if I say, "Hi," it actually

18:11

will immediately fail. And so, this

18:14

model just doesn't work right now, and

18:16

so I can't fully cover it because there

18:20

is some backlash. Apparently, it's not

18:22

that much better than DeepSeek V4 Flash,

18:24

and they're also getting some flak for

18:28

their pricing. So, apparently the

18:29

pricing has come up, and they went to

18:31

usage-based pricing. But luckily for us,

18:33

58 minutes ago Gemini officially

18:36

released 3.7 Flash. And apparently, this

18:39

is a very fast model, and it's brand

18:42

new. And I do notice here, though, a red

18:44

flag is that the models they're

18:46

comparing against, right? You can see

18:48

Gemini Flash is being compared to Claude

18:50

Sonnet and GPT Terra. So, this is the

18:53

third best model by Claude and the

18:56

second best model by OpenAI. That's what

18:59

they're comparing it to. And so, they

19:00

haven't released a new pro model in a

19:02

while, and so this 3.7 Flash is a fast

19:06

mid-tier model. And if you want to test

19:09

it out, again, I like to do it inside

19:11

Cursor. I have it running because I'm

19:13

using open router. And if you were to

19:16

sign up for open router and just ask

19:18

cursor how to set it up, you get it set

19:20

up in like 2 minutes inside cursor. But

19:22

I can say, "Hey." Or you could use it

19:24

inside the anti-gravity. I'm sure they

19:28

have it inside anti-gravity. You can use

19:30

the new Gemini 3.7 model. Since it's

19:33

only a mid model,

19:35

and it's not that good, it's just like

19:37

pretty fast. I don't think I'll be

19:39

testing it too much. But if you want to

19:41

test it out, you can. Okay, so for the

19:43

final big thing that I want to talk

19:44

about today is I want to talk about what

19:47

I call the chat verse work problem. Many

19:50

people are confused about how to use

19:54

AI agents, specifically the Claude

19:56

co-work and the GPT work features inside

20:00

these platforms.

20:02

Signal said, "It's not clear to me

20:04

that OpenAI realizes how strange the

20:06

distinction between chat and work feels

20:08

in practice and how fragmented the

20:10

entire chat GPT experience has become.

20:13

If you leave it on work, every simple

20:15

question or search becomes an

20:17

expedition. It starts thinking,

20:18

planning, and using tools when you just

20:20

wanted a quick answer." What he's

20:22

talking about is in chat, you now have

20:25

chat and work. And these are two very

20:28

different products. Obviously, you know

20:30

what chat GPT is. But GPT work is a

20:33

bigger thing, right? Here, you can

20:35

actually do things on Slack. You can

20:39

have it control your Gmail. You can have

20:41

it fully control your notion. You can

20:44

literally get, if you set up the right

20:46

scheduled tasks, you can have it fully

20:49

reply and send out emails, and you can

20:51

get it to run like it is a full agent

20:54

platform. And what he's talking about

20:56

here is just hard to understand the

20:58

distinction between chat and work. I

21:00

made a full video on GPT work. It's

21:02

incredibly powerful, but the distinction

21:05

between chat, work, and then the Codex

21:07

app is pretty confusing right now.

21:11

He goes on to say that there's no good

21:13

default. Chat is often too limited for

21:15

real tasks. Work is too slow and

21:17

cumbersome for normal queries.

21:18

Constantly switching between them makes

21:21

the entire product infinitely more

21:23

complex. Also, the lack of sync between

21:25

mobile and desktop plus what's a local

21:28

chat versus a cloud chat is a mess. And

21:31

again, these are things that I talk

21:32

about in my last video. It is relatively

21:35

confusing the difference between a local

21:37

chat and a cloud chat. Chat GPT went

21:40

from the most usable, simple consumer

21:41

experience to confusing AF in such a

21:44

short period of time. Pretty nuts. I do

21:47

think it's incredibly powerful, but I do

21:48

think he's talking about a very

21:52

difficult problem, which I call the

21:56

chat versus work problem. And I tweeted

21:59

this about the brand new Grokbot

22:01

yesterday.

22:02

I showed you the Grokbot platform

22:04

earlier. And then I tweeted this

22:06

specifically around how the platform is

22:09

set up. I said that I think Grokbot is

22:11

behind Codex and GPT work for a lot of

22:14

reasons. There's a lot of things that

22:15

GPT work can do that Grokbot cannot do.

22:18

But I will say their biggest innovation

22:20

is personifying the chat sessions. An

22:22

agent or a Grokbot is basically a named

22:25

session or a named chat session with a

22:28

mini system prompt. For example, on

22:31

Grokbot, if you just click on the name,

22:33

right, this description is just a mini

22:35

system prompt. And all of the plugins

22:38

that you can create, right, I can add

22:39

Gmail, that is a plugin. I can also use

22:41

skills. And all of these skills are

22:45

global to all of my agents. They

22:47

basically just made each chat session

22:49

have its own system prompt and your its

22:52

own name. So that when I want to create

22:54

a content script, I'll just go to my

22:56

content agent. Or if I want to scrape

22:58

from social media, I'll go to my content

22:59

agent. If I want to do my weekly update

23:02

agent, I'll go to my weekly update

23:04

agent. If I need something handled in my

23:06

partnership bot, um I will go to my

23:09

partnership bot. Not only that, but like

23:11

if something happens inside Slack in

23:13

this Slack channel, it will

23:15

automatically ping me. And it stays

23:18

organized by these chat sessions. And

23:21

so, I think the biggest innovation of

23:23

Grokbot was kind of making this analogy

23:26

easier to understand. And so, I think

23:28

we're going to see a lot of innovation

23:30

over the next 3 months as the frontier

23:32

labs try and figure out what is the best

23:34

setup for people to use AI agents in

23:36

their business. And I think there's just

23:38

going to be a lot of innovation. Because

23:40

this is something I would have never

23:41

thought I wanted, but once I used it, I

23:44

was like, "Okay, this actually makes

23:46

more sense." The automations should not

23:48

live at a global level, they should just

23:49

live within the bot. I have a feeling

23:51

that the other labs are going to copy

23:53

this design because I do find it's a lot

23:56

easier to get started, and it just

23:58

intuitively makes sense that you have a

23:59

bot. Each bot has its own computer, and

24:02

it has its own routines. I find it to be

24:05

an easy interface to pick up and

24:07

understand. Another thing is that

24:08

workspace agents are coming soon to chat

24:10

GPT. And I know this because I quote

24:13

tweeted their workspace agents. And so,

24:14

workspace agents are only on the GPT

24:17

team plans. It allows you to create

24:19

these like agents, right? You can create

24:21

agents. And these are as powerful as GPT

24:22

work, except they do have their own

24:24

system prompts. You can message these

24:26

custom agents, and you can even add them

24:27

to Slack and message them through there.

24:29

And it's only available to teams. And

24:31

so, I tweeted this. I said, "Wish this

24:33

wasn't only a Teams plan and only on

24:35

web. This should be on desktop as well."

24:37

And the Someone from the OpenAI team,

24:40

Andrew, said, "Yes." So, this indicates

24:43

that this is coming very soon to the

24:45

chat GPT desktop app, and it won't just

24:48

be a Teams plan, which I think is really

24:50

cool. This is This is the most slept on

24:52

OpenAI product yet, in my opinion.

24:55

And then finally, one update is that

24:57

codex you can get and download on Linux.

25:00

So codex the codex app or the chat GPT

25:03

app where I use codex

25:06

you can get this on Linux now. So that

25:08

is a new update on the codex side. And

25:11

yeah, that's basically everything. So

25:13

this was an agent native update. I'm

25:15

Riley Brown. Thank you guys so much for

25:16

watching and please like, please

25:18

subscribe. It helps you out a ton. I'll

25:20

see you here for the next video.

Interactive Summary

This video covers a major week in AI, focusing on the release of SpaceX's Grok 4.6 model and the introduction of their new 'Grok Bot' platform. The host discusses how Grok 4.6 is challenging current frontier models in terms of performance and cost-efficiency. A significant portion of the video is dedicated to the 'chat versus work' problem, where the host highlights how Grok Bot improves upon the user interface of existing tools like Claude Co-work and GPT Work by personifying chat sessions into specialized agents. The video also touches on updates from Anthropic, Gemini, and DeepSeek, alongside future developments expected in AI agent platforms.

Suggested questions

3 ready-made prompts