HomeVideos

Grok 4.6 is Fable now

Now Playing

Grok 4.6 is Fable now

Transcript

542 segments

0:00

So, huge news. Grok 4.6 just got

0:03

released as I'm recording this. But,

0:05

here's like the big headline. xAI caught

0:09

up with Anthropic. Now, the benchmarks

0:11

for Grok 4.6 High are showing that

0:13

they're equivalent or or close to Fable

0:16

5 Max. So, this is the first Grok

0:20

release that is at the time of release

0:22

is just neck to neck with the best

0:24

OpenAI model and the best Anthropic

0:27

model. 4.5 closed the distance and 4.6

0:31

but I'm right there neck to neck. Now,

0:33

of course, this was foretold, so to

0:35

speak. This was kind of predictable.

0:36

xAI's or Spec Space xAI's purchase of

0:40

Cursor probably had a lot to do with

0:41

this. In fact, as they say here, the

0:43

biggest difference between 4.5 and 4.6

0:46

was a much longer supplemental training

0:49

run. So, we believe that 4.5 and 4.6,

0:52

they are kind of the the V9 model or

0:55

base as it's referred to. It's a 1.5

0:57

trillion parameter base model and 4.6

1:00

had a much longer supplemental training

1:02

run versus 4.5. So, they used a lot of

1:05

the data from Cursor. And as they say

1:07

here, curated model generated data for

1:09

reasoning and advanced technical

1:11

concepts, high-quality engineering data,

1:13

and an improved optimizer and training

1:15

recipe. This produced a stronger

1:17

foundation for the supervised

1:19

fine-tuning and reinforcement learning

1:20

stages that followed. Now, of course,

1:22

most of us at this point view benchmark

1:24

numbers with some suspicion cuz we've

1:26

seen some places before kind of like

1:28

just trying to game the system, get high

1:29

scores on and those benchmarks, but the

1:32

models themselves kind of failed basic

1:34

tasks when people actually got their

1:35

hands on them. Is that the case here? I

1:37

don't think so. I've only played around

1:39

this model for a few hours at this point

1:41

and I got to say from my first initial

1:43

experiences just right out of the gate,

1:46

it's a beast. One of the first prompts

1:48

that I did was I asked to kind of do a

1:50

clone of Portal 2. So, kind of being

1:52

able to fire different portals, blue and

1:54

whatever, orange, right? You walk into

1:56

one, you walk out the other. Something

1:57

that a lot of the previous models up to

2:00

very recently just weren't able to do. I

2:02

used Grok Build to do this. It worked

2:04

probably for two to three hours or

2:06

thereabouts, and I don't want to say one

2:08

shot cuz I had a few suggestions and I I

2:10

steered it a little bit during the

2:11

process, including asking it to add some

2:14

sound effects and voice overs with 11

2:16

Labs, but pretty much yeah, it one shot

2:20

one sort of level from Portal 2. Created

2:23

all the mechanics, created all the

2:25

models, which you're going to see in

2:26

this demo in just a second. You know,

2:28

all the three objects it's created. The

2:30

conservation of momentum, the ability to

2:32

see through the portals, the ability

2:34

ability to see the character on the

2:36

other side of the portal. It's important

2:38

to understand that like this is not a

2:41

low-end model. This feels like a

2:43

high-end smart model. It created the

2:46

entire room in one shot, including the

2:49

puzzle that you're supposed to solve for

2:50

multiple elements. So, take a look. So,

2:53

this is chamber 07. Grok 4.6 build,

2:56

basically a replica of Portal 2.

2:59

>> Chamber 07, dual gate certification.

3:02

Please proceed.

3:03

>> And it's looking pretty good. We've got

3:05

reflections, we got lighting. Test

3:07

chamber 007. Blue gate.

3:09

>> Blue gate established.

3:10

>> Select the blue gate, and here is the

3:14

>> Both gates are live.

3:15

>> Let's see here. Oh, wow. All right. So,

3:18

we have a character model with

3:19

reflections. We're carrying that gun.

3:22

Uh that is pretty nifty. Okay, let's go

3:25

through here. Oh, wow. Okay, that's

3:27

pretty smooth, I got to say.

3:29

All right. So, if we wanted to go

3:31

through here and have my character here.

3:34

Wow, okay. So, I can see myself there.

3:35

So, what happens if I walk through it?

3:36

Oh, hey. That is

3:39

>> one, emerge from the other. Momentum is

3:42

conserved.

3:43

>> is conserved as we go through it.

3:47

>> Orange gate established.

3:48

>> Oh, look at that. Okay, so that thing

3:50

should

3:51

I did not expect that to happen.

3:53

I mean, it's supposed to happen. I just

3:54

didn't expect it to.

3:56

Okay, so this is our

3:57

>> Wait, you mad at Cuba quiet. Do not

3:59

introduce it to the acid.

4:01

>> Do not introduce it to the acid. The the

4:03

green bubbly stuff is the acid. Okay, so

4:05

basically I guess what we want to do is

4:08

slide this kind of like what is that

4:10

game where you slide the thing across

4:12

the ice? Loogie? Loogie? Something like

4:14

that. Okay, we go through here and as it

4:15

slides

4:16

>> Yes.

4:17

>> Hey, that is pretty good. I got to say.

4:19

I'm statistically a success apparently.

4:21

Well, thank you. So, there's a lot of

4:23

really incredible and amazing things

4:25

there. It's very impressive. Of course,

4:27

keep in mind this was just the first I'm

4:29

pretty sure this is the first prompt I I

4:30

threw its way. I was trying to see if it

4:32

works in Grok bot, which is the other

4:35

big release from yesterday. But, I

4:37

wasn't able to confirm that it was yet

4:39

available in Grok bot. So, I did Grok

4:41

build instead. That's where it built all

4:43

of this. And it was built with Grok 4.6

4:45

on the high settings. Later I realized

4:47

there's an extra high setting. So, I I

4:48

do want to test around it with that, you

4:50

know, turning up to 11. But, this was

4:52

Grok 4.6 high and this is very similar

4:56

to what I would expect Fable 5 to come

4:59

up with on on the first attempt. So,

5:01

again, Fable 5 level model from xAI.

5:05

But, wait, there's more because pricing

5:07

starts at $2 per million input tokens

5:09

and $6 per million output tokens.

5:12

Additionally, there's a fast variant,

5:14

which is twice the price. My next test

5:16

that I'm running within I'll post it

5:17

more about this later. So, I'm trying to

5:19

get it to build from the ground up fresh

5:21

without reusing any of the existing code

5:23

cuz I tried this before for a different

5:24

model. So, this is a starting fresh.

5:26

Here I'm using cursor by the way to to

5:28

do this. Important point on cursor and

5:30

Grok build because right now during

5:33

launch week, so that's right now, you

5:35

have double the usage for these new

5:37

models if you're using cursor or Grok

5:39

build. So, if you have those, do take

5:41

advantage. And so, here in cursor I used

5:43

a Grok 4.6 extra high fast. So, again,

5:46

fast is that model that it's twice the

5:48

speed, but it uses up twice the usage.

5:51

Right now, you get, you know, twice the

5:53

the usage, so it kind of evens out. But,

5:55

here I'm trying to get it to do a fully

5:57

AI automated run kind of streamer. So,

5:59

in this case, it's going to be playing

6:00

Pokémon Red and playing the game using

6:03

voice to narrate what it's doing.

6:05

Probably a HeyGen avatar kind of a video

6:07

avatar that's going to explain

6:08

everything while it's playing the game.

6:10

It's also going to be able to talk to

6:11

chat. And so, for that, I'm making sure

6:13

that it adds some sort of a censorship

6:15

so that people don't jailbreak it, which

6:17

could be bad. So, we'll see. We'll see

6:19

if this model survives contact with, you

6:21

know, you all people. So far, the

6:23

results are good. So, here's kind of

6:25

what we have so far. Just keep in mind,

6:27

this is a bare-bones prototype. If I

6:29

click start game, so notice it loads up

6:32

Pokémon Red. It goes quickly through the

6:34

initial screen, and I know it's got past

6:38

this right now. And so, right now, it

6:39

doesn't have all the logic and all that

6:41

stuff quite yet. Oh, the last time it

6:43

did get out of the house. Right now, it

6:44

seems to be stuck on something. But, the

6:45

point is, on its first attempt, it was

6:47

able to take a Pokémon Red emulator,

6:50

hook it up to make sure that it's able

6:51

to have enough of a harness to control

6:53

the game. It created a fake chat. It

6:55

created text-to-speech. So, that little

6:57

model that's over there, as it's

6:58

talking, that gets gets a real

7:00

voiceover. And so, all this was built

7:02

within the first hour or two of Grok 4.6

7:05

going live. So, again, I'm not claiming

7:07

anything about this model yet. I'm

7:08

saying just straight out of the gate,

7:10

this thing is feeling pretty beefy. So,

7:13

looking at these numbers, I'm not

7:15

doubting them. I think this is kind of

7:17

on the level that it's going to operate

7:18

at. This is Cognition posting about

7:20

Frontier Code 1.1. So, notice Grok 4.6

7:23

is right behind the Fable 5 and Opus 5

7:27

series of models. It's it's up there.

7:29

What caught my eyes, Elon responded to a

7:31

few of these messages saying that Grok

7:32

4.7 will exceed all current models. So,

7:36

the next release, 4.7, which we're

7:38

expecting in 3 to 4 weeks, will be

7:41

number one on the [clears throat] chart

7:43

assuming no one else releases new models

7:45

which is not going to happen. But as of

7:47

right now if everybody else was stuck in

7:49

time for three to four weeks, it would

7:50

be number one on the leader boards. And

7:52

for that model they're adding a massive

7:54

amount of SpaceX company data in

7:57

supplemental training. Okay, so so far

7:58

the only thing that's kind of like is

8:00

confirmed so far is what we know about

8:02

Grok 4.6. Some of the things that are

8:03

coming in the future, these are

8:05

speculations of certain announcements by

8:07

Elon, some of the kind of rumors that

8:09

are floating around. But we're expecting

8:10

4.7 to be a 2.1 trillion parameter

8:13

model. Again, that's not verified by

8:15

official sources as far as I know, but

8:17

here is the important thing to

8:18

understand. So, 4.5 wasn't that long

8:21

ago. Now we got hit with 4.6. We have

8:24

4.7 coming out in say three to four

8:27

weeks, and Grok 5 is targeted to be

8:30

released before the end of the year. So,

8:32

as you can imagine, that's a kind of a

8:33

crazy release schedule. Why is this

8:36

happening? xAI reorganized engineering

8:38

to ship major updates every two to three

8:42

weeks, it seems like. So, kind of keep

8:43

all this in mind, right? So, they caught

8:45

up, they got to Feat of Strength level,

8:47

and they are just gunning to release,

8:49

release, release. The pricing is $2 per

8:52

million input, $6 per million output, so

8:54

it's a half the price of comparable

8:56

OpenAI and Anthropic frontier models.

8:59

And we also have this SpaceX company

9:00

data that's going to be used in

9:02

supplemental training, but there's more,

9:05

and this is kind of the the other side

9:06

of the coin. This is the other big

9:08

thing. So, this is Grok Bot. And if I

9:12

understand correctly, 4.6 isn't running

9:14

this quite yet. I I did a few tests

9:16

here. I I don't think 4.6 went live on

9:18

this thing yet. Now, as you recall from

9:20

before, we talked about the different

9:21

benchmarks. One of them was GPT Val. So,

9:24

GPT Val takes real jobs, real projects

9:27

that's done by real people out there

9:29

that work in specific industries, and

9:31

these models are asked to complete those

9:32

projects. And then real humans sort of

9:34

like verify. So, these humans have a lot

9:37

of background working in those

9:38

industries, and the industries are from

9:40

engineering to like AutoCAD to some

9:43

hospitality and travel industry like a

9:45

travel agent. There's finance, video

9:47

editing, like it's a broad range. It

9:49

really represents a large part of the

9:52

job market. And the 4.6 model beats out,

9:56

you know, Fable 5 and everything else.

9:58

Why is that important? Because it goes

9:59

hand in hand with this with Grok Bot.

10:02

So, Grok Bot is xAI's always on, always

10:06

online, always available agent. So, it's

10:08

on desktop apps. I believe it's on

10:10

Linux, Mac OS, it's on Windows with

10:13

Android versions coming soon. And the

10:15

whole pitch is these are your AI

10:17

teammates. So, this is not like a

10:19

chatbot. These This is a team of agents

10:21

that work for you. One of the first

10:22

things that it kind of suggested you do

10:23

is you create, you know, let's say a

10:25

chief of staff. That's the role that it

10:27

kind of recommends. And so, I I jump in

10:29

here and I give it this this task.

10:30

Initially, I was want to see how far I

10:32

could take that AI streamer. So, I told

10:33

it, "Do this. Create whatever agents you

10:35

need." It does. So, everything here,

10:37

most of these agents except one, it

10:39

created automatically by itself. Then it

10:41

begins running the various tasks and

10:44

handing off the sort of the delegates

10:46

work to those sub-agents. So, you're not

10:48

talking to the million of AI agents

10:50

working in the background. You're

10:51

talking to your chief of staff or

10:53

whatever you want to call it. It

10:54

delegates, it thinks about it, it asks

10:56

you questions, but those questions are

10:58

might be the ones that are surfaced by

10:59

the agents. You're You have one point of

11:01

contact. Or you can have multiple points

11:02

of contact, but the point is that you're

11:04

not talking to every independent agent

11:06

on here. If you're wondering, "Well, how

11:07

could this be always on? Isn't this on

11:10

your desktop?" Well, no. Each of these

11:12

agents, Grok Bot, it has its own virtual

11:16

machine in the cloud that it can use 24

11:18

hours a day. You can actually see the

11:20

screen here. So, here it is. This is

11:21

what's available for it to use. So, you

11:22

can open up Chrome, terminal, whatever

11:25

you want. What's interesting about this

11:26

is there's actually a button to teach it

11:28

a task. This is kind of brilliant

11:31

because there's going to be a lot of

11:32

things that these AI agents have to

11:34

navigate that they're going to struggle

11:35

with. I had one of my AI agents not

11:37

through Grok, it was I think it was

11:38

OpenAI or Claude, I I forget which, but

11:40

I had it go online and go through an

11:42

online messaging app like WhatsApp or

11:44

Slack or what whatever it was to try to

11:46

see if somebody had mentioned something.

11:48

While I was doing that, it accidentally

11:50

sent a message. Now, this wasn't a big

11:51

deal, it caught the error immediately,

11:53

deleted it, and the message was just the

11:54

command it was run to to try to search

11:56

through all the messages by keyword. So,

11:59

it was something like search whatever

12:00

keyword, right? So, it wasn't a big

12:02

deal, but it could be bad in in a

12:04

different circumstance, right? Because I

12:05

didn't ask it to message people, but

12:07

with these teach a task, I can actually

12:09

show it, it gets recorded, and then

12:11

Grok, these bots, they're able to

12:13

execute those skills. Now, of course,

12:15

this is great if you're trying to work

12:16

with it, and it's kind of great for XAI

12:18

because it's basically a lot of people

12:20

kind of contributing data to these

12:23

models. And by the way, of course, you

12:24

have the choice to to opt out or not, so

12:26

that's entirely up to you, but for the

12:28

people that choose to opt in, of course,

12:30

their data can be used to further

12:32

improve these models. One other agent

12:33

that I built was my X news real-time

12:36

person that went through and basically

12:37

checks real-time news on X. So, X slash

12:40

Twitter, it goes out there and it starts

12:41

searching for me. It's super fast at

12:44

finding those news, so I just ask a

12:45

question, you know, like, "Did So, I

12:46

just ask it questions like, "Did anybody

12:48

talk about this? Did Elon say something

12:50

about this? Did somebody blah blah what

12:52

whatever you want?" It goes through, it

12:53

searches it, and it just explains it to

12:55

you. So, one of the things I asked it

12:55

is, "Grok 4.6, is it live on Grok bot?

12:58

Is this where we're using it?" And I

13:00

asked the other model, too, but it

13:01

wasn't able to say which model it's

13:03

running on, so I don't think as of that

13:04

time Grok 4.6 was live here yet as far

13:08

as I know. So, it found the official

13:09

posts, it found some comments by Elon,

13:12

and it said that so we're not 100% sure,

13:14

it's likely coming very very soon. Do

13:15

you want me to watch X for when they

13:18

turn it on?" I said, "Yes." It

13:20

automatically created routines and it's

13:21

going to keep checking, you know, three

13:23

times a day and for the next 10 days,

13:25

after which he'll notify me one way or

13:26

another. So, this is kind of awesome cuz

13:29

now I have this bot whose job it is to

13:31

keep searching X and when we get new

13:33

updates, it notifies me. Since this is

13:35

always-on, on 24/7, it's not like when I

13:38

shut down my computer, it goes to sleep

13:39

and forgets it. This will notify me for

13:42

sure. So, the big point here is that

13:44

these are meant to be more like your

13:45

colleagues and coworkers that you

13:48

delegate work to. So, this is kind of

13:51

digital labor, right? So, this has been

13:53

positioned by XAI as digital labor as

13:55

actual workers. So, a company that has

13:57

access to this basically is able to get

14:00

a lot of the cognitive labor handed off

14:02

to these agents that are online 24/7,

14:04

that are able to do research, you know,

14:05

when you're sleeping, put together

14:07

projects, hand them back to you. And

14:09

this is kind of I would say how they

14:11

design it is innovative. Now, not like

14:14

where they're breaking new frontiers or

14:15

whatever, but there's a lot of little

14:17

things here and there that just seem

14:18

like when you start using you're like,

14:19

"Oh, yeah, this feels right." One is

14:23

kind of they pop up with these boxes

14:24

more and more where you get to choose

14:26

your own answer. Your initial impression

14:28

might be like, "Oh, that's for, you

14:29

know, people that are just starting out.

14:30

Maybe they don't quite know exactly what

14:32

they're trying to do." But working with

14:33

this, I realized that there's a lot when

14:35

you try to go fast, this is very

14:37

helpful. Cuz when you give it a project,

14:38

it might have a bunch of questions for

14:40

you. Each time it asks you a question,

14:42

you have to kind of sit there and and

14:44

read it and think about it and then type

14:45

out an answer or just speak an answer.

14:47

There's been a number of times when it

14:49

shows me something like this and just at

14:51

a glance, I kind of know, "Oh, yeah, I

14:53

know exactly which one of those options

14:54

I want." Sometimes you don't even have

14:56

to really consciously read the question.

14:58

You get what it's asking you. So, at

14:59

some point it asked me this question,

15:01

right? And the answers were like, "Add

15:02

Heygen or just use text-to-speech." Like

15:05

I as soon as I glanced at it, I knew

15:07

exactly what it was asking without

15:08

needing to read the question. It was

15:10

going, "Do you want to build out the

15:11

video avatar or just leave that for

15:12

later?" Like I knew and I was like,

15:14

"Yeah, go ahead and build that out."

15:15

Things like this might seem tiny, but it

15:17

makes things go a lot faster. It's less

15:19

of a cognitive load, right? You're not

15:21

wasting cycles having to read the I mean

15:24

in your own brain, so to speak, reading

15:25

through everything. You're kind of like

15:26

cuz you know what you want. The model

15:28

doesn't. And so, any little polish we

15:30

can add to make that transfer of

15:32

knowledge faster and smoother and takes

15:34

less effort, it's a big deal, especially

15:37

as more and more work begins to get done

15:39

like this. So, love this so far. The X

15:42

News AI agent that adds the routine and

15:45

then just says, "Yep, I'll go ahead and

15:46

I'll start checking it three times a day

15:48

for the next 10 days." I didn't ask it

15:50

for that. It says here, "Do you want me

15:51

to watch X for when they flip it?"

15:53

Right? When it goes on. Yes. It's like,

15:55

"All right, got it." And it adds the

15:56

routine. This is beautiful. This is one

15:59

of the things where I like this so much

16:01

better than Fable 5, as much as I love

16:03

Fable 5, but the way that it talks half

16:05

of the time, half the answers, I'm like,

16:07

"What are you saying?" I mean, at this

16:09

point there's like entire memes by it.

16:10

It's got this weird way of speaking, you

16:13

know, and yeah, it's kind of funny, but

16:14

after a while it kind of gets old. It's

16:15

just like just tell me what I need to

16:18

know. What are we talking about? So far,

16:21

going through these and working with

16:22

Grok, it's just so much faster. You just

16:24

kind of fly through it. So, definitely

16:25

encourage everybody to download this.

16:27

It's super simple to set up. When you do

16:30

so, create your kind of chief of staff

16:32

or whatever name you want to use for it,

16:33

then pin it to the top, and then try to

16:36

use this as much as possible, and have

16:37

it create its own structure underneath

16:39

it so that you're sort of talking to

16:40

just this thing. And then and maybe

16:42

there's other sort of side agents like

16:44

this X News one that you have off to the

16:46

side that that it's its own thing. I

16:47

don't need it going through the chief of

16:48

staff, but for a lot of this work, have

16:51

one agent in charge of it. And of

16:53

course, this is easily cross-device,

16:55

right? So, I can go from here to using

16:57

it on my phone to wherever. I'm not

16:59

relying on the computer being on at any

17:01

given moment cuz it has its own

17:03

computer. So, between Grok bot and Grok

17:06

4.6 coming out and the new models that

17:08

are coming out, 4.7 and Grok 5, I mean,

17:11

those are not out yet, but so far it

17:12

looks like XAI is just in a great great

17:16

position to start winning. Check it out

17:18

within Grok bot if you have access to

17:20

Cursor. You get some double the usage

17:22

during launch week or if you're using it

17:24

with Grok build, same thing. But, I got

17:27

to say this is

17:28

this is kind of exciting. It seems like

17:30

Google is lagging slightly behind, but

17:33

we still have, you know, multiple

17:35

frontier models because here comes xAI

17:37

to fill in Google's spot. So, congrats

17:39

to xAI, congrats to Cursor. I was always

17:42

very impressed with Cursor and all the

17:43

stuff that they've been able to do. So,

17:45

them joining forces with xAI, which has

17:48

just incredible amounts of compute and

17:50

and talent and capital in order to

17:52

really put these two things together. I

17:54

feel like they're going to have a lot of

17:55

success in the future. But, let me know

17:57

what you think. If you've had a chance

17:59

to do more testing of Grok 4.6 more than

18:01

I have so far, let me know. What do you

18:03

think of it? Is it, in your opinion, as

18:05

good as Fable 5 and GPT 5.6? Will you be

18:09

using Grok bot? Let me know in the

18:10

comments. If you made this far, thank

18:12

you so much for watching. My name is Wes

18:13

Roth. I will see you in the next one.

Interactive Summary

This video details the launch of xAI's new model, Grok 4.6, highlighting that it has reached parity with top frontier models like Fable 5. The presenter demonstrates Grok 4.6's capabilities by building a playable Portal 2 replica and discusses the introduction of 'Grok Bot', an agent-based platform designed for continuous, autonomous task management. The video also covers the aggressive release schedule xAI has adopted, including upcoming versions 4.7 and 5, and the advantages of integrating with Cursor for development workflows.

Suggested questions

3 ready-made prompts