HomeVideos

GPT-5.6 Feels Like the Beginning of AI 2.0

Now Playing

GPT-5.6 Feels Like the Beginning of AI 2.0

Transcript

624 segments

0:00

GPT5-6 was released by OpenAI, and it is

0:03

a great model. But, even better than it

0:06

being a great model, I think it takes us

0:09

out of the turn-by-turn work with AI.

0:12

What I mean is I think we're seeing a

0:14

brand new bend in the curve. We're

0:16

seeing one of those transformation

0:17

points where we start using AI in a

0:20

truly different way because of the

0:22

capabilities that have been given to us

0:25

with this kind of incremental increase

0:26

that we're seeing. I want to go through

0:28

what GPT5-6 is, how strong it is, the

0:31

three different variants, Luna, Terra,

0:33

and Sol, how they compare against

0:36

something like Fable or even Sonnet, so

0:38

that you kind of understand where it

0:40

fits and which one you might use, and

0:41

whether or not you should even be

0:42

interested in it. But, I really do also

0:44

want to cover what it means, I think,

0:46

going forward. This is a truly important

0:49

moment, I believe. These models, Fable

0:52

and 5-6, these are really that moment

0:54

where we're saying, "We have new

0:56

capabilities. We don't know how to use

0:58

them yet, and they're a little bit raw,

1:00

but they're going to give us something

1:01

brand new." So, we're going to go

1:03

through and build this. I'm going to

1:04

kick this all off with something really

1:05

kind of fun that I haven't even done

1:07

yet, so we'll see how it works out.

1:09

We're going to kick off a new request

1:10

against Fable and GPT5-6 Sol to see who

1:14

can build the best version of this

1:15

application, or not even best, just see

1:18

what both of them come up with with one

1:20

big build. We'll let that cook the whole

1:21

time we're going through the video, and

1:23

then we'll take a look at it at the end.

1:25

I hope. In any case, let's dive right

1:27

in. Okay, so very quickly, I'm going to

1:30

show you I'm in the Codex application

1:32

right now. For now, what I'm going to do

1:34

is I'm going to put in a new request

1:35

here. This really is pointing at a file

1:38

that's inside of this repo that we're

1:40

in. I'm also putting it in the work

1:41

tree. So, what I'm going to do is kick

1:43

that off, and this is using uh 5-6 Sol

1:46

extra high. So, this is GPT5-6. Okay.

1:49

So, I'm also going to do this in Claude

1:52

Code against Fable 5 extra high, and get

1:55

both of those kicked off and running

1:57

while we're sitting here doing this

1:58

video. So, let's dive into what they

2:00

released or what they announced with 5

2:03

6. All right, I'm giving myself only

2:06

about a minute maybe a minute and a half

2:08

to talk about this. GPT 5 6 has

2:10

released. There are three different

2:12

sizes. There is Saul, the big one.

2:14

That's the big big one. Terra, which is

2:16

the middle-sized model. And Luna, which

2:18

is the smallest model of them all. They

2:20

of course kind of degrade in their costs

2:23

in that same kind of way. Luna is much

2:24

cheaper than Saul. Okay, better better

2:26

faster faster cool neat very capable

2:29

faster cheaper better better better at

2:30

following instructions. Okay, and to

2:33

round it out. I know I know it's crazy.

2:35

It's the same thing every single time.

2:36

We'll do a little bit of measurement in

2:37

a second so you can see what it's done.

2:39

I want to point out though Saul, the big

2:41

one, is a 530. And that's kind of $5 per

2:44

million input 34 million token output.

2:47

And Terra is a 2 and 1/2 15. And Luna is

2:50

a 1 6. I would definitely take a look at

2:53

Luna. I wouldn't normally say that, but

2:56

it's a cheap enough model and even Terra

2:58

is measured at a price cost here. 2 and

3:00

1/2 and 15 is roughly the Sonnet price

3:03

cost. So, it's really interesting to

3:06

compare Terra against Sonnet. And Terra

3:09

really kind of outpaces Sonnet in all of

3:11

these charts in a major way. So, that

3:13

cost is kind of interesting. Okay, let's

3:16

briefly go through my own benchmark. I

3:18

run a benchmark. I say this every single

3:20

time. I give it a big request that has

3:22

like a hundred items in the request and

3:24

I also really put in some meaningful

3:26

details that are not necessarily

3:28

requested items, but uh durable intent.

3:31

The reasons I would need these things to

3:33

be the way that they're being requested.

3:35

And so, on the x-axis is the item

3:37

itself. I asked for a button. Is button

3:39

in the final plan? Not necessarily after

3:42

it's built just just the first step. So,

3:44

this is does it survive first impact? On

3:46

the y-axis is kind of does my intent,

3:49

the reason, survive the plan? So, what

3:51

we'll see this is the previous

3:52

information. Here was GPT-5.5, previous

3:54

model, Opus-4.8, pretty close to it, and

3:57

Fable-5 hanging up here. So, Fable-5

4:00

really did a great job. Let's come in

4:02

here and add some of the new models. If

4:04

we look at Luna, you'll see where Luna

4:06

hangs out. Really fantastic. Great on

4:09

planning, not bad on uh intent recovery.

4:12

So, the Y is still a little bit low,

4:14

quite honestly, but it is a much, much

4:16

smaller model. And remember, Sonnet down

4:19

here, Sonnet-5, the brand new Sonnet

4:21

that released, is hanging out

4:23

languishing way down here on my chart.

4:25

So, let's take a look at the Tera model,

4:27

which is supposed to compare, and you'll

4:29

see Tera way up here. So, Tera is doing

4:32

a great job, quite frankly, at uh

4:34

basically maxing. We're at 81.4%. This

4:37

is the highest that I've seen for making

4:40

the passing the intent through, being

4:42

aware of the intent. And this is three

4:43

runs for each one of these models at

4:45

different levels, and I think these are

4:46

all extra high for each one of these.

4:48

So, it's really worth saying this is a

4:50

fantastic model. And the last big one

4:52

that we're all wondering about, which is

4:54

the big behemoth, Saul. How did Saul do?

4:56

Saul is sitting right in here. So, let

4:58

me zoom in a little bit so that we can

5:00

see the family that we're dealing with.

5:02

Saul, GPT-5.5, and then Fable-5. So,

5:05

Saul and Fable-5 have roughly the same

5:08

intent recovery, but Saul doesn't have

5:10

as much planning recovery. But then,

5:12

we're really only talking about a point

5:14

here, and I on any given one of these

5:17

axes, if I'm if I'm really talking about

5:19

anything less than a point, I would say

5:20

my fidelity is probably not tight enough

5:22

to care about that. So, that cluster

5:24

Saul very close to Opus and Fable,

5:28

Opus-4.8. So, really, really great

5:30

performance. Let's see how they do.

5:32

Okay, first, let's check in on what

5:34

everybody's working on in the the big

5:36

prompt that we put in, the game that

5:38

we're asking each one of them to build.

5:39

You can see um Saul and all of the GPT

5:43

models, including uh ChatGPT, which is

5:46

now what we call Codex Desktop, Chat GPT

5:48

Desktop, uh which has work and Codex

5:52

inside of it. Like I said, different

5:53

different video for that. Actually does

5:56

computer use and evaluates the the

5:58

browser itself. It actually has a

6:01

browser integrated over here that we

6:02

would be able to see it if we were

6:04

looking at it over here. So, this is

6:06

actually what it's showing us a picture

6:07

of and it's running through the system

6:09

right there. Same thing over here for

6:12

where Opus or where Fable is, we can

6:15

open the browser and see if they they

6:17

have not yet taken over to use the

6:19

browser to take a look. I'm going to

6:21

guess that they're not actually working

6:23

on the UI of it just yet. Fable so far

6:25

is at 137,000

6:28

tokens and Saul is at 120,000 tokens.

6:32

So, really really equitable. Let's let

6:33

those cook. Now, let's go take a look at

6:35

the Saul model and the others and how

6:38

they built a previous application design

6:40

that I've put together so that you can

6:41

really see apples-to-apples comparisons.

6:44

Okay, so I have an application that I

6:46

get new models to kind of build and it's

6:49

a native application that has a bunch of

6:52

design assets, both images as well as

6:55

true design systems and assets inside of

6:57

it that should describe the application

7:00

itself. I will show you the finished

7:02

product in a second so that you kind of

7:04

understand the fidelity that we're going

7:05

for and then what I'm going to do is

7:07

take you through the builds of each one

7:09

of the models, showing you what's best

7:11

about each one of those models. What's

7:12

probably important to understand very

7:15

simply is this is a native application.

7:17

So, it's kind of highly performant. It

7:19

uses kind of advanced techniques for

7:22

dealing with animations and blends and a

7:25

bunch of things like that. But, that

7:27

isn't something that's going to be super

7:28

easy to understand through a simple

7:30

video like this. But, what I'll show you

7:32

is what really is good for some of the

7:34

models and really a distinction with the

7:37

exact same model, a better way to get

7:40

better results or at least in my

7:41

experience. All right, let's just take a

7:43

look. Okay, so let's kick off the video

7:46

for Luna, and you can see it will launch

7:49

the application and do a circle around

7:51

the outside ring, and then a circle

7:53

around the inside ring. It is pretty

7:55

complete. I will pause it here so that

7:58

we can see what the application looks

7:59

like. The lines are clean. It has all of

8:02

the all of the circular parts to it.

8:04

This is a great start. Next one, let's

8:07

dive into Terra. Same concept, do a ring

8:10

around the outside, and then around the

8:12

inside.

8:14

And you can see none of the animations

8:16

are in place. Let me show you what the

8:18

application

8:20

should actually look like. So, this ring

8:22

here is Terra, and if I bring up the

8:25

actual application as built, this is an

8:28

application that I finished out. So,

8:31

this is not the first single shot

8:33

experience that I got back, but this is

8:36

from Fable, I think extra high gold. So,

8:40

a while ago I built this using Fable,

8:42

and it got me very very close to this.

8:45

And if I go around the outside, you'll

8:46

notice the smooth animations between

8:48

each one of the movements. And then the

8:50

same thing with the inside, you can see

8:53

some of the moving icons as they jump

8:56

around from position to position, as

8:58

well as the elegant movement between all

9:00

of the the rings and the selections. So,

9:03

that's what we're really looking for

9:05

here, and certainly not what we're

9:07

seeing when we look at the build of

9:10

Terra. You can see it's just popping

9:12

around the outside, it's popping around

9:13

the inside, and even worse, it's worth

9:16

pointing out that the bays or the

9:19

sections are not centered on the icon.

9:21

So, this icon should be inside of this

9:24

bay, or moreover, the selection and the

9:26

lines on the graph itself that you might

9:29

not be able to see on this the dial

9:31

itself are going through the center of

9:33

the items instead of the center of them

9:36

as bays. And okay, so that was the small

9:38

and the medium. Obviously, what we have

9:40

to do next, let's take a look at the

9:43

large. We will start here.

9:50

None of the GPT-56 models had any of the

9:53

animations that were necessary. And

9:56

that's curious because all of the Fable

9:58

and Opus models did have this. Unlike

10:02

these GPT-56, they did deal with the

10:05

animations between the different

10:06

elements. So, I would say the winner out

10:08

of all of these is all of them or kind

10:11

of

10:12

sort of none of them. None of them are

10:14

terrible. Okay, and let's round this out

10:16

quickly by showing you one of the things

10:18

that's really different. So, when I do

10:21

56 soul,

10:23

but I use the exact same prompt that I

10:25

asked for the first one to be built and

10:27

instead use {slash} goal in front of it.

10:30

So, I set this up as just a goal with

10:33

the exact same request that I was giving

10:34

the others. This is the outcome. And

10:37

this already, you can see very quickly,

10:39

has a much better better design style to

10:42

it. Simple, light, straight lines

10:44

between everything. Really good job. So,

10:47

this is a huge step forward. And all I

10:49

added with {slash} goal. And this is

10:51

something that I want to call out. This

10:53

{slash} goal concept is very, very

10:54

strong and it's worth you giving it a

10:57

shot in some of your builds. If you're

10:59

not quite getting the results you want

11:01

or you're always finding that you have

11:03

to push back in a couple times to get it

11:05

closer to what you asked for, the

11:07

{slash} goal, I find in both this and

11:09

Fable and all of the other models that

11:11

I've tested, always seems to perform a

11:14

little bit better or very, very

11:16

frequently at least. Okay, and so when I

11:18

talk about {slash} goal, the one thing

11:20

that might come to mind, and I think it

11:22

does for a lot of people, is if I'm

11:24

going to do {slash} goal, what is that

11:26

going to do to my token count? So, as

11:29

you might see up here is soul itself.

11:32

The Soul straight is what I call it.

11:35

Took 31 minutes. Soul with goal took 28

11:38

minutes. So, 3 minutes less and 700,000

11:41

tokens versus 550,000 tokens.

11:44

That's a big difference. And the goal is

11:47

both faster and fewer tokens. Taking a

11:50

look at Here is Fable, Fable 5, and

11:53

Fable 5 gold. So, you can see Fable 5

11:57

took 27 minutes. Fable 5 gold took 48

12:03

minutes. So, just putting gold on Fable

12:05

exploded it. Now, again, these are

12:07

non-deterministic, so it would probably

12:10

build at different times. If I had built

12:12

this three or four times each, I'd get

12:13

better numbers. These are single-shot

12:15

definitions. The build is only done

12:17

once, to be fair. So, you have to kind

12:19

of grain of salt this and kind of take

12:21

this with a trend line. And the one that

12:23

I would call out is this is the one we

12:25

were just looking at. And this is the

12:26

one that you've heard about the release.

12:28

So, this is GPT-56 Soul. It's the big

12:30

new GPT model.

12:32

When you put gold in and it got much

12:34

better, remember it got markedly better,

12:37

it's both

12:38

faster, so it took 28 minutes instead of

12:42

31 minutes, as we noticed, and it's also

12:45

using less tokens. And that's what we

12:46

noticed. That does not hold across the

12:48

board. The 5-6 Terra gold here took a

12:52

little bit longer and used more tokens

12:54

than the 5-6 Terra here, but it did,

12:57

once again, have a better

12:58

representation. So, it's really worth

13:00

saying /gold is not what you think it

13:03

is. It is not an absolute token burner

13:06

that will do 5-10 times the amount of

13:08

tokens that you're currently using. I

13:10

have not experienced this. I will say,

13:12

it is worth saying, the Fable here,

13:14

where it took 280,000 tokens, it only

13:18

did it for 28 minutes and was not a very

13:19

good fidelity. In fact, this is the one

13:21

that things didn't line up. It was

13:23

actually the only build that I have

13:25

that's kind of broken. And but the five

13:28

the fable five goal one it is the one

13:30

that I ended up pushing forward because

13:32

it was so good. Now, it took twice as

13:34

long and we can say it took twice as

13:35

many tokens roughly. I would say it did

13:38

what the fable five initial straight

13:40

build should have done in the first

13:41

place. But in any case, {slash} goal, I

13:44

really find it to be tremendously useful

13:46

for the reasons of how goal works, what

13:49

it's really all about. And goal is just

13:51

basically saying check your work at the

13:53

end. If you're not finished, keep going.

13:56

That's making it a little bit simpler

13:58

than it really is, but that's basically

14:00

the concept is it creates a loop out of

14:02

self-evaluating whether or not it is has

14:05

reached the goal that it should stop at.

14:07

So, okay. I know you all can't hear

14:09

that, but

14:10

there is some music going on. So, I have

14:13

a feeling one of our video games is

14:16

done.

14:17

We got to go take a look at those. Okay.

14:18

[laughter]

14:19

So, here we are. This is what we were

14:21

building is a battleship kind of clone

14:24

game. And this is Claude Code's build of

14:27

it.

14:28

This is pretty awesome.

14:34

And get to my battle stations.

14:38

All right, and it's my turn. I have 41

14:40

seconds. Looks great. I'll attack there.

14:43

I missed.

14:45

Oh, yeah.

14:50

Oh.

14:53

Oh, come on.

14:55

Oh,

14:56

>> [laughter]

14:56

>> wait. Wait. Who the heck am I playing?

14:58

All right. How about here?

15:02

Yeah, you missed me, buddy.

15:05

Oh, yeah. You're dead now. You're dead.

15:08

Here we go.

15:12

Oh, and it just stopped. Not sure why it

15:14

just stopped. Might be reloading.

15:18

Hard to really tell, but I think it's

15:19

complete. So, let's close that.

15:24

It is still thinking. So, I did open it

15:26

in the middle of its still working.

15:30

That's damn good for still working.

15:32

Okay, let's take a look at Codex's.

15:35

So, this one would be the Sol build

15:38

first run. Let's take a look at what

15:40

first run looks like for the Sol build.

15:42

All right, I come in

15:44

two commanders. Oh, wow, have you done

15:45

person versus person already? That's

15:48

kind of awesome. I'm going to do

15:50

algorithmic. Yes, let's do algorithmic.

15:54

At position overlaps.

15:59

How do you rotate? R.

16:06

Well,

16:09

interesting. So, I can't seemingly move

16:11

the ships after we've put them in.

16:14

What if I rotate? Yeah.

16:17

Okay.

16:19

And a destroyer.

16:21

I can't even get back to the back row.

16:24

Okay.

16:25

So,

16:26

maybe not perfect yet.

16:28

No sound effects. The other one did a

16:29

great job with sound effects.

16:45

Very difficult to tell

16:48

that I'm hitting or missing.

16:51

So, I would say not bad. It's It's

16:54

definitely interesting. Okay, I'd say

16:56

it's fine.

17:23

>> [music]

17:26

>> Awesome, it's even testing the computer

17:30

Actually, in this case it's computer

17:33

against AI and even AI was taking too

17:37

long, so it's trying to evaluate why AI

17:39

is taking so long to make its turns.

17:42

I mean, really fantastic. Even the test

17:44

cases here are kind of fantastic to

17:46

watch.

17:47

All right, I want to talk to you about

17:48

the four

17:50

uh elbows that I've seen in AI and what

17:52

I mean by kind of AI 2.0. So, what we

17:55

saw about 2 and 1/2 years ago was being

17:58

able to code with these systems very

18:00

meaningful. Maybe maybe three three

18:03

years ago, uh being able to code

18:06

meaningfully. We had tab to complete

18:08

another things like that for a while

18:09

that you might be able to finish a line

18:10

or fill in the variable that is closest

18:13

to your line, those kinds of things.

18:15

Those were great, but it wasn't quite

18:16

the same as being able to go to chat GPT

18:18

and ask for a function and it gives you

18:20

gives you kind of the whole function

18:22

back and you drop it into your code and

18:23

then adjust it. It was absolutely not

18:26

finished, it was not perfect by any

18:28

stretch of the imagination, but it

18:30

meaningfully could move you forward and

18:31

scale you a little bit. That was kind of

18:33

really useful. And then about a year and

18:35

a half ago, a little bit more than that,

18:37

uh Codex was the really the first one

18:39

that cracked the harness to be able to

18:41

do kind of agentic file editing. So, it

18:44

could all of the sudden magically create

18:47

the file. When you said, "I need this

18:49

new function." or "I need this new

18:50

class." it would actually go to write a

18:52

whole new file for you. And I understand

18:54

today, that's not all that fascinating,

18:56

but back then it was mind-blowing to

18:59

watch it create files. Not to mention go

19:02

to files you had and edit the correct

19:04

lines in a file. That was magical. So,

19:07

that was the second elbow and it

19:08

tremendously changed the way that we

19:10

worked with it such that if you were

19:12

still copying and pasting code from chat

19:14

GPT or something like that, which a lot

19:16

of people still were for quite a while,

19:18

you were actually doing it wrong and you

19:19

weren't scaling like you should. All

19:22

right. Then we get to kind of the end of

19:24

last year, in the November time frame,

19:27

where we started understanding these

19:29

things were now agentically working,

19:30

creating files, things like this. But we

19:32

were pretty much still stuck giving it

19:34

small sets of requirements. I could give

19:36

it four, five, 10 requirements at the

19:38

most together. It would lose some of

19:40

those. That's part of my planning

19:42

benchmark. That's uh before that was

19:43

period where I started that. How much

19:45

planning, how many items are actually

19:47

making it through the planning phase.

19:48

That's why it existed.

19:50

So, we came up with policies and

19:52

practices of how you might put together

19:54

different kind of request buyers files.

19:56

This was the time we started with PRDs,

19:58

spec-driven development, those kinds of

20:00

things, to-do lists, all of this. Those

20:03

substantially helped. Those were very,

20:05

very meaningful that you could then put

20:07

in requests that were 20 or 30 kind of

20:10

items. And around November last year, we

20:13

started noticing you could put in 100

20:15

items. Now, my planning benchmark was

20:17

showing back then 60% of the items were

20:20

making it through. Um maybe 70% at that

20:23

point, but still that was substantial

20:25

compared to the days just previous, a

20:26

couple months previous to that where we

20:28

were doing two or three items or 10

20:30

items and usually getting it right. All

20:33

right. Then, fast track forward to now.

20:36

I think we're at another one of those

20:38

elbows. It's been a little bit longer,

20:40

maybe I'm not really sure what date we

20:42

describe to this, but I think these

20:44

models bring us to what I might describe

20:47

as AI 2.0. This is that moment that we

20:50

are no longer asking for kind of

20:52

task-level work. We can now ask for

20:56

objective-level work. And you've seen

20:58

slash goals, you've seen workflows,

21:00

you've seen fan out, you've seen a lot

21:01

of mechanisms these days inside of these

21:04

harnesses, directly available for you to

21:07

be able to ask for much bigger work,

21:09

much broader work than just the tasks

21:12

themselves.

21:13

Indeed, the thing we just did with

21:14

Battleship, I did not describe any of

21:17

the individual tasks. In there, I was

21:19

describing the basic outlay of a

21:21

Battleship game. I didn't even really

21:23

describe Battleship, of course, but that

21:25

would be a pretty known target. And I

21:27

was just saying, I want a 3D game that

21:29

fits this feeling and should be

21:32

highly animated and you can very clearly

21:35

see what is the opponent doing and

21:37

thinking. All of these kinds of things.

21:38

Really, it was much more intent-driven

21:41

and objective-driven than it was

21:43

feature-level driven. That is actually a

21:46

change. And I understand this is just a

21:48

single piece of software, so it's it

21:50

doesn't quite stretch it as far as we

21:51

actually can go right now. I go much

21:53

further than this in a lot of my

21:55

experimentation. My experimentation, I'm

21:57

several levels above that at this point.

22:00

It doesn't work all that well yet. It

22:01

does work, but you have to do a lot of

22:03

work to make it work. This would be the

22:05

copy and paste day. So, we're getting

22:07

there, but the models are making a

22:08

substantial change. So, I think what

22:10

we're actually seeing, better model,

22:12

sure, better better faster faster

22:14

cheaper cheaper. Maybe not cheaper

22:16

cheaper. Yes, that is something that

22:18

we're seeing right now, but in reality,

22:20

what I think we're seeing is this is a

22:22

moment that we will look back at very

22:24

much like we did in the Opus planning

22:27

phase with PRDs and the cursor days with

22:30

agentic writing and chat GPT being able

22:33

to write

22:34

kind of meaningful code that we would be

22:37

able to copy and paste. All of those are

22:38

very very meaningful moments. I think

22:41

we've just seen another one. It's raw.

22:43

It's the beginning, but I would say it's

22:45

your chance to start diving in and

22:47

understanding how that works. These

22:49

models are really capable of offering

22:51

that to us. All right, enough of this. I

22:53

hope you saw something that was

22:55

interesting, at least. My final takeaway

22:58

of all of this would probably be

23:00

Uh, saw Battleship, right? You got to

23:02

see it with your own eyes. At the same

23:04

time, Battleship built by Fable far

23:08

better far superior than it was using um

23:12

the GPT models, the new 5 6 GPT models.

23:15

GPT 5 6 was great. It worked very very

23:17

well.

23:18

But, to be fair, it was maybe a half of

23:21

the time or now a third of the time. So,

23:24

Fable for the win still. I have to say

23:26

Fable for the win.

23:28

But, admittedly,

23:29

three times longer is actually three

23:31

times longer. All right. I hope any of

23:34

this helped in any way. Good luck out

23:36

there moving up to kind of

23:37

objective-based requesting cuz 2.0 is

23:39

here. We're really moving forward into

23:42

the new world of AI, especially from a

23:43

development or kind of a a work

23:45

standpoint where we don't have to say,

23:48

"Go to my email client, search for these

23:50

five things, bring back the in" That

23:52

would be task level. You being able to

23:54

objectively say, "I want emails from my

23:57

boss." That's objective level. It will

24:00

figure out the details. So, kind of

24:01

start working on that space if you can.

24:03

You'll be surprised at how much better

24:05

these things have gotten in that space.

24:07

All right. Thanks for coming along for

24:08

the ride on this one, and I'll see you

24:09

in the next one.

Interactive Summary

This video explores the release of OpenAI's new GPT-5/6 series models—Luna, Terra, and Sol—analyzing their performance against existing competitors like Fable. The presenter conducts a benchmarking experiment to evaluate planning and intent recovery, and demonstrates the effectiveness of the '/goal' prompt engineering technique for complex software development tasks. The video concludes that these advancements mark a shift towards 'AI 2.0,' where interaction moves from task-level prompting to objective-based requesting.

Suggested questions

3 ready-made prompts