HomeVideos

I Made Codex and Claude Code Build the Same App. One Clearly Won.

Now Playing

I Made Codex and Claude Code Build the Same App. One Clearly Won.

Transcript

696 segments

0:00

So, I had Claude Code and Codex build me

0:01

the exact same app. I gave them the

0:03

exact same prompt, but the results are

0:04

extremely different. Not only the

0:06

outputs, but the way that they actually

0:08

got to the output, very different. One

0:10

of them took 3 days, one of them took 5

0:11

hours, one of them spent $3,000, one of

0:13

them spent $800. So, today I'm going to

0:15

break down the two outputs, why they're

0:17

so different, and it's very clear to me

0:18

now what Codex is better at and what

0:20

Claude Code is better at. So, by the end

0:22

of this video, hopefully you have some

0:23

clarity. Let's not waste any time and

0:24

get straight into the video. All right,

0:25

so real quick, before we start going

0:26

over the outputs, let me just show you

0:28

guys the actual prompt. I utilized a

0:30

{slash} goal, and I did the {slash} goal

0:31

inside both Codex and Claude Code, gave

0:33

them the exact same prompt. So, here it

0:35

is. I'm not going to read the whole

0:36

thing, but here it is. I basically said

0:38

build a production-ready, originally

0:40

branded Typeform alternative. So, we're

0:42

kind of trying to clone Typeform here. I

0:44

said orchestrate specialized agents

0:45

throughout three phases: the research

0:47

phase, the build phase, the verify

0:49

phase. Now, looking back, I'm not saying

0:51

this is the most optimal prompt. I

0:53

honestly, if I was to redo this, I would

0:56

probably add another phase between

0:57

research and build that would be of all

0:59

about planning and mapping out the whole

1:01

flow. I honestly think that that would

1:02

result in a much better output from both

1:04

of these systems, but we'll take a look

1:06

at them in a sec. Anyways, I'm not going

1:08

to read this whole thing out. You can

1:09

screenshot it, you can reuse it as you

1:10

want. I do want to call out what I put

1:12

at the end though here, which was do not

1:13

stop at a prototype or first successful

1:16

build, continue researching, building,

1:17

testing, breaking, fixing, and retesting

1:19

until the app is genuinely complete. So,

1:21

I basically just wanted something that

1:23

was like production-ready, maybe we

1:24

could go to market the next day. Cool.

1:26

So, like I said, I gave both coding

1:28

agents, Codex and Claude Code, the exact

1:31

same prompt. Okay, so let's take a look.

1:33

I have not looked at either of these

1:34

yet. This first one I had to do on my

1:35

Mac because I was actually out of town,

1:37

I was traveling when I kicked off this

1:39

first one. And then I had the idea to

1:40

like, oh, you know, I should have the

1:41

other one do the same thing. So,

1:43

anyways, this is the first version that

1:45

we have. I haven't actually tested any

1:46

of this yet, I haven't clicked through,

1:47

so we're getting the real raw reaction.

1:49

My first reaction here is that this is

1:51

called Realform, and it looks pretty

1:52

solid. Like, honestly, I, because I've

1:55

built a lot of sites with AI, I can tell

1:56

this was AI generated. You know, we've

1:58

got the hero image over here, which is

2:00

still pretty a pretty cool UI element.

2:02

We've got the text over here. We've got

2:04

this little pill. It just feels to me

2:05

like AI, but the background, the depth,

2:08

it's not too bad, right? Like an average

2:10

person probably wouldn't look at this

2:11

and just assume, oh, AI vibe coded site.

2:14

So, we've got experience, which will

2:15

take us down a little bit. We've got

2:17

reliability, which takes us down again.

2:19

We have pricing, which goes down to the

2:21

bottom. We can build a private draft,

2:23

which will take us to a sign-up page.

2:25

I'm going to blur this out right here

2:26

because it says it basically gives away

2:27

which agent built this version. And we

2:30

can also click right here to start

2:32

building. So, let me just sign up real

2:34

quick for an account, and then I'll show

2:36

you what the experience inside looks

2:37

like. And by the way, this whole sign-up

2:38

page, not too bad at all. I like this

2:41

vibe. Okay, so we are in a demo

2:42

workspace because it didn't have any

2:44

keys to make like real authentication

2:45

and real whatever. So, this is a demo

2:47

workspace. Let's take a look at how this

2:50

thing runs. So, we can make a product

2:52

launch survey. We have the ability to

2:55

structure how this works. This is the

2:57

welcome screen of our form. So, let's

2:59

plan a launch people remember. We can

3:00

play with this text right here. We can

3:02

put variables, so we can put a score or

3:05

a price. Interesting. Okay, so this is

3:07

how we like insert different elements,

3:08

and we can

3:09

play with the sizings on them. Okay,

3:11

very interesting. I don't know how

3:14

I don't know how we remove these. So,

3:15

like if I wanted to remove these

3:16

elements right here, these two boxes,

3:19

I'm not exactly sure how I do that. Do I

3:20

click delete there?

3:22

Delete. Oh, this will delete the entire

3:23

thing. No, I don't want to do that. We

3:25

can add text right here. Okay. So, I

3:27

think this is a little bit, honestly,

3:29

like looking at this, it's a little bit

3:30

overwhelming. Like I'm kind of confused.

3:32

There's a lot going on in this UI. I can

3:35

change the label of the button.

3:37

Cool. We can upload an image for the

3:39

welcome screen. Okay, so I put in an

3:41

image, but it actually says the image

3:43

preview is unavailable. So, this just

3:45

goes to show that even though there was

3:47

testing, there's so many different

3:49

scenarios that agents might not think to

3:51

actually test, especially if you're not

3:53

the one driving those tests. So, that's

3:55

kind of like a bug that we already

3:56

found, right? You can put alternative

3:57

text, you can change the focus and the

3:59

placement of everything. Okay, so that's

4:00

the home screen. Let's see what's on

4:02

page one. So, this first thing is what

4:04

would you most like to improve? And once

4:07

again, this is just too much going on. I

4:09

mean, clearly this is a multiple choice

4:10

type of answer. As you can see, we could

4:12

have made this also short text. And

4:15

yeah, let's change that to short text.

4:16

There's still all of these variables up

4:17

here, which I don't love. I mean, this

4:19

is supposed to be the title. I just

4:20

think that there's too much going on,

4:21

right? Like, this is a description.

4:23

Please answer in one sentence max.

4:25

People can put in their answer right

4:26

there. But, we've got short text, long

4:28

text, email, phone number, website,

4:29

contact info, opinion scale. Okay, let's

4:32

see what that one looks like. It's

4:33

asking me to confirm, and I don't know

4:34

if you guys realized, but like

4:36

when this pops up to confirm, it kind of

4:39

like pops up over here on this left

4:40

side, which also kind of feels like a UI

4:42

bug.

4:43

That's not perfect. We can do a matrix,

4:45

we can do a file upload. So, there's a

4:47

lot of things that it is letting us do.

4:49

So, I think that this agent did really

4:51

good about thinking about kind of like

4:54

from the admin side, the data you want,

4:56

and maybe the way you want the data to

4:57

be collected. But, from a UI

4:59

perspective, like from the user

5:00

perspective making forms, this is

5:02

confusing. Like, I think this would

5:04

probably have me churning out really

5:05

quickly because of how simple other

5:07

things truly are, like Typeform. We can

5:09

have multiple pages. Okay, so I think

5:10

we're starting to get the gist here. We

5:12

can set up a workflow, right? So, we can

5:14

make kind of like logic-based conditions

5:16

within the form, which obviously is very

5:18

important. We can set up a theme. So,

5:20

corner radius, we can do different

5:21

colors, which

5:23

honestly though, if I look at this, like

5:25

this isn't letting me click.

5:27

So.

5:30

That's really interesting, right? Like,

5:31

this doesn't really seem to be even

5:33

changing anything. So, that's not good

5:34

at all. Accessibility, we can see if

5:37

there's anything blocking. Publishing,

5:40

we can connect this to a webhook, email

5:42

notification, partial response. We can

5:44

share the actual link to the form. You

5:46

know, I think because it assumed that

5:48

this is basically just a demo, it's

5:49

showing me the UI, but it didn't

5:51

actually maybe build out like the full

5:53

functionality of this. So, I don't love

5:55

that, right? Now, I do have to say that

5:57

I do think that there's promise here,

5:58

but because it built that out purely

6:01

with the idea that it's a demo, even

6:02

though I told it, "Hey, I'm looking for

6:03

this to be like a a real genuinely

6:06

complete product." That's not great.

6:08

Okay, so now we're hopping over to my PC

6:10

where we're going to test out this other

6:11

version. So, this version's called

6:12

Formora. Immediate reaction is that this

6:16

was not really being thought about from

6:17

like a design perspective. I don't even

6:19

know what this is supposed to be. I

6:19

mean, this looks like it's supposed to

6:20

be a landing page, but this thing is

6:21

just brutal. It's brutally ugly, right?

6:24

Anyways, build beautiful one question at

6:26

a time forms, share a link, and watch

6:27

the answers roll in. Self-hosted,

6:29

private, and fast. So, let's go ahead

6:30

and get started for free. I'll go ahead

6:32

and create an account. My name will be

6:33

Bob. Oh, oh, no. My email will be bob@

6:37

test.com, and my password will be 1 2 3

6:39

4 5 6 7 8. And that will be it cuz it's

6:42

at least eight characters. Sign up.

6:44

Okay.

6:45

Cool. So, this

6:47

Yes, I agree. This looks vibe coded.

6:50

But, at least as a user, I know exactly

6:52

what to do. I'm not staring at something

6:54

and immediately overwhelmed. I've got my

6:56

account down here,

6:57

which

6:59

Okay, that is another UI bug, right?

7:01

Like it's not letting me I don't know

7:03

what this button is for. Okay, it opens

7:05

settings, but I'm not able to I don't

7:08

see log out.

7:10

I assume that log out would be down

7:11

here, but we can't actually access it

7:13

even if I change the sizing. So, that's

7:16

one bug that didn't get caught,

7:17

unfortunately. We have different

7:18

workspaces. So, right now I'm in a

7:20

workspace up here, but I can create

7:22

another one. So, let's call this one

7:24

business. We can create another

7:25

workspace, and we can switch between

7:27

those. Okay, that's pretty cool. That's

7:28

pretty slick. Let's go ahead and create

7:31

a new form. Okay, this is very nice.

7:33

This is way less intimidating. I

7:35

actually feel like I know what's going

7:36

on. This looks way more like a Typeform.

7:38

Let's see. Default display mode is

7:39

either conversational or stacked. So,

7:42

I'm assuming we can either have it be

7:44

one at a time or we can have all

7:45

questions showing in one form. So, we

7:48

can also have a progress bar showing up.

7:50

We can have question number showing up

7:51

and this is responding. We can have

7:53

keyboard hints. Um we can have auto save

7:57

and we can have capture partial

7:58

responses. Okay, cool. And over here

8:00

we're able to just customize the stuff

8:01

easily. So, page one, first question

8:03

goes here. What is your

8:06

mood? Okay, cool. That pops up.

8:08

Description.

8:10

Happy. Okay. This is called question

8:12

one. We can make it required. We can

8:14

have max characters, pattern, blah blah

8:15

blah. Um where do we control the field

8:17

type? Is Oh, right up here. Short text.

8:20

Okay, this is showing is Oh, I think

8:21

it's because it's like a default. So, if

8:23

we add another one, there we go. So,

8:25

when we add a new type of content,

8:27

that's where we control what it is.

8:29

Email, phone number, website, drop-down,

8:31

picture choice, net promoter score,

8:33

opinion scale, rating, ranking. Okay.

8:36

So, these are pretty cool. Let's see if

8:37

we go ahead and do like an opinion scale

8:39

what that looks like. Um

8:42

Okay, so it made it but it kept the same

8:45

questions already. So, Oh, okay, no it

8:47

didn't. I just had to switch over to it.

8:49

Okay, so this is saying how much do you

8:50

agree?

8:51

Now, one thing I just noticed though is

8:53

this is keeping the number here as one.

8:56

So, like this should obviously be two.

8:59

Um how much do you agree though? That

9:00

looks good. Opinion scale. Okay, let's

9:02

try to add something else. Let's add a

9:03

website. Um right here. This is a bit of

9:06

a bug. I've clearly selected this one,

9:07

which is what is your website, but it's

9:09

not showing that. I have to click off of

9:10

it and then I have to click back on it.

9:11

So, that's another bug. Man, like you

9:13

really cannot underestimate how much

9:16

testing has to go into finding bugs.

9:19

There are so many different bugs and

9:20

it's not as simple as just telling Cloud

9:22

Code, "Hey, {slash} goal, build me an

9:23

app." But anyways, this is giving a

9:24

keyboard hint. That's pretty cool. Once

9:26

again, I'm a little confused why it's

9:28

not fixing these page numbers or

9:30

question numbers. It's showing that all

9:31

of them are just question one. But what

9:33

I've noticed is that like up here it

9:34

says one of three.

9:36

And then if I go here it says one of

9:37

two, and up here it says one of one. So,

9:41

there's some other logic bugs going on.

9:43

But anyways, this is honestly a much

9:45

better experience. Let me go ahead and

9:47

do Let's see what else we got. We got

9:49

logic. So, once again, if they answer a

9:51

certain thing, we can route them to a

9:53

different place. We can add these

9:54

branches. We have design, so we can

9:56

change the way the form looks. Okay,

9:58

that's pretty cool. You can also upload

9:59

your own themes, which is pretty cool.

10:01

Let's just say we like this dark theme

10:02

for now. Now, one thing I don't like

10:03

about this design is once you click

10:05

click design,

10:06

you can't easily navigate back. You'd

10:07

have to like,

10:09

you know, hit the back arrow. But once

10:10

you click design, it puts you in this

10:12

new UI where you There's not like a

10:13

button to just navigate back. So, that's

10:15

that. We can share it. We can look at

10:17

our results. Okay, same thing. I guess

10:19

we can click up here to navigate. We can

10:21

go to webhooks.

10:23

Same thing. It has this issue where once

10:26

you navigate into one of these versions

10:27

or one of these buttons, you can't nav

10:29

back very easily. So, that's a little

10:30

bug we need to fix. We can change the

10:32

theme from up here, which is pretty

10:33

cool. Anyways, let's go ahead and

10:34

publish this. So, when we publish this,

10:37

it gives us a link. And let me just call

10:39

this form real quick test one.

10:42

So, when I publish this,

10:43

um I'm able to then open up a new page,

10:47

and here's the form. Okay, this looks

10:48

like Typeform, right? So, I am

10:50

happy-ish.

10:52

I agree three. My website is

10:56

um

10:56

google.com,

10:57

and that is our response. Let's go ahead

11:00

and see. If we go to results, oh, nice.

11:03

We got an actual result. I guess we're

11:05

just seeing Oh, because this is just the

11:07

insights of drops, right? This is If I

11:08

go to summary, now I can see the actual

11:10

answers. If I go to responses, I can see

11:12

each individual submission and where it

11:14

came from as well, which is pretty

11:15

interesting. But that's not bad. And if

11:17

I go back to my dashboard, we can see

11:18

our forms. Um I think that I I thought I

11:21

titled this one. I guess it didn't save.

11:23

So, I'll just call it one. Oh, I have to

11:25

publish it. Okay, republish it.

11:27

Dashboard, it's called one. If I go to a

11:29

different workspace that doesn't exist.

11:31

Okay, cool. So, like, this obviously is

11:33

far far better than the first one. The

11:35

first one from a design perspective I

11:36

think was better, but as far as

11:38

functionality and doing what I wanted it

11:40

to do, this one definitely takes the

11:42

cake. So, now I'm curious, what do you

11:44

guys think? Which one do you think was

11:46

made by which AI?

11:49

Well, let's just start taking a look at

11:51

the results here. Okay. So, we've got

11:53

Claude Code, we've got Codex. There's a

11:54

few things that we're going to go over.

11:56

First of all, let's do the reveal. Which

11:58

one did Claude Code make?

12:00

Claude Code made Formora and Codex made

12:03

Realform. So, Claude Code made this one.

12:05

Claude Code made this one, which

12:07

from a design perspective wasn't as

12:08

good, but from a, you know, design of

12:12

the actual like

12:13

method, the functionality, the actual

12:15

output, definitely much better. And I

12:18

will be honest with you guys, I wasn't

12:19

expecting that. I was expecting Codex to

12:22

have built a better one just based on

12:24

the way that I've been using Codex and

12:26

Claude Code and how much I've been using

12:28

Codex lately, I was honestly expecting

12:30

Codex to win this challenge, but so far,

12:32

when we're just looking at the actual

12:33

like output so far,

12:35

I think Claude Code is taking the cake.

12:37

Obviously, that's not to say that Claude

12:38

Code's just better right out because it

12:40

has to do a lot with this prompt, and

12:42

I'm going to explain what I mean by that

12:43

as we keep digging in here. But anyways,

12:45

let's take a look at the cost. Claude

12:46

Code costed me, if I was using API

12:48

billing, 832 bucks, which means that

12:50

Codex costed almost $3,000, which is

12:53

just insane. When it comes to the output

12:55

tokens, Claude Code used a little over 2

12:57

million while Codex used almost 11 and a

13:00

million output tokens. Now, when it

13:02

comes to Codex, it used GPT 5.6 Soul.

13:05

That's what I was using to drive. I was

13:06

using it on high, and it orchestrated

13:08

all of its sub-agents and everything

13:10

like that to use GPT 5.6 Soul. So, all

13:13

of this is 5.6 Soul.

13:15

With Claude, it did this breakdown where

13:17

it had Fable 5 going, it had Opus 4.8

13:18

going, and it had Opus 5 going. Now,

13:20

look at this. You see here that Opus 4.8

13:23

was the main orchestrator, which really

13:25

threw me off because I didn't start the

13:26

session on Opus 4.8. I started the

13:29

session on Fable 5 and I gave that slash

13:31

goal. I think that it for some reason

13:33

had some sort of security safeguard

13:34

check and it reverted back to Opus 4.8

13:37

and then Opus 4.8 became the main

13:38

orchestrator and it started spinning up

13:40

Fable 5 agents to do the work, which I

13:42

thought was really really interesting,

13:44

but it was still able to be pretty

13:45

efficient. I mean, 832 bucks with 2

13:48

million output tokens compared to what

13:50

Codex did over here with 11 million

13:52

output tokens, man, that's really

13:54

interesting. Time-wise, Claude Code took

13:56

5 and 1/2 hours while Codex took 61

13:59

hours, almost 62 hours. So, that's like

14:00

2 and 1/2 days. That I was genuinely

14:03

shocked when Claude Code was like, "Hey,

14:04

I'm done." and I had been running Codex

14:06

for almost

14:07

a day and a half already, you know, cuz

14:09

I was traveling. I kicked it off on

14:10

Codex. I was like, "Oh, when I get home,

14:11

I'm going to send the same prompt off to

14:13

Claude Code and just see like, you know,

14:14

how they compare." I was really shocked

14:16

to see how fast Claude Code finished

14:17

here. It's interesting because in

14:19

previous testing, I've always felt like

14:20

Codex was a bit better with being

14:22

efficient and being quick, but obviously

14:25

this very different things and I think

14:27

it has to do with the way I prompted

14:28

once again. So, we'll dig into that in a

14:30

sec. What did the shape look like? All

14:31

right, so now how about the shape? Well,

14:33

Claude Code did this with one

14:35

orchestrator, 35 sub agents and about

14:37

2,800 tool calls, whereas Codex did this

14:39

with one orchestrator, 126 sub agents

14:42

and 32.5k tool calls. So, I we're kind

14:45

of understanding the gist of what

14:46

happened here. Codex worked longer, it

14:48

used more agents, it used more tokens,

14:50

it used more tools and Claude Code did

14:53

less and honestly, I think it would did

14:54

better so far. Now, what's really

14:56

interesting is the tests. We clearly saw

14:58

bugs in both of them, which was a little

15:00

bit disappointing,

15:01

but

15:03

this tells us a lot about the way that

15:04

these models work.

15:06

So, Claude Code did 296 unit tests, 199

15:09

test cases and 102 browser tests,

15:11

whereas Codex did 2,300 unit tests, 341

15:14

test cases and 391 browser tests. So,

15:18

I've always felt like

15:20

Claude Code or I guess Fable is kind of

15:23

the wise owl. I like to use it for being

15:24

creative, for planning, for

15:25

brainstorming, for helping me figure out

15:27

the path. Whereas Codex has never felt

15:30

that good at that for me. For Codex, it

15:33

feels like the

15:35

just it's going to be obedient, it's

15:36

going to do what you say, and it's going

15:37

to do it well. It's going to run tests,

15:38

and it's going to make sure that the job

15:40

has been done. Which means for me, when

15:41

I prompt Claude code, I like to give it

15:43

a prompt like this. I give it a

15:45

high-level goal. I say, "Hey, this is

15:47

what I want. This is what good looks

15:48

like. This is when you stop." And with

15:50

Codex, it almost feels like you need to

15:51

be a little bit more specific. You need

15:53

to be more like, "Here's kind of like

15:54

step one, step two, step three, step

15:56

four." I just made a video about getting

15:57

out of the model's way, and that video

15:59

was based on a talk that Boris Cherny

16:01

had done. And obviously Boris Cherny is

16:03

the creator of Claude code. But clearly

16:05

in this experiment with Codex, it just

16:07

didn't seem to interpret what I meant

16:09

well enough, and it didn't seem to be

16:11

creative enough to explore enough to

16:13

figure out what sort of experience I was

16:15

looking for at the end of the slash goal

16:17

prompt, even though it worked so hard

16:19

and so long. I feel like this was a

16:21

waste of money here. So I thought that

16:22

was really interesting. And then what I

16:23

did is I basically inspected the entire

16:25

sessions, I inspected everything they

16:26

did, and I consolidated all that, and

16:28

then I had Codex look at both of those

16:31

and tell me which agent did better. And

16:34

agent A is Claude code in this case, cuz

16:35

I I kept this anonymous. Codex said that

16:38

Claude code did better, which was really

16:40

interesting. Now, before I dig into

16:41

this, guys, I'm not trying to bash on

16:43

either one of these tools. I use them

16:46

both on the daily. I will be honest with

16:48

you guys, lately for knowledge work,

16:50

I've been using Codex to drive my

16:51

sessions. I've been using Codex probably

16:52

80% of the time and Claude code probably

16:54

20% of the time. But I think it's really

16:56

important for you guys to realize that I

16:57

still like them both because I use them

16:59

both for different scenarios. And I like

17:01

to, as the models improve, as new

17:03

updates come out, I switch around a lot.

17:05

I'm not just going to choose one and

17:06

say, "Hey, this is my driver for the

17:07

rest of my life." You know what I mean?

17:08

So anyways, let's take a look at these

17:10

categories that Codex said that Claude

17:12

code won in. So product judgment and

17:14

scope. Agent A, Claude code won with

17:17

judgment and scope, and that aligns with

17:18

the way that I feel about it. Cloud Code

17:20

made clear must slash defer decisions

17:22

and focused on valuable differentiators.

17:24

While Codex pursued 135 capabilities

17:27

including several expensive operational

17:30

features with less restraint and I feel

17:32

like that's exactly what we saw. Codex's

17:33

version was overwhelming, wasn't

17:35

thinking about the user, wasn't thinking

17:36

about the experience. Codex was just

17:38

building just to build and and it just

17:40

was too much. Then, the next one,

17:41

architecture and execution. This one

17:44

went to Codex. Cloud Code's contract

17:46

first waves produced zero merge

17:48

conflicts, but Codex built the more

17:50

operationally mature system with

17:51

immutable revisions, offline recovery,

17:54

migration safety, concurrency handling

17:56

and cloud boundaries. So, maybe the back

17:58

end at scale, the infrastructure Codex

18:00

was building was much better and that's

18:01

why I love to do a lot of development

18:03

and planning with Cloud Code and I love

18:05

to do security reviews, bug finds, bug

18:07

fixes, all those types of things with

18:09

Codex. You guys have probably seen the

18:11

Codex plugin for Cloud Code where you

18:12

run the adversarial review and it's

18:14

really really helpful and it almost

18:16

always finds things that my Cloud Code

18:18

workflow missed out on. Bugs, edge

18:20

cases, things like that. Now, the next

18:22

category, testing and reliability, Codex

18:23

won by a wide margin. Cloud Code

18:26

performed strong security and data

18:27

correctness testing, but Codex added

18:29

cross-browser testing, property tests,

18:31

fault injections, blah blah blah. Codex

18:33

was able to do things like it tested on

18:35

so many different types of browsers, it

18:36

tested on mobile and we did not see that

18:38

happening with Cloud Code. As you can

18:40

see here with the tests, there are just

18:42

obviously significantly more tests that

18:43

were being run by Codex. And yes, it

18:46

worked for a lot longer and spent more

18:47

money, but still the model harness in

18:50

this case seemed to just do better with

18:52

the testing. Now, I guess some of you

18:53

could argue like, okay, well, maybe

18:55

Cloud Code needed less tests because it

18:56

did a better job building in the first

18:58

place and that is also a valid argument,

19:00

but I'm just trying to show you what I

19:01

found here. And the last one here was

19:02

basically around efficiency. Cloud Code

19:04

got a 9.0 out of 10 while Codex got a

19:06

5.5 out of 10. Cloud Code finished in 5

19:09

and 1/2 hours for roughly $447 using 35

19:11

agents. See, I know that I said earlier

19:14

832, there was a mismatch somewhere. The

19:16

point being

19:17

Cloud Code spent a lot less. I inspected

19:19

the session logs and I was mainly for

19:21

the most part getting this answer. So, I

19:23

think that this must have been a little

19:24

bit of a hallucination somewhere along

19:26

the way. I'm not sure where, but

19:28

according to the slash usage stats of

19:29

the session and all the sub agents,

19:32

the number was more like 800. But

19:33

anyways, Codex took 2 and 1/2 days and

19:36

spent way more money, way more sub

19:38

agents. So,

19:39

Cloud Code did this about 11 times

19:41

faster and 6.6 times cheaper. And that

19:44

once again kind of goes against what I

19:46

thought was going to happen because a

19:47

lot of my tests in the past when I've

19:49

done Cloud Code versus Codex, Codex has

19:51

been more efficient with tokens and has

19:53

also been quicker. But, you know, 5.6

19:56

Soul is new. You got Fable over here.

19:57

You got Opus 4.8. You just There's so

19:59

many different variables and it's always

20:01

like you're pulling a lever to slot

20:02

machine. You just don't know what you're

20:03

going to get with these models. But

20:05

anyways, I hope that this experiment was

20:06

insightful to you guys. I hope that this

20:08

at least made you think about the way

20:10

that you think about these two model

20:11

harnesses and the way that you think

20:12

about prompting these things cuz it's

20:15

always changing and that's why it's so

20:16

important to kind of be hands-on doing

20:19

little experiments like this because you

20:21

never know how it's going to fit into

20:22

your workflow. I think that it's great

20:23

to be following along with Boris Trchony

20:25

and Andrej Karpathy and all of these

20:26

thought leaders in the space and people

20:28

that are actually developing these

20:29

tools, but something I said in my

20:30

previous video was like you should not

20:32

be just taking their advice and blindly

20:34

applying it because they do different

20:35

things with it. They have different

20:36

motivations. It would be like if you're

20:38

a professional triple jumper and you're

20:40

taking advice on your jumping from a

20:42

high jumper. Like maybe there are some

20:43

similarities there and maybe the

20:44

fundamentals and you know, some of the

20:46

foundational things are consistent, but

20:48

like at the end of the day it's a

20:49

completely different sport. It's just a

20:50

completely different ball game and you

20:51

probably want to be taking advice from

20:53

people that are also triple jumpers.

20:55

That's a real sport, right? Okay, yeah.

20:56

So, I knew that this was obviously like

20:58

in track and field, but I just wanted to

20:59

make sure that it triple jumpers sounds

21:01

like a weird

21:02

term. But anyways, guys, that is going

21:04

to do it for today. So, if you enjoyed

21:05

the video you learned something new,

21:06

please give it a like. It helps me out a

21:08

ton. And as always, I appreciate you

21:09

guys making it to the end of the video,

21:10

and I will see you all in the next one.

21:12

Thanks, everyone.

Interactive Summary

The video compares two AI coding agents, Claude Code and Codex, by tasking both with building a Typeform alternative using an identical prompt. While Codex invested significantly more time, money, and testing, Claude Code produced a more usable product, though both agents exhibited UI bugs and limitations. The experiment highlights the differences in behavior between the two models: Claude Code demonstrated better product judgment, while Codex focused heavily on architectural depth and exhaustive testing. The author concludes that model behavior varies based on the specific prompt and use case, suggesting that users should perform their own experiments rather than relying on one specific tool.

Suggested questions

3 ready-made prompts