GPT-5.5 vs Opus 4.8: The Real Winner Was /goal Mode
691 segments
Okay, today I'm excited about this one.
I haven't done one of these in a while.
This used to be all the rage about 6-8
months ago, but it has kind of fallen
off recently and I want to bring it
back. I think it's so much fun. It's
worth taking a look at. What I'm going
to do today is I have one kind of big
requirement. You'll see it in a second.
It's going to be a really cool fun
application to try to get agents to
build and I'm going to pit two agents
against one another. We're going to see
can GPT-55
or Opus 4-8 build this best. That's
really today's challenge. Which one
builds it best? And oh, by the way,
along the way I'll show you how I
designed the product and a couple other
things and hint at what I'm going to do
with this kind of series moving forward.
It's going to be kind of fun, I believe.
Okay, first before we go any farther,
I'm going to show you what we're going
to build. I'm going to show it to you
here for just a second so you can kind
of get the concept, but this is going to
be a native application on my Mac. So,
this is going to be a real normal
launcher application, not this web kind
of expression that you're seeing. What
it's going to be is an application that
you can line different
different kind of apps around a
a radial diagram here or kind of a
radial dial that you can select from.
You can use the arrow keys to select
different items or to go between the
different pages that you might have. And
then when you go into these things, you
should be able to edit them and put
different things in different locations,
different apps, that kind of stuff. So,
it's a launcher, but it's a lot more
than that because it also has a
mechanism to be able to open files with
certain applications. If you use this in
finder, you should be able to open those
files with a particular application. You
should be able to run websites. You
should be able to do just about anything
with this launcher. It's really a
complex request set with this kind of
expression, this diagram that you see
here as part of the request that I'm
going to be able to give to these very
advanced models right now to see how
they do to build it. That's really the
whole point. This is going to be pretty
hard to build especially in Swift, so I
want to see how it does. But moreover, I
want to compete models against one
another and now I want to get that
kicked off. Let's just get that started
now so that we can see that kind of
running and then I'll tell you how I got
this designed. Really easy and really
cool. Okay, this is going to be easy.
What I'm going to do is I'm going to use
the Codex desktop application which is
Codex CLI essentially. I'll get that one
kicked off. I'm also going to use the
Claude Code desktop application and get
that one kicked off. So, both of those
are off and building. It's of course
worth mentioning we're on Opus 48 extra
in Claude Code and we're on GPT 55 extra
high in Codex. But for fun, I am also
going to add some new ones here which
are goal items. So, I've built two new
work trees, same kind of thing for
Codex. I will do the same work here
using {slash} goal and then I will tell
it to
run a {slash} goal. You'll see it sends
in as a goal. I'll do the same thing
over here for Claude Code. All right, we
go. And the goal is, no pun intended,
those will actually both run using the
{slash} goal feature which is kind of a
more open-ended building process. So,
interesting to see with such a complete
version of definition around a product
that we want built is goal any better
than using just a direct representation
of a single agent that that master agent
then can spread out as much as it needs
to. I don't know. Let's find out. So,
one of the things I want to show you
just as kind of lagniappe here, a little
bit extra, is how I designed this
application because I've actually asked
for this system to be built before and
it kind of comes out pretty trash. And
so, I had to come up with a creative way
to allow a system to imagine a radial
diagram or radial system that would be
attractive and elegant and have some
animations to it which was not
necessarily an obvious thing for the
systems to come up with. Let me show you
the steps that I took. All right, so
first things first, I was using Claude
design. Claude design, great product if
you have Anthropic as a subscription,
you have access to Claude design. And so
my first request was kind of a big
request here, basically saying, "What we
want, we want radial rings, maybe three
different layers of rings with different
apps on them." So a lot of this was
because of the original request and how
spread out it was. It's okay that it's
this spread out, but I asked for six
variations on the theme, so we might be
able to choose from them. And this is
kind of what it gave me to begin with.
So you can see variations on the same
theme here, but very, very similar
between all of the different variations.
Then it came back, it created this, and
then I pushed back in a little bit and
said, "Well, keep this design." So this
document that we're looking at here, I
wanted to keep all of the ones that it
was already showing me and make sure
that it was going to essentially
version. And I said, "Do this again with
this extra information on the same
page." And so what it did is it put more
items down here. I wanted to kind of
create sections or something that might
create the modes that you see in the the
final design that we have. So okay, gave
me six more, all fine is what I would
say about these, not great. I wasn't
excited about them. I had mentioned I
wanted kind of a AAA gaming expression
so that it had an opportunity to be
highly graphical, and this feels a lot
more um mechanical than I would have
liked. So what did I do next? Okay, so
the next thing I did is I came out to
Chat GPT and thought, "Okay, let me use
the image model." The image model is
highly creative. Let me start there and
see if I can get ideation from here to
kind of go broad first. And so I gave it
pretty much the same original kind of uh
request that I said, but maybe a few
other things. And what it came up with
was relatively interesting. It certainly
wasn't uh extremely attractive, but it
these were very interesting graphically
compared to what we were just looking
at. And I thought, okay, I really like
the simplicity of one layer of icons and
one layer of kind of the modes that you
might be on and a button in the middle
you might be able to press to kind of
switch the modes. Really neat, very
simple, allows for a lot more icons in a
more creative pattern. So, quite
frankly, I took a screenshot just of
this bottom one. I grabbed my screenshot
utility and took a screenshot right here
of just this bottom one and passed it
on. Let's go back to Claude design.
Okay, here in Claude design, what I was
looking at for this one, I gave it the
the full screenshot here and later came
back and just gave it the smaller
screenshot and said, "Here's what I want
you to push in on. Think bigger and
better than you've been thinking
before." And I told it also to version
this file. So, I wanted to keep this
file and I wanted it to kind of create a
new version. And so, what it came up
with was more like this. You can it's
getting closer to what we were looking
at
from the ChatGPT display, better, but
still not quite there. So, then I took a
screenshot of this one and said, "Okay,
this is a very good start. I'm going to
start a brand new context." So, I
started a brand new conversation up here
and went to a brand new context, gave it
just that one image and said, "Okay,
from this one image, let's start pushing
in with a much simpler kind of
design message that you'll see here."
So, with that it started creating this
kind of version, which I thought was
fantastic. This is a great first step.
You can see it already has animation
around the outside ring. It was very
word heavy, these kinds of things. So,
this was a great start. I went back and
forth with it a few more times and got
to a place to where what we were looking
at earlier is the design that we have.
And I told it that I want arrow keys,
these kinds of things. So, there was
some iteration that I did here, working
as I was discovering and exploring more
things, working on the assets it was
giving me. I would push in and say,
"Well, let's duplicate instead of having
words in the middle, let's duplicate the
icon that maybe they're selected on at
any given moment. Let's put that into
the middle
down here so that it kind of feels like
as you move between nodes, you're
reinforcing the image. And let's use the
arrow keys to allow somebody to swim
around the outside and be able to hit
enter." All of that kind of stuff. So,
that's where they stopped. I then
packaged this up. There are ways to
share things from Claude Design. And in
here, you can use a send-off message.
You can just copy this prompt and give
it to your agent. I have found that to
be hit or miss, unfortunately. I'm sure
they're still working on this. I've had
the agent come back multiple times and
say, "Well, because of OAuth or
something else, I can't quite get to
that. Can you help me?" But, to solve
that, you can click this box to download
it, and then this will be the prompt
that you'll want to give your agent when
you give it the downloaded zip. You just
drop the zip on it, give it that, and it
has all the resources that it needs. So,
that really is kind of the the
methodology that I took taking this
this kind of ring mechanism and putting
it into Claude code or somewhere else to
be able to design from. So, that's where
we start. That's the resources that are
inside of this original build
environment. But, I want to show you one
other thing really briefly. This one
shocked me. Okay. So, I'm sharing a new
product here. Not It's new to me. It's
not new to many of you called Open
Design. A team has gone out and kind of
looked at what Claude Design was doing
and decided to put together almost
exactly the same product in an open way.
You can come in here and put in your own
subscriptions or your own
API keys, anything that you need. It
really is very, very open. And what I
ended up getting into was a version of
the same build. I will share this at a
later time cuz I want to do a comparison
of open design to to Claude design. I
just wanted to show this as quickly as I
could because if you're looking for
design tools and maybe you don't have a
Claude subscription or haven't found
great use out of it, definitely take a
look at open design. All right, let's
kick off. This is the apps that they're
building. They're popping up on our
screen trying to give us their version
of the build. And by the way, they
already look freaking great. [snorts]
Okay, taking a very quick peek.
>> [laughter]
>> Uh this is already fantastic. These
models are amazing these days. Oh,
excellent. Our first settings panel.
Interesting. See what we see here.
Ooh, all right.
>> [laughter]
>> Too many things happening at once, but
exciting. I kind of want to see that
settings panel again. Excellent. Kind of
awesome.
>> [laughter]
>> I'll just let them cook. I mean, windows
are popping up and closing all over the
place. A building four basically the
same application,
uh multiple builds all at the same time
is
maybe not the best idea.
But
it's awesome. Oh boy.
Uh so, it looks like it it won't be
surprising to many of you. It looks like
Claude is running quite a bit behind
GPT-55.
That tends to not be necessarily
surprising. GPT-55, both the goal and
the straight, are both complete, by the
way. And this is the straight version,
not on goal Opus build. Okay, we're
getting close. All except the goal
builder for Claude code is complete, but
the goal builder for Claude code, I
won't spoil it, is pretty impressive.
And it's also reporting the way that
it's been running its captures has been
off screen intentionally, apparently.
We'll see how that works out, but all of
the screen captures that it's been doing
and everything else, it's it's taking a
picture of what it renders and not the
rest of the system, which is really
quite nice. So, it's not accidentally
taking screenshots of my my desktop
while it's building. Oh, all right.
Well, we just got my very That's my very
first hint, and you saw it as well. It
looks kind of fantastic. I can't wait to
compare it to the non-gold version.
Uh yeah, very exciting. But, we're at 38
40 minutes at this point. That's worth
noting. Uh the the other ones I'll have
a whole table for you to see how this
worked out. For Codex, landed much much
faster, but it's not just about speed.
Okay, really briefly, while we're
waiting, let me talk to you about
{slash} goal. I'm not going to go deep
on this. I can do another video on it
sometime to really talk more about it.
What am I really doing when I say I have
a set of files, if you looked inside of
this repo, there's maybe 12 15 files and
that screenshot and some sample code
from the example that we were looking at
in design. That's all that's in here and
kind of describing what it should do.
So, it's kind of a complex set of files
that represent the request. Makes sense.
I have put that in to an agent like you
normally would. That agent will take a
look at that and then kind of put it
through a planning agent, figure out
what its plan is, work through that
plan, and work through it. So, that's
what's happening, and agents are very
very good at this these days. So, you
really don't need to do the planning
mode yourself. You don't really need to
break it down to a PRD, something like
that. But, there is another aspect of
work these days inside of some of these
agent systems, you know, these
harnesses, which is a goal. The concept
of saying I don't have an absolute set
of definitions of the things that I want
you to build. And by the way, this
system does not have an absolute set of
definitions. That's its intention is to
leave it a little bit vague and allow
the building system to kind of figure
out what else it might need to build,
but it does say you should be able to
open websites and applications, things
like that. A goal is there when you
don't have a series a known set of steps
to get to something and it allows the
agent system the harness harness system
to kind of figure out the path forward
to reach the end goal. The idea is you
would say I have a goal goal. Here's
what I'd like you to achieve. Here's the
stuff you can measure and here's the
level of measurement that you can use to
determine that you're done. So some
mechanism to kind of give it a read on
whether or not it's it's ready. I've
given it as you saw here, make sure that
it's professional release capable kind
of code. That's a very soft measurement
for it to hit, but it is going to use
that and try to keep running until it
feels like it meets that that target. So
that's really the reason I'm using goal
is to figure out on this somewhat
open-ended build, does it build better
when it's in goal mode or does it build
just the same as it would when you're
just doing a straight shot and let that
main agent determine the the path
forward to building it. So that's all
we're really interrogating. Okay, we're
finally done. Here we are, we're running
the first straight one from Claude Code
first. I'm going to give that a shot.
Let's just run it real quick and see. It
seems to have built. So let's try option
space one more time. Okay, I mean
>> [laughter]
>> maybe it doesn't perfectly fit the
brief. The idea is there, certainly. Up
and down moves around the different
areas. If I want to add an action, it
brings up an actions menu kind of with a
broken panel, interesting. Okay, so this
is straight. Remember, this is the
non-goal version on Claude Code. Let's
do straight on Codex and see if it did
any better cuz I'd say no. Okay, here we
go to run the Codex one straight. Built,
okay, not in the background, that's is
fine. What is our hot key here? There we
go. Hot key is control option space. All
right, if I click away from it, I put it
it works under my cursor. It obviously
is immediately better. It works there
perfectly well. It works around
perfectly well. My mouse cursor is
giving it focus to things. Oh, wow.
Yeah, this is really great. If I click
into one of these items, it brings up
the settings panel. Yeah, all kinds of
different excellent. These are the
different internal modes. Like the work
ring is this one, the web ring is that
one, and these are the different rings
or different modes themselves. So, very
cool. I think really entirely
meaningfully valuable right now. Like
this seems to work. I need to use it a
fair bit, and this is not a full test of
every single feature that's inside of
there, but really the first thing that
I'm trying to make sure is that it's
even close to really the intention that
we were asking for. Now, it's separated
each one of these objects. It's
separated the outside ring is not
animating. There's a lot of interesting
little details that it did miss that are
defined in the actual that that sample
that we handed off, but it got a lot
right. Like if this is where where we
started, I think I'd be very very happy
with it. So, that one that one gets a
pass. Good work. Okay, and so here we
are. We're going to stay in Codex once
again. I know this gets complicated.
We're staying in Codex and we're doing
the gold version of Codex. Let's see how
this one worked. Woo.
Yeah. Much better.
So, remember we were just looking at
Codex, and this again is Codex as the
circling right. Excellent. Golly. All
right, has all of that really right.
Let's get rid of it and move the mouse,
bring it back, and it does pay attention
to spaces on the screen. So, if it's off
screen, it stay it's forcibly staying on
screen. That's great.
Uh not that you'll be able to see this.
Let me try this on a different monitor.
Worked on a different monitor. Let's see
if that No, works great. Always going to
my cursor. Wow. [laughter] Okay, this
Yes, was the goal better? Yes.
Perceivably? Yes. Meaningfully? 100%
it's better. Let's take a look at the
settings that it has. Let's add one and
see what their setting mechanism looks
like.
Uh the different systems, what are on
each one of these different systems, how
you might kick them off. Oh, here's all
the parameters to be able to kick each
one of those off. So, yeah. I think this
is great. I I would have to play with
this. Obviously, it's pretty complex
piece of software, but yeah, I already
kind of want it. And it did load Google
when I selected that.
>> [laughter]
>> Of course. So, if I'm I'm in web and I
go to whatever's in brightness,
brightness is apparently weather.
Or I guess it's sunshine.
>> [laughter]
>> Okay. So, yes, I'm playing. I'm sorry.
Let's move on to the very last one,
which is Claude and kind of its goal
mechanism. I I have high hopes for this
one. All right, here we are in the
Claude goal version and it actually gave
me the way to run it as well as
information on how to kick it off once I
do. It's the only one that did that, so
give them real credit for that. All
right. All right, look for the ring icon
in the menu bar. Very cool. Okay, let's
Did you see that disclosure?
Oh, that's pretty sweet. The the animate
out. So, all of the icons are animating
out and it's subtle, but the icons blend
away from one another into the next icon
rather than being then rather than
snapping between them. See if I can get
one that's really pretty full. Nope,
it's got that loop back. So, this would
be moving something to a certain
kind of position in a circle. They move
it back to 0°
and it snaps back counteractively. So,
that's a very common CSS problem, common
everything. So, that's interesting to
see that they didn't see that. They
don't have the outside animating, but
they do have very tight lines.
Everything's closed in really correctly.
I mean, this feels very very close to
professional. Let's take a look at It
did dismiss the ring when we went into
the picker. This already looks much
better. Quite frankly, this feels like
it represents the app that I just saw in
some meaningful way.
are carried through here. You can choose
different colors for I really really
like what's going on here.
And was this appreciably different?
1,000%. If you remember, the last one
was janky all over the place. Really
didn't even work. So, is goal solving
some problems for me? Is it circling
back in? Yes. Now, I have a secret for
you. I want to share something with you.
Let's take a look at effective cost. How
many tokens did these burn? Okay.
I'll put this in a chart at the end, but
first, let me show you how it reports it
here. If we go to Codex first, Codex was
the fastest one. As you can see, the
straight version, so non-goal took 28
minutes to build, and in the end, it was
about 400,000, 500,000 tokens. Somewhere
in the half a million range. That was
the non-golden version, right? And it
only took roughly 28, 30 minutes. If we
go to the goal version here,
that's right. The goal version told me
straight up, goal finished faster, and
it was appreciably better, if you
recall. And with less token burn. That
is very interesting. Not sure about
that. That's surprising, but it was
faster. So, very interesting. So, 25
minutes, the best version over here
using goal. So, let's go take a look at
Claude code. When we look at the
straight version, this is just non-goal
straight on using it, it took, let's
see, 36 minutes. So,
8, 10 minutes longer, something like
that, to do the work, but that's
appreciable when we're talking half an
hour. Um and really it talked about uh
there are 36 million tokens, but it's
misleading cuz 35 million of them are
cash read, so it did the math out of
this, and it figured out that it really
was about 1.17 million tokens. So, this
is the straight one, 1.17
million tokens, which is kind of
incredible in 36 minutes. And that was
the one that was janky we couldn't even
use. So, very surprising. Last one, the
one we liked the most, let's take a look
at the goal one. Now, goal tells us that
it actually was about half a million
tokens in 48 minutes. Now, by far the
longest. This is another 12, 15 minutes,
or something along those lines. So, it
is getting long to do this work, but it
did burn not quite the least tokens, but
it burned the same amount as the gold
tokens or or or the straight version of
Codex. So, it kind of is in the middle,
but not nearly the 1. 17 that we saw
from the straight version, and it was by
far the best. So, if you're thinking
goals are bad because they're actually
going to go and eat all your tokens,
this is my example. I don't think that's
actually the case. So, it's worth
saying, sometimes when you have
open-ended kind of requests, things that
what you're asking for has enough
variability in it that you want teams
assigned to it, and you want the agents
to be able to figure out how to get
there, goals might be your way. You
might give them a shot. I really found
it to be much, much better here. And now
I have a cool application. Okay. So, I
hope you found that interesting. I
haven't done one of these bake-offs in a
while. This is really a lot of fun. I
always love building something with
multiple models. I'm sorry I gave you
four different builds. That's almost
impossible to keep in your head.
Obviously, goals in both cases were
were measurably, markedly better. Like,
it would be very, very easy to discern
for anybody to come in and go, "Well,
this one feels more like what was asked
for, at least, or this one works versus
really doesn't work.
So that's worth recognizing. There's a
lot of other features I didn't test to
figure out how integrated into the
operating system and those kinds of
things that I'll go off and test off
camera. I'll report back if if these
goals really didn't perform the way they
seem to have because maybe they made a
mess of everything else. I'll report
that back for sure. But at the same
time, really worth it to me. I would say
no question Claude wins. Claude wins
almost hands down. Um but Codex the
GPT-5 was really very excellent. If that
was the application I had to use, I'd be
fine. I could push back into the
settings panel a little bit and say
maybe work a little bit on this and do
maybe one or two other passes and get it
to where Claude got the the final build
using goal. So
you know, everything's very very close
these days. But here's all right. I
wanted to share with you the thing that
I'm going to do with this.
I run a benchmark and I share videos
about different models when they come
out. I actually have multiple models
that I've tested. The benchmark is open.
I'll put a link in the uh
in the description below. So you can go
look at the scores now. You don't have
to wait until my videos come out to
understand where MiniMax and
Grok and all of these other models
score. I haven't scored everything yet,
but I've scored a fair number of them.
I'll report those in other in other
videos. I am planning on doing a video
that has MiniMax and Kimmy and a couple
other models that people are very
interested in. I've had people ask me.
So I want to show those numbers, but my
goal is to actually allow a couple of
those models every now and then to try
to build this application so that we
also get a sense of how does it build? I
won't be able to do that with everything
cuz even these took half an hour to
build and an hour in the longest case
roughly. So I can't do it with
everything, but I can do it on a spot
check basis to where we can really
figure out it's not as good a model, but
how not a good as it is as it is.
How is it not You know what I'm saying?
All right. Thanks for coming along for
the ride on this one, and I'll see you
in the next one.
Ask follow-up questions or revisit key timestamps.
This video features a comparative 'bake-off' between AI coding agents (Codex and Claude) to build a complex, radial-menu desktop application for macOS. The creator tests both standard ('straight') prompting and 'goal'-oriented prompting to see which approach produces a superior application. The video details the design process using tools like Claude Design and ChatGPT, compares the results based on functionality and token consumption, and concludes with an assessment of which models and methodologies yielded the most professional outcome.
Videos recently processed by our community