The Best AI Coding Models Right Now (OpenHands Index Update)
161 segments
We have our OpenHands Index. We have
Juan joining us today and I will share
the link to OpenHands Index, but I'm
going to stop presenting so that Juan
can share his screen and go over the
latest and the greatest models.
>> Yes, thank you so much, Amy. So,
um here's the index and you will see it.
>> Yes.
>> Okay, thank you. And um
so we have a a graph um
uh vertical axis is a score, horizontal
axis is a cost. So, it the best model
would be here, which doesn't exist, and
the worst model would be here, which we
wouldn't pay attention to.
Um I want to start by talking about
Fable um by saying that I'm not going to
talk about Fable. Uh Fable
did a great performance at a very high
cost and then disappear. So,
we cannot use it, so we will only talk
about models we can actually use.
Um all models we can actually use uh on
the Pareto line, that means the line
that's has uh the best uh
cost for a given performance. Um we can
see that Opus 48 and Opus 47 are very
close.
So, one could ask which one to use.
Like, they are Opus 48 seems a little
bit more expensive uh but performs a
little bit better. I'm going to break it
down per benchmark so you can decide
which one is better for you for your use
case.
And then on the Pareto line, we have a
new contender, which is Minimaxen 3, and
it's performing very good. Um I will
also break it down across all
benchmarks.
So, first,
issue resolution. This is your standard
hey, OpenHands fix this issue. We can
see that Opus 48 is on the Pareto line
and Opus 47 is performing worse at a
higher uh cost. Um so, if you're just
calling OpenHands late to fix issues,
like that's most of what they do, uh
going for Opus 48 uh would be your
default, not Opus 47.
And Minimaxen 3 uh performs very well.
If if you can pay attention, you notice
that Minimaxen 3 is like in the same
line as Opus 47. It actually performs
like a hair thin worse. They're both at
76. Um
Opus 46 is 76.8 and M M Minimax 3 76.4.
So, less than uh
than half a point. Um but it's very good
news that uh open source models are
performing uh are catching up with uh a
model that at the moment was amazing
like Opus 46.
Moving on towards uh testing, we uh
we then start testing separate from uh
each resolution. And here we can see
that uh Opus performs uh
a tiny bit better at Opus 4 uh although
48 is a tiny bit better at Opus 47.
Still enough to be the the the straight
to the choice. Um and Minimax performs
almost as well as Opus 48. So, if you
are doing test, if you're like just
doing test all day, um
you might want to use Minimax instead of
Minimax 3 instead of Opus 48 because
it's very very close and it's 1/4 of the
cost.
Um moving on towards front end, this is
all your UI fixes, um polishing
interface, anything graphical, and
looking at images and describing what's
inside the image. So,
we have our Pareto line. You can use
Opus. Minimax is also in the if you
can't going to see it because it's uh at
at
it's almost the same as Deep B4 Pro, but
they are both on the Pareto line. Um
my advice on front end is pick the one
you actually need. So, like
if you're testing images, do you need a
full description or do you need like uh
a high-level one? If you need a full
description, I would continue going for
Opus. If you're okay with something uh a
little more high-level, not that much
detail, um
Minimax 24 perform very well at uh much
more accessible cost.
Um then we are going to move to green
field. We call green field like by
coding, like when you are starting
something from a scratch. New tool, new
application, uh starting from zero. Um
Minimax is also on the Pareto line. Um
and we have
Opus
here. So, uh Opus 48 8
Oh,
I seem to Okay.
>> [snorts]
>> Uh
Opus 4 8 performs
uh a little bit better than um
than Opus 4 7 at a little bit more cost.
So, uh
we can see a history here. Opus 4 6 um
was more expensive than Opus 4 7, but
they they perform the same. Now, Opus 4
8 performs better, but at the same cost
as Opus 4 6. And curiously, Fable did
not perform that well. We have some
theories, but we're not going to talk
about them because we cannot choose
Fable anyway, so what's the point? Um
So, if you're about to go with 4 8, it
looks great. GPT-5 4 also keeps looking
great. GPT-5.5
It did have some regression.
Um
So, yeah. Either Opus 4 8 or GPT-5.5 or
Minimax and
Deep Seek Deep 4 Pro, they're all
looking great.
And last but not least, information
gathering. This is the one you do when
you have to create a a few a few
databases, feel know if something
happened in the real world, get an
official link, get evidence. Um GPT-5 is
performing the best here. Uh that's my
go-to when I really really need to find
something out. Um Minimax M3 is not in
the borderline first first of the
benchmarks in which it's not in the
borderline. Uh it's very close, but
Queen 3 uh 6 plus uh beats it.
So, you can use either of those.
Um we have improved our skills on that
end to actually get the link. So, all of
them should perform better as well um
than they used to before. Um
So, yeah. So, uh my advice, use Opus 4 8
if you really need performance for
almost everything, maybe not for
information gathering, and Minimax M3 uh
looks great overall. Um
So, yeah. Uh that's that's the news on
the index for this week.
Ask follow-up questions or revisit key timestamps.
Juan presents an update to the OpenHands Index, comparing various AI models based on their performance and cost. He highlights that models on the 'Pareto line' offer the best balance between quality and expense. Throughout the analysis, he evaluates different models across specific tasks like issue resolution, testing, UI development, greenfield projects, and information gathering, recommending Opus 4.8 for high performance and Minimax 3 as a cost-effective alternative for many use cases.
Videos recently processed by our community