HomeVideos

The Best AI Coding Models Right Now (OpenHands Index Update)

Now Playing

The Best AI Coding Models Right Now (OpenHands Index Update)

Transcript

161 segments

0:00

We have our OpenHands Index. We have

0:02

Juan joining us today and I will share

0:04

the link to OpenHands Index, but I'm

0:06

going to stop presenting so that Juan

0:08

can share his screen and go over the

0:10

latest and the greatest models.

0:13

>> Yes, thank you so much, Amy. So,

0:16

um here's the index and you will see it.

0:19

>> Yes.

0:19

>> Okay, thank you. And um

0:22

so we have a a graph um

0:25

uh vertical axis is a score, horizontal

0:27

axis is a cost. So, it the best model

0:29

would be here, which doesn't exist, and

0:30

the worst model would be here, which we

0:33

wouldn't pay attention to.

0:34

Um I want to start by talking about

0:36

Fable um by saying that I'm not going to

0:37

talk about Fable. Uh Fable

0:39

did a great performance at a very high

0:41

cost and then disappear. So,

0:43

we cannot use it, so we will only talk

0:45

about models we can actually use.

0:47

Um all models we can actually use uh on

0:50

the Pareto line, that means the line

0:51

that's has uh the best uh

0:53

cost for a given performance. Um we can

0:55

see that Opus 48 and Opus 47 are very

0:59

close.

1:01

So, one could ask which one to use.

1:03

Like, they are Opus 48 seems a little

1:05

bit more expensive uh but performs a

1:07

little bit better. I'm going to break it

1:08

down per benchmark so you can decide

1:10

which one is better for you for your use

1:12

case.

1:13

And then on the Pareto line, we have a

1:15

new contender, which is Minimaxen 3, and

1:17

it's performing very good. Um I will

1:19

also break it down across all

1:20

benchmarks.

1:21

So, first,

1:23

issue resolution. This is your standard

1:25

hey, OpenHands fix this issue. We can

1:27

see that Opus 48 is on the Pareto line

1:30

and Opus 47 is performing worse at a

1:33

higher uh cost. Um so, if you're just

1:36

calling OpenHands late to fix issues,

1:38

like that's most of what they do, uh

1:40

going for Opus 48 uh would be your

1:42

default, not Opus 47.

1:45

And Minimaxen 3 uh performs very well.

1:48

If if you can pay attention, you notice

1:50

that Minimaxen 3 is like in the same

1:52

line as Opus 47. It actually performs

1:54

like a hair thin worse. They're both at

1:56

76. Um

1:58

Opus 46 is 76.8 and M M Minimax 3 76.4.

2:03

So, less than uh

2:04

than half a point. Um but it's very good

2:06

news that uh open source models are

2:08

performing uh are catching up with uh a

2:11

model that at the moment was amazing

2:13

like Opus 46.

2:15

Moving on towards uh testing, we uh

2:18

we then start testing separate from uh

2:19

each resolution. And here we can see

2:21

that uh Opus performs uh

2:23

a tiny bit better at Opus 4 uh although

2:25

48 is a tiny bit better at Opus 47.

2:28

Still enough to be the the the straight

2:30

to the choice. Um and Minimax performs

2:33

almost as well as Opus 48. So, if you

2:35

are doing test, if you're like just

2:37

doing test all day, um

2:39

you might want to use Minimax instead of

2:41

Minimax 3 instead of Opus 48 because

2:43

it's very very close and it's 1/4 of the

2:45

cost.

2:47

Um moving on towards front end, this is

2:50

all your UI fixes, um polishing

2:53

interface, anything graphical, and

2:55

looking at images and describing what's

2:56

inside the image. So,

2:58

we have our Pareto line. You can use

3:00

Opus. Minimax is also in the if you

3:03

can't going to see it because it's uh at

3:05

at

3:06

it's almost the same as Deep B4 Pro, but

3:08

they are both on the Pareto line. Um

3:10

my advice on front end is pick the one

3:13

you actually need. So, like

3:14

if you're testing images, do you need a

3:17

full description or do you need like uh

3:18

a high-level one? If you need a full

3:20

description, I would continue going for

3:21

Opus. If you're okay with something uh a

3:23

little more high-level, not that much

3:25

detail, um

3:26

Minimax 24 perform very well at uh much

3:29

more accessible cost.

3:31

Um then we are going to move to green

3:33

field. We call green field like by

3:34

coding, like when you are starting

3:36

something from a scratch. New tool, new

3:38

application, uh starting from zero. Um

3:42

Minimax is also on the Pareto line. Um

3:45

and we have

3:47

Opus

3:49

here. So, uh Opus 48 8

3:53

Oh,

3:54

I seem to Okay.

3:55

>> [snorts]

3:55

>> Uh

3:56

Opus 4 8 performs

3:59

uh a little bit better than um

4:02

than Opus 4 7 at a little bit more cost.

4:04

So, uh

4:05

we can see a history here. Opus 4 6 um

4:07

was more expensive than Opus 4 7, but

4:09

they they perform the same. Now, Opus 4

4:11

8 performs better, but at the same cost

4:13

as Opus 4 6. And curiously, Fable did

4:16

not perform that well. We have some

4:17

theories, but we're not going to talk

4:18

about them because we cannot choose

4:20

Fable anyway, so what's the point? Um

4:23

So, if you're about to go with 4 8, it

4:24

looks great. GPT-5 4 also keeps looking

4:26

great. GPT-5.5

4:29

It did have some regression.

4:31

Um

4:32

So, yeah. Either Opus 4 8 or GPT-5.5 or

4:35

Minimax and

4:36

Deep Seek Deep 4 Pro, they're all

4:37

looking great.

4:38

And last but not least, information

4:40

gathering. This is the one you do when

4:42

you have to create a a few a few

4:45

databases, feel know if something

4:47

happened in the real world, get an

4:48

official link, get evidence. Um GPT-5 is

4:52

performing the best here. Uh that's my

4:53

go-to when I really really need to find

4:55

something out. Um Minimax M3 is not in

4:58

the borderline first first of the

4:59

benchmarks in which it's not in the

5:00

borderline. Uh it's very close, but

5:02

Queen 3 uh 6 plus uh beats it.

5:06

So, you can use either of those.

5:09

Um we have improved our skills on that

5:11

end to actually get the link. So, all of

5:13

them should perform better as well um

5:16

than they used to before. Um

5:18

So, yeah. So, uh my advice, use Opus 4 8

5:21

if you really need performance for

5:22

almost everything, maybe not for

5:25

information gathering, and Minimax M3 uh

5:27

looks great overall. Um

5:30

So, yeah. Uh that's that's the news on

5:32

the index for this week.

Interactive Summary

Juan presents an update to the OpenHands Index, comparing various AI models based on their performance and cost. He highlights that models on the 'Pareto line' offer the best balance between quality and expense. Throughout the analysis, he evaluates different models across specific tasks like issue resolution, testing, UI development, greenfield projects, and information gathering, recommending Opus 4.8 for high performance and Minimax 3 as a cost-effective alternative for many use cases.

Suggested questions

3 ready-made prompts