HomeVideos

WTF Ox Alpha - A Free Unlimited AI Model

Now Playing

WTF Ox Alpha - A Free Unlimited AI Model

Transcript

224 segments

0:00

Well hello there so recently there's been a mystery online

0:03

there's a new ai model that's been kind of making the waves on

0:08

threads and x basically the entire ai community and this

0:12

model is called ox alpha so in this video i'm going to tell you

0:15

what ox alpha is and how you can use it and also kind of the

0:19

theories on where this ai model is from so let's get right into it

0:23

now if you're new to this channel my name is john i'm a senior

0:26

staff software engineer at meta and on this channel i go over

0:29

tech news ai content and a bunch of tutorials on doing agentic

0:34

coding so if you're interested in that please subscribe a lot of

0:37

you guys are not subscribed i think like 90 80 percent of you

0:40

guys are not subscribed but just watch all my stuff do me a

0:43

favor it was my birthday recently subscribe all right with that

0:46

intro out let's get right into it so wtf ox alpha a quick rundown of

0:51

ox alpha is that it's actually a stealth model it's basically just

0:55

got released anonymously a few days ago, I think like August

0:59

20th or so. Here are the high level specs. It has roughly

1:02

a million, you know, context token. It has like an output

1:05

maximum of like 131 tokens, which is pretty large. It supports

1:10

multimodal like input and videos, but it only generates text,

1:13

which is like interesting. It doesn't do image model output.

1:16

And right now it's totally free. And that's kind of the very

1:20

interesting

1:21

like point. I think the most interesting take on this entire model

1:24

is the marketing, like the stealth model launch. I think the entire

1:28

industry and the whole AI community, they really like this

1:31

interesting like approach. There's some story behind it,

1:35

in my opinion. But yeah, so as I was saying, it's free,

1:37

it's mysterious, and it has a lot of capacity right now,

1:40

20 trillion tokens. So one of the interesting takes here is that

1:44

because there's so much free tokens available during this like

1:47

launch window, is that it has to be one of the major labs.

1:50

That's kind of like the only places

1:52

where the capacity can come from. So people are thinking

1:55

maybe it's Gemini, you know, from Google or Grok.

1:58

Elon's been hinting at the new like Grok model, like I think 4.7

2:01

or some Chinese lab that is getting a lot of support from the

2:04

government or something, right? It got really good usage

2:07

right away. I think it had 260

2:10

something K unique users. And I think OpenCode like reported

2:14

that it was like number two on the high, like traffic wise,

2:17

like 5.4% of OpenCode traffic, which is really large.

2:20

And as you can see right here, it like spiked up and then

2:23

it dropped. And kind of like interest. I think a lot of people are

2:26

like testing it in and out here and then people kind of

2:28

lost interest.

2:29

And here is OpenCode's posts on like kind of the,

2:33

like capacities. This really went viral, which was

2:35

really interesting. But the real story, the thing that made this

2:38

like really insane beside it being like free to use, I think was

2:43

some of the initial benchmarks. And I think a lot of people ran

2:46

some of the DeepSuite benchmarks and it got like 80%,

2:49

which is like incredible. And people are kind of losing their

2:52

minds on this. But it turns out,

2:55

turns out that it was actually just a very small data set.

2:59

Only about eight out of tens has scored that much. And when

3:02

they did like a larger data set, it actually scored on the 58.4%.

3:07

So if you look at this chart, if it's 56%, it would say right around

3:12

here in between Opus 4.8 and QEM3 Max. It's definitely

3:16

nowhere near the frontier, right? But if it was 80%, it would be

3:19

better than Opus 5, which is like pretty good, right? Sold in

3:22

Opus 5. One important thing about the whole initial numbers

3:25

was that it wasn't really an official benchmark that was run.

3:29

It was just like community run. There was some people just

3:31

running it on their own. And then that kind of came out with a

3:35

lot of like weird results. But it really helped all of this mystery

3:38

behind it. And like kind of the initial small sample size made it

3:42

like really pop and people were like kind of freaking out about it.

3:45

This is a guy named Ben. He actually is one of the like early

3:48

posters about it. He was like claiming that it was like 80%,

3:52

but after like a lot of testing, it was a much smaller.

3:54

And, you know, he kind of like went through it a bunch, giving a

3:58

lot of takes that it has better voice than Cloud or HGPT, has like

4:02

decent design handling. It runs code for a long time, but it's like

4:06

very slow and it has some weirdness. But yeah, so a lot of

4:09

people were seeing these like initial takes and it just turned out

4:12

it was like completely wrong. And that's kind of where I think a

4:15

lot of why the initial usage was like really high,

4:18

but then it kind of like tapered down now.

4:20

And eventually people like tested it on their own using open

4:23

code or like open code to like open router or something

4:26

like that. And they, everyone kind of came to the similar

4:28

conclusions as I did. It's not really frontier level at all. So there

4:32

was a lot of takes like this where it basically said it scores along

4:36

Kimmy 2.6 and it's like about as good as Luna, maybe worse.

4:40

And here's like another take they're saying it's right around

4:43

like little above luna maybe and like around terra nowhere near

4:47

the frontier at all that's kind of high level where it is a lot of

4:51

people are kind of like guessing where things are because i

4:54

think the interesting thing is that gemini still hasn't released

4:57

their like latest models there's been delays

4:59

and you know grok

5:01

4.7 or so the elon was hinting at that so there's like these like

5:04

kind of moments that people are so waiting on and they

5:07

thought it was this and because of the initial kind of like wrong

5:10

benchmark scores from the community people are like jumping

5:13

their gun and you know claiming that oh the only people that

5:16

can do this kind of stuff is people with a lot of compute and you

5:20

know google and only like the major labs have

5:23

compute worth to spare one of the like other really interesting

5:27

things that i

5:28

found while doing my research on this topic is that there's like

5:31

this whole side

5:33

of people using a lot of like interesting engineering techniques

5:37

to try to like reverse engineer get the models to like reveal

5:42

like what model it is but yeah so there's all these like interesting

5:45

things like something called tokenizer

5:48

and also like java class path to try to like get the information to

5:52

be released i found these two articles that was really

5:54

interesting i really recommend you guys to take a look there's

5:57

this article by asim sharin and essentially it used an

6:01

uncensored like ai model to like hunt down what ox alpha was

6:06

and basically it's saying that it found it to be glm5 there's a lot

6:10

of interesting strategies here but essentially at a high level this

6:14

is kind of a rough shot of what that i think he used quen model

6:19

but basically he just had to go

6:21

and like hit these models multiple times

6:24

and try to like compare with different models and

6:27

basically, here's like, kind of like a rough outline. And there's

6:30

like things like tokenizer differentials. And then there's like,

6:33

you know, kind of like looking at the emojis and things like that,

6:36

and trying to like match it with different how the model outputs

6:40

on this normal and trying to like reverse it to some other model

6:43

that would like kind of match its behavior. And this is kind of

6:46

how a lot of the watermarking stuff is going to work in

6:48

the future, where it's mainly like a stats based where you're

6:52

going to be able to try to know based on how the

6:55

model answers, what that model is like, what the original

6:59

source model is like, if you distill the model, it will have a similar

7:03

characteristics and behaviors of

7:06

like the model that it is still from. So that's kind of the strategies

7:09

is doing. And basically, at the end of all of this, what he came

7:12

down to the conclusion of this author is that it's the JLM five

7:17

from z.ai. And there were some other people that making the

7:21

same claim using like different techniques, like this is so

7:24

AI generated, but like, the model will never say that it's like a

7:28

JLM model, but it's servers apparently did. And based on

7:33

this data, the server is just servers where it's coming from the

7:36

origin server is from JLM, there's nothing conclusive here.

7:40

But like from what I researched, and like kind of the leaks that

7:43

I've seen, and people doing these kind of different research,

7:45

it really does feel like it's a Chinese model, and most likely JLM

7:49

because I think like Quinn has already released a

7:51

model recently, the sick as well, and also Kimmy as well.

7:55

So like, all the major ones that I know of have already

7:57

released something. So I think JLM is the only one. I believe

8:00

we're waiting for. But yeah, it's still unknown. We don't know

8:02

anything else besides,

8:04

like what I've been saying. So yeah, so the highest theories is

8:08

that is JLM is probably the strongest fit, maybe is Gemini,

8:12

I've seen some random posts, there's like engineers who

8:15

work at. Google are hinting at something.

8:17

But I think honestly, for that, I think they're hinting at the models

8:20

that's been delayed, I'm sure like they're saying it's gonna be

8:22

the best and same with Elon, you know, it's like fits the timeline

8:26

of what he was saying for the next rock model. So that's the

8:29

only reason but I think from all the other data, it definitely feels

8:32

like JLM.

8:34

Okay, editor John here, minor update as news just got released

8:39

that the OX alpha model actually is. JLM 5.3 flash.

8:44

So yeah, I guess I was right. But yeah, it sucks because I was

8:48

like working on this.

8:49

Anyways i think there's some good takes

8:51

on this video that's not just about the model itself and the

8:55

mystery but yeah i hope you guys enjoy the rest of the video so

8:58

that's basically it i think the glm model is the ox alpha most

9:02

likely we'll see if i'm right if we'll see i'm wrong maybe it's a

9:05

completely new

9:06

like model that distilled from glm i don't know it could be

9:09

anything but

9:10

besides the model itself i think one of the interesting lessons

9:14

and i've been kind of feeling this lately that the ai war this whole

9:18

like industry it's being driven so much by social media i think

9:24

there's so much attention that is happening with these and

9:27

people essentially move from model to model like very quickly

9:32

and then they switch sides very easily like i think even a few

9:35

months ago

9:36

i think around when fable came out originally like clarko was at

9:40

the peak in my opinion of

9:42

very favorable public opinion, but right now because of the

9:46

cost of Fable being so expensive and I don't know, I just feel

9:49

like some of the sentiment has been

9:51

shifting to maybe use, you know, open AI models.

9:55

Sol is my favorite model right now. Also, like if you remember a

9:59

few months ago, like Google was

10:01

like top dog again, but right now they slipped a little bit,

10:04

you know, they haven't launched their latest model.

10:07

There's been a huge shift, a bunch of people leaving Google

10:10

right now, you know, Jeff Dean left for Discovery Loop, took a

10:14

lot of the great AI researchers and distributed

10:16

systems engineers. Also like Demis, the head of like Gemini,

10:20

now he's like Alphavis chief scientist instead. So like, we don't

10:24

know what happened there, but there's like a move and. I think

10:28

Sergey's coming back to lead Gemini. This industry is insane.

10:31

It's like, it just moves so quickly. It's changing so fast and

10:34

people have different opinions all the time. Like everyone sees

10:37

me as like the cloud code guy. I made a really popular,

10:40

cloud code video, but I'm a codex guy right now.

10:43

I use codex app the most. I still use cloud code. I still love

10:48

all of these tools, but you know, my daily driver right now

10:50

is codex. And then this whole stealth model job, I think is really

10:54

interesting because it's like a very fun

10:57

marketing strategy. Basically

10:59

this model released would have been perfect if it actually

11:02

delivered on the benchmark. I think from a

11:04

marketing perspective, they did really

11:07

good, but it didn't meet the benchmarks. So it kind of fell flat on

11:10

its face. The AI war is going to be won based on social media,

11:14

as funny as that is. Also how good the model actually is.

11:17

The media pushes people to use things, but they stay for the

11:21

model quality. So you kind of need both. But yeah, I hope you

11:24

guys enjoyed this video. I'm going to cover more news, AI tech

11:27

news in this channel. So don't forget to subscribe. Recently,

11:30

I started posting on threads and the X a lot more, and I did

11:33

an AMA

11:34

and I answered like 20 something questions or so that people

11:37

asked me. I don't know why, but people seem very interested in

11:41

like takes from some random like questions that only like L7

11:45

senior staff engineers could answer, I guess, I don't know.

11:48

I'm just like a normal dude, but I guess people were very

11:51

interested because that post really blew up. Check that video

11:54

out and let me know in the comments below what kind of other

11:56

videos that you would love to see on this channel. But until I

12:00

see you guys on the next one, bye.

Interactive Summary

This video explores the mystery surrounding 'Ox Alpha', an AI model that launched stealthily and sparked significant community speculation. The host details its features, the viral initial benchmarks that were later debunked, and the various techniques used to reverse-engineer its origins, eventually revealing it to be GLM-5.3 Flash. The video concludes with broader reflections on the fast-paced nature of the AI industry, the influence of social media on model adoption, and the importance of balancing marketing with actual performance.

Suggested questions

3 ready-made prompts