WTF Ox Alpha - A Free Unlimited AI Model
224 segments
Well hello there so recently there's been a mystery online
there's a new ai model that's been kind of making the waves on
threads and x basically the entire ai community and this
model is called ox alpha so in this video i'm going to tell you
what ox alpha is and how you can use it and also kind of the
theories on where this ai model is from so let's get right into it
now if you're new to this channel my name is john i'm a senior
staff software engineer at meta and on this channel i go over
tech news ai content and a bunch of tutorials on doing agentic
coding so if you're interested in that please subscribe a lot of
you guys are not subscribed i think like 90 80 percent of you
guys are not subscribed but just watch all my stuff do me a
favor it was my birthday recently subscribe all right with that
intro out let's get right into it so wtf ox alpha a quick rundown of
ox alpha is that it's actually a stealth model it's basically just
got released anonymously a few days ago, I think like August
20th or so. Here are the high level specs. It has roughly
a million, you know, context token. It has like an output
maximum of like 131 tokens, which is pretty large. It supports
multimodal like input and videos, but it only generates text,
which is like interesting. It doesn't do image model output.
And right now it's totally free. And that's kind of the very
interesting
like point. I think the most interesting take on this entire model
is the marketing, like the stealth model launch. I think the entire
industry and the whole AI community, they really like this
interesting like approach. There's some story behind it,
in my opinion. But yeah, so as I was saying, it's free,
it's mysterious, and it has a lot of capacity right now,
20 trillion tokens. So one of the interesting takes here is that
because there's so much free tokens available during this like
launch window, is that it has to be one of the major labs.
That's kind of like the only places
where the capacity can come from. So people are thinking
maybe it's Gemini, you know, from Google or Grok.
Elon's been hinting at the new like Grok model, like I think 4.7
or some Chinese lab that is getting a lot of support from the
government or something, right? It got really good usage
right away. I think it had 260
something K unique users. And I think OpenCode like reported
that it was like number two on the high, like traffic wise,
like 5.4% of OpenCode traffic, which is really large.
And as you can see right here, it like spiked up and then
it dropped. And kind of like interest. I think a lot of people are
like testing it in and out here and then people kind of
lost interest.
And here is OpenCode's posts on like kind of the,
like capacities. This really went viral, which was
really interesting. But the real story, the thing that made this
like really insane beside it being like free to use, I think was
some of the initial benchmarks. And I think a lot of people ran
some of the DeepSuite benchmarks and it got like 80%,
which is like incredible. And people are kind of losing their
minds on this. But it turns out,
turns out that it was actually just a very small data set.
Only about eight out of tens has scored that much. And when
they did like a larger data set, it actually scored on the 58.4%.
So if you look at this chart, if it's 56%, it would say right around
here in between Opus 4.8 and QEM3 Max. It's definitely
nowhere near the frontier, right? But if it was 80%, it would be
better than Opus 5, which is like pretty good, right? Sold in
Opus 5. One important thing about the whole initial numbers
was that it wasn't really an official benchmark that was run.
It was just like community run. There was some people just
running it on their own. And then that kind of came out with a
lot of like weird results. But it really helped all of this mystery
behind it. And like kind of the initial small sample size made it
like really pop and people were like kind of freaking out about it.
This is a guy named Ben. He actually is one of the like early
posters about it. He was like claiming that it was like 80%,
but after like a lot of testing, it was a much smaller.
And, you know, he kind of like went through it a bunch, giving a
lot of takes that it has better voice than Cloud or HGPT, has like
decent design handling. It runs code for a long time, but it's like
very slow and it has some weirdness. But yeah, so a lot of
people were seeing these like initial takes and it just turned out
it was like completely wrong. And that's kind of where I think a
lot of why the initial usage was like really high,
but then it kind of like tapered down now.
And eventually people like tested it on their own using open
code or like open code to like open router or something
like that. And they, everyone kind of came to the similar
conclusions as I did. It's not really frontier level at all. So there
was a lot of takes like this where it basically said it scores along
Kimmy 2.6 and it's like about as good as Luna, maybe worse.
And here's like another take they're saying it's right around
like little above luna maybe and like around terra nowhere near
the frontier at all that's kind of high level where it is a lot of
people are kind of like guessing where things are because i
think the interesting thing is that gemini still hasn't released
their like latest models there's been delays
and you know grok
4.7 or so the elon was hinting at that so there's like these like
kind of moments that people are so waiting on and they
thought it was this and because of the initial kind of like wrong
benchmark scores from the community people are like jumping
their gun and you know claiming that oh the only people that
can do this kind of stuff is people with a lot of compute and you
know google and only like the major labs have
compute worth to spare one of the like other really interesting
things that i
found while doing my research on this topic is that there's like
this whole side
of people using a lot of like interesting engineering techniques
to try to like reverse engineer get the models to like reveal
like what model it is but yeah so there's all these like interesting
things like something called tokenizer
and also like java class path to try to like get the information to
be released i found these two articles that was really
interesting i really recommend you guys to take a look there's
this article by asim sharin and essentially it used an
uncensored like ai model to like hunt down what ox alpha was
and basically it's saying that it found it to be glm5 there's a lot
of interesting strategies here but essentially at a high level this
is kind of a rough shot of what that i think he used quen model
but basically he just had to go
and like hit these models multiple times
and try to like compare with different models and
basically, here's like, kind of like a rough outline. And there's
like things like tokenizer differentials. And then there's like,
you know, kind of like looking at the emojis and things like that,
and trying to like match it with different how the model outputs
on this normal and trying to like reverse it to some other model
that would like kind of match its behavior. And this is kind of
how a lot of the watermarking stuff is going to work in
the future, where it's mainly like a stats based where you're
going to be able to try to know based on how the
model answers, what that model is like, what the original
source model is like, if you distill the model, it will have a similar
characteristics and behaviors of
like the model that it is still from. So that's kind of the strategies
is doing. And basically, at the end of all of this, what he came
down to the conclusion of this author is that it's the JLM five
from z.ai. And there were some other people that making the
same claim using like different techniques, like this is so
AI generated, but like, the model will never say that it's like a
JLM model, but it's servers apparently did. And based on
this data, the server is just servers where it's coming from the
origin server is from JLM, there's nothing conclusive here.
But like from what I researched, and like kind of the leaks that
I've seen, and people doing these kind of different research,
it really does feel like it's a Chinese model, and most likely JLM
because I think like Quinn has already released a
model recently, the sick as well, and also Kimmy as well.
So like, all the major ones that I know of have already
released something. So I think JLM is the only one. I believe
we're waiting for. But yeah, it's still unknown. We don't know
anything else besides,
like what I've been saying. So yeah, so the highest theories is
that is JLM is probably the strongest fit, maybe is Gemini,
I've seen some random posts, there's like engineers who
work at. Google are hinting at something.
But I think honestly, for that, I think they're hinting at the models
that's been delayed, I'm sure like they're saying it's gonna be
the best and same with Elon, you know, it's like fits the timeline
of what he was saying for the next rock model. So that's the
only reason but I think from all the other data, it definitely feels
like JLM.
Okay, editor John here, minor update as news just got released
that the OX alpha model actually is. JLM 5.3 flash.
So yeah, I guess I was right. But yeah, it sucks because I was
like working on this.
Anyways i think there's some good takes
on this video that's not just about the model itself and the
mystery but yeah i hope you guys enjoy the rest of the video so
that's basically it i think the glm model is the ox alpha most
likely we'll see if i'm right if we'll see i'm wrong maybe it's a
completely new
like model that distilled from glm i don't know it could be
anything but
besides the model itself i think one of the interesting lessons
and i've been kind of feeling this lately that the ai war this whole
like industry it's being driven so much by social media i think
there's so much attention that is happening with these and
people essentially move from model to model like very quickly
and then they switch sides very easily like i think even a few
months ago
i think around when fable came out originally like clarko was at
the peak in my opinion of
very favorable public opinion, but right now because of the
cost of Fable being so expensive and I don't know, I just feel
like some of the sentiment has been
shifting to maybe use, you know, open AI models.
Sol is my favorite model right now. Also, like if you remember a
few months ago, like Google was
like top dog again, but right now they slipped a little bit,
you know, they haven't launched their latest model.
There's been a huge shift, a bunch of people leaving Google
right now, you know, Jeff Dean left for Discovery Loop, took a
lot of the great AI researchers and distributed
systems engineers. Also like Demis, the head of like Gemini,
now he's like Alphavis chief scientist instead. So like, we don't
know what happened there, but there's like a move and. I think
Sergey's coming back to lead Gemini. This industry is insane.
It's like, it just moves so quickly. It's changing so fast and
people have different opinions all the time. Like everyone sees
me as like the cloud code guy. I made a really popular,
cloud code video, but I'm a codex guy right now.
I use codex app the most. I still use cloud code. I still love
all of these tools, but you know, my daily driver right now
is codex. And then this whole stealth model job, I think is really
interesting because it's like a very fun
marketing strategy. Basically
this model released would have been perfect if it actually
delivered on the benchmark. I think from a
marketing perspective, they did really
good, but it didn't meet the benchmarks. So it kind of fell flat on
its face. The AI war is going to be won based on social media,
as funny as that is. Also how good the model actually is.
The media pushes people to use things, but they stay for the
model quality. So you kind of need both. But yeah, I hope you
guys enjoyed this video. I'm going to cover more news, AI tech
news in this channel. So don't forget to subscribe. Recently,
I started posting on threads and the X a lot more, and I did
an AMA
and I answered like 20 something questions or so that people
asked me. I don't know why, but people seem very interested in
like takes from some random like questions that only like L7
senior staff engineers could answer, I guess, I don't know.
I'm just like a normal dude, but I guess people were very
interested because that post really blew up. Check that video
out and let me know in the comments below what kind of other
videos that you would love to see on this channel. But until I
see you guys on the next one, bye.
Ask follow-up questions or revisit key timestamps.
This video explores the mystery surrounding 'Ox Alpha', an AI model that launched stealthily and sparked significant community speculation. The host details its features, the viral initial benchmarks that were later debunked, and the various techniques used to reverse-engineer its origins, eventually revealing it to be GLM-5.3 Flash. The video concludes with broader reflections on the fast-paced nature of the AI industry, the influence of social media on model adoption, and the importance of balancing marketing with actual performance.
Videos recently processed by our community