Your AI Bill Is About to Collapse
269 segments
So, the biggest AI labs and the smartest
money are all circling cheap Chinese
open models right now, ironically.
Export controls, memory chip supply
chains, sovereign AI, basically all of
it. And honestly, most of that has
nothing to do with running your business
day-to-day. I almost skipped the whole
thing for that reason, but then I
noticed the one lesson inside sitting
there. And I think it's the same thing
I've been telling you for about a couple
weeks now, and except this time, the
biggest companies on Earth are proving
it with their own money. So, I guess
maybe you should listen about [music]
this one. So, it's about time we stop
shopping for the smartest model and
start controlling the cost of getting
the job done. Let me walk you through
it.
For a long time, the story [music] was
very straightforward. The frontier
models, the biggest and smartest and the
most expensive ones, state-of-the-art
models were miles ahead. And if you
wanted good output, then you just pay
for the best, I think. And I think that
gap is closing way, way much faster than
most people actually realize. The open
models, the ones anyone can download and
run for free if you have the right
[music] device, are getting stronger
week by week, and I'm not exaggerating.
So, to give you a fresh example here,
there's a small resource-starved Chinese
lab called DeepSeek. I know you guys
know about this. And a few days ago,
they shipped something called the Spark.
And what that means is when an AI writes
one word at a time, and every word has
to kind of look back at everything
before it. And that looking back is
usually the slow part. The common fix is
to just hire a fast, cheap, and turn
model to draft a few different words
ahead while the big boss, like big boss
model, checks all of them at once. And
the boss keeps the final say, so the
quality never drops there. So, DeepSpark
made that intern very smarter. Where it
nudges each word based on the one right
before it. It stops drafting the second
it gets unsure. And it watches how busy
the servers are and adjusts on the fly.
Isn't that crazy? The reported result is
speed [music] gains and several times
the output on the same hardware with no
loss in quality, and they released the
whole entire thing open source and
already baked that into their latest
model. So, I'm going to break this one
down properly in an upcoming video, so I
will not go very deep around it here. My
point being here is a tiny lab with a
fraction of the money keeps making
getting up open model faster and cheaper
than the giants conglomerates like
opening a cloth ever thought was
possible and that trend is not slowing
down at all. So, the number you should
be optimizing changes [music] here. It
stops being tokens used for benchmark
scores and it actually becomes completed
outcomes per dollar. Now, what counts as
outcome depends on your business,
obviously. So, I cannot really hand you
one clean metric here. But, my rough
rule of thumb is just, you know, sales
is everything. If the AI is not moving
that, it's just decoration. Now, this is
where I think most people leave money on
the table here. Um I know that you might
think the smart move is to use the best
model, say Favored-5, recent model, for
the whole thing, planning, building,
review, iterating, all of it. One brain
start to finish feels really safe,
right? Honestly, that's the expensive
mistake.
The hard part of almost any build is
just the front end. The architecture and
the planning and that is exactly where
you want your sharpest model. So,
Favored-5 earns its spot right there.
But, the middle, the actual building and
the implementation is mostly bulk work
and a good enough local model like
Minimax M3 or GLM 5.2, Gemini 2.7, they
can just carry [music] that part just
fine. Then, you bring Favored-5 back in
for the review pass just once to catch
what the cheaper model missed. After
that, you drop back down again to
iterate and keep building on it with the
local model. Same job, better result and
here is the part people usually miss
about. So, you stop burning through your
usage limits because you are not
spending your most expensive model on
work that never needed it in the first
place. So, being fixated on the idea of
I have to use that best model for every
single step is not exactly I'm not
actually the safe choice here. So, my
take is that the choice is that runs you
out of runway and makes you miss the
smarter [music] play.
So, why don't we think of it like if you
had one brilliant, expensive specialist
on your team, you would never have them
standing at the copier all day at a
reception desk. You would just put them
on the hardest problems and let a
cheaper hire handle the repetitive
stuff, right? Obviously. Reaching for
the most powerful model on every tiny
task is the exact same waste. You are
literally paying specialist rates for
photocopier and I went a lot deeper on
this whole routing idea in my recent
video, including the architecture first
laws that you guys just walked you
through. [music] So, if you want the
full version, why don't you go watch
this video up top, okay?
Now, the caveat, the catch that I can
think of is because good enough plus a
little custom tuning is a real winning
combo, but only if you get the timing
right. So, tuning a model too early can
actually work against you. I don't know
if you ever heard about this one. And
here's my take on it. You are optimizing
a guess. Early on, you do not really
know your real task yet, right? So,
tuning locks into these assumptions and
when the assumption is wrong, you just
do not get a wrong answer. get a wrong
answer that the model is now very
confident about. You are teaching it
your current mess. Early on, you only
have a few examples and they are very
noisy, you know, not as optimized, neat
enough. And tuning amplifies whatever
you feed it, so mistakes, all of it,
right? You give up the free upgrades.
The base models get smarter every few
months with zero effort on your part.
So, a model you tuned too early
basically freezes in place while the
frontier walks right past it. I hope
that makes sense. The everyday version
of the same mistake is just wearing
bigger clothes. You lock your whole
setup one fixed model, rigid prompts, a
built pipeline before you have even
watched how you ever use it. Both are
the same error at the core, I think. You
are optimizing before you
>> [music]
>> validated. And notice how Coinbase did
it in the right order. They watched
their real users first, found that most
of their people never even hit the
limit, and only then moved their
default. [music] So, you observe first,
then you optimize, not the other way
around.
You know, there's one thing I kind of
want to say to you plainly here. Anyone
who claims that their way of using AI is
the one true way is wrong, completely
wrong. And the reason is nobody has
figured this out fully yet. It's still
that new technology. So, the real skill
to build is not memorizing someone
else's system. I think it's developing
the eye and the mind to tell the
difference between an opinion and a
fact. Whatever comes out of anyone's
mouth, including myself, my mouth on
this [music] stuff, mine included, is
just a current opinion. It's not settled
truth. So, I hope you see it that way,
okay?
And the risk almost nobody prices in
right now, a frontier model that you
rent can change [music]
under you at any time. It can get
restricted, you know, you we've already
experienced it with Stable 5, or
re-priced, wrapped their new approval
steps, or just shuts up completely. And
you don't have control any of that. The
vendor does. [music] But a model you can
download and run on your own machine
cannot be switched off on you unless you
lose electricity, right? It might be
slower and might be less capable on the
really hard task, but it's yours. For a
small business, I think that continuity
is the quiet thing that actually is
really important. The scary scenario
here is not that your competitor gets a
smaller models, you know. It's you
waking up one morning to find the tool
your whole workflow depends on got
pulled or tripled in price overnight.
[music] So, owning the workflow and
being able to swap the engine underneath
it whenever you want is how you protect
yourself from that. And we're getting
there, guys. The open-source models,
they're getting crazier and crazier and
crazier, you know?
Okay, now, I want to kind of zoom out
with me for a second because this whole
thing rhymes with something a lot older.
You know, cheap AI is just cost collapse
and cost collapses have happened many
times before this one. My verdict on how
they play out usually, when an expensive
product gets cheap [music] the moment
the resources to make it up become
widely available and from there only two
things can happen I think. If the
product is actually crazy good, the
value relocates to the general public.
So everybody gets access and everybody
benefits from it. Whereas if the product
was mostly just hype around it and it
[music] it'll just disappear into thin
air the second it gets cheap enough for
everyone to see clearly. And cheap AI is
running that exact same test right now
on you. So the only real question that
matters is the one I keep coming back to
lately. Is your AI use producing real
outcomes for you or is it just software
dressed up as progress? You already know
my field number on this by now if you go
to my website, seven out of 10 owners
who buy an AI system or pilot never use
that thing ever. So this cost collapse
does not fix that for you. It widens the
gap between the people who use AI for
real and the people who only collect
subscriptions. So pick the right side of
the gap, own your workflow and watch
your sales instead of your token count.
That's what's happening around in the
world I think. Anyway, that's my take on
it. If there's anything that you'd like
to see more from me, why don't you
comment down below and hype this video
if you don't mind because that's how
this channel runs. You ask, I build, we
all learn. See you in the next video.
Ask follow-up questions or revisit key timestamps.
The video discusses the shifting landscape of AI development, emphasizing a transition from solely relying on expensive 'frontier' models to leveraging cost-effective, open-source alternatives. The creator argues that businesses should stop focusing on benchmark scores and instead optimize for 'completed outcomes per dollar,' using powerful models only for complex architectural tasks while employing smaller, cheaper models for repetitive work. The video also warns against premature model tuning and stresses the importance of owning one's workflow to avoid risks associated with vendor lock-in and price volatility.
Videos recently processed by our community