HomeVideos

Your AI Bill Is About to Collapse

Now Playing

Your AI Bill Is About to Collapse

Transcript

269 segments

0:00

So, the biggest AI labs and the smartest

0:02

money are all circling cheap Chinese

0:04

open models right now, ironically.

0:06

Export controls, memory chip supply

0:08

chains, sovereign AI, basically all of

0:10

it. And honestly, most of that has

0:12

nothing to do with running your business

0:14

day-to-day. I almost skipped the whole

0:15

thing for that reason, but then I

0:16

noticed the one lesson inside sitting

0:19

there. And I think it's the same thing

0:20

I've been telling you for about a couple

0:22

weeks now, and except this time, the

0:24

biggest companies on Earth are proving

0:26

it with their own money. So, I guess

0:27

maybe you should listen about [music]

0:28

this one. So, it's about time we stop

0:30

shopping for the smartest model and

0:32

start controlling the cost of getting

0:34

the job done. Let me walk you through

0:35

it.

0:38

For a long time, the story [music] was

0:40

very straightforward. The frontier

0:42

models, the biggest and smartest and the

0:44

most expensive ones, state-of-the-art

0:45

models were miles ahead. And if you

0:47

wanted good output, then you just pay

0:49

for the best, I think. And I think that

0:50

gap is closing way, way much faster than

0:53

most people actually realize. The open

0:55

models, the ones anyone can download and

0:58

run for free if you have the right

0:59

[music] device, are getting stronger

1:01

week by week, and I'm not exaggerating.

1:03

So, to give you a fresh example here,

1:05

there's a small resource-starved Chinese

1:07

lab called DeepSeek. I know you guys

1:08

know about this. And a few days ago,

1:10

they shipped something called the Spark.

1:12

And what that means is when an AI writes

1:14

one word at a time, and every word has

1:17

to kind of look back at everything

1:18

before it. And that looking back is

1:21

usually the slow part. The common fix is

1:22

to just hire a fast, cheap, and turn

1:25

model to draft a few different words

1:27

ahead while the big boss, like big boss

1:30

model, checks all of them at once. And

1:32

the boss keeps the final say, so the

1:34

quality never drops there. So, DeepSpark

1:36

made that intern very smarter. Where it

1:38

nudges each word based on the one right

1:41

before it. It stops drafting the second

1:43

it gets unsure. And it watches how busy

1:45

the servers are and adjusts on the fly.

1:48

Isn't that crazy? The reported result is

1:49

speed [music] gains and several times

1:51

the output on the same hardware with no

1:54

loss in quality, and they released the

1:56

whole entire thing open source and

1:58

already baked that into their latest

2:00

model. So, I'm going to break this one

2:01

down properly in an upcoming video, so I

2:03

will not go very deep around it here. My

2:05

point being here is a tiny lab with a

2:07

fraction of the money keeps making

2:09

getting up open model faster and cheaper

2:12

than the giants conglomerates like

2:14

opening a cloth ever thought was

2:16

possible and that trend is not slowing

2:18

down at all. So, the number you should

2:19

be optimizing changes [music] here. It

2:22

stops being tokens used for benchmark

2:24

scores and it actually becomes completed

2:26

outcomes per dollar. Now, what counts as

2:29

outcome depends on your business,

2:31

obviously. So, I cannot really hand you

2:33

one clean metric here. But, my rough

2:35

rule of thumb is just, you know, sales

2:37

is everything. If the AI is not moving

2:39

that, it's just decoration. Now, this is

2:41

where I think most people leave money on

2:43

the table here. Um I know that you might

2:45

think the smart move is to use the best

2:47

model, say Favored-5, recent model, for

2:49

the whole thing, planning, building,

2:51

review, iterating, all of it. One brain

2:54

start to finish feels really safe,

2:56

right? Honestly, that's the expensive

2:58

mistake.

3:01

The hard part of almost any build is

3:03

just the front end. The architecture and

3:06

the planning and that is exactly where

3:07

you want your sharpest model. So,

3:09

Favored-5 earns its spot right there.

3:12

But, the middle, the actual building and

3:14

the implementation is mostly bulk work

3:16

and a good enough local model like

3:17

Minimax M3 or GLM 5.2, Gemini 2.7, they

3:20

can just carry [music] that part just

3:22

fine. Then, you bring Favored-5 back in

3:24

for the review pass just once to catch

3:27

what the cheaper model missed. After

3:29

that, you drop back down again to

3:31

iterate and keep building on it with the

3:33

local model. Same job, better result and

3:34

here is the part people usually miss

3:37

about. So, you stop burning through your

3:38

usage limits because you are not

3:40

spending your most expensive model on

3:42

work that never needed it in the first

3:44

place. So, being fixated on the idea of

3:46

I have to use that best model for every

3:48

single step is not exactly I'm not

3:50

actually the safe choice here. So, my

3:52

take is that the choice is that runs you

3:55

out of runway and makes you miss the

3:56

smarter [music] play.

4:00

So, why don't we think of it like if you

4:02

had one brilliant, expensive specialist

4:04

on your team, you would never have them

4:06

standing at the copier all day at a

4:09

reception desk. You would just put them

4:10

on the hardest problems and let a

4:12

cheaper hire handle the repetitive

4:14

stuff, right? Obviously. Reaching for

4:16

the most powerful model on every tiny

4:17

task is the exact same waste. You are

4:20

literally paying specialist rates for

4:22

photocopier and I went a lot deeper on

4:24

this whole routing idea in my recent

4:26

video, including the architecture first

4:28

laws that you guys just walked you

4:29

through. [music] So, if you want the

4:30

full version, why don't you go watch

4:31

this video up top, okay?

4:35

Now, the caveat, the catch that I can

4:37

think of is because good enough plus a

4:39

little custom tuning is a real winning

4:42

combo, but only if you get the timing

4:44

right. So, tuning a model too early can

4:47

actually work against you. I don't know

4:49

if you ever heard about this one. And

4:51

here's my take on it. You are optimizing

4:53

a guess. Early on, you do not really

4:55

know your real task yet, right? So,

4:57

tuning locks into these assumptions and

5:00

when the assumption is wrong, you just

5:01

do not get a wrong answer. get a wrong

5:03

answer that the model is now very

5:05

confident about. You are teaching it

5:08

your current mess. Early on, you only

5:10

have a few examples and they are very

5:12

noisy, you know, not as optimized, neat

5:15

enough. And tuning amplifies whatever

5:17

you feed it, so mistakes, all of it,

5:19

right? You give up the free upgrades.

5:21

The base models get smarter every few

5:23

months with zero effort on your part.

5:24

So, a model you tuned too early

5:26

basically freezes in place while the

5:29

frontier walks right past it. I hope

5:31

that makes sense. The everyday version

5:32

of the same mistake is just wearing

5:34

bigger clothes. You lock your whole

5:36

setup one fixed model, rigid prompts, a

5:39

built pipeline before you have even

5:41

watched how you ever use it. Both are

5:43

the same error at the core, I think. You

5:45

are optimizing before you

5:46

>> [music]

5:46

>> validated. And notice how Coinbase did

5:48

it in the right order. They watched

5:49

their real users first, found that most

5:51

of their people never even hit the

5:53

limit, and only then moved their

5:55

default. [music] So, you observe first,

5:57

then you optimize, not the other way

5:58

around.

6:01

You know, there's one thing I kind of

6:03

want to say to you plainly here. Anyone

6:05

who claims that their way of using AI is

6:07

the one true way is wrong, completely

6:10

wrong. And the reason is nobody has

6:12

figured this out fully yet. It's still

6:14

that new technology. So, the real skill

6:16

to build is not memorizing someone

6:18

else's system. I think it's developing

6:21

the eye and the mind to tell the

6:23

difference between an opinion and a

6:25

fact. Whatever comes out of anyone's

6:26

mouth, including myself, my mouth on

6:28

this [music] stuff, mine included, is

6:30

just a current opinion. It's not settled

6:33

truth. So, I hope you see it that way,

6:34

okay?

6:37

And the risk almost nobody prices in

6:39

right now, a frontier model that you

6:41

rent can change [music]

6:42

under you at any time. It can get

6:45

restricted, you know, you we've already

6:47

experienced it with Stable 5, or

6:48

re-priced, wrapped their new approval

6:50

steps, or just shuts up completely. And

6:53

you don't have control any of that. The

6:55

vendor does. [music] But a model you can

6:56

download and run on your own machine

6:58

cannot be switched off on you unless you

7:01

lose electricity, right? It might be

7:02

slower and might be less capable on the

7:04

really hard task, but it's yours. For a

7:06

small business, I think that continuity

7:08

is the quiet thing that actually is

7:10

really important. The scary scenario

7:12

here is not that your competitor gets a

7:14

smaller models, you know. It's you

7:16

waking up one morning to find the tool

7:18

your whole workflow depends on got

7:21

pulled or tripled in price overnight.

7:22

[music] So, owning the workflow and

7:24

being able to swap the engine underneath

7:26

it whenever you want is how you protect

7:29

yourself from that. And we're getting

7:30

there, guys. The open-source models,

7:32

they're getting crazier and crazier and

7:34

crazier, you know?

7:37

Okay, now, I want to kind of zoom out

7:39

with me for a second because this whole

7:41

thing rhymes with something a lot older.

7:43

You know, cheap AI is just cost collapse

7:46

and cost collapses have happened many

7:48

times before this one. My verdict on how

7:51

they play out usually, when an expensive

7:53

product gets cheap [music] the moment

7:54

the resources to make it up become

7:56

widely available and from there only two

7:59

things can happen I think. If the

8:00

product is actually crazy good, the

8:03

value relocates to the general public.

8:05

So everybody gets access and everybody

8:07

benefits from it. Whereas if the product

8:09

was mostly just hype around it and it

8:11

[music] it'll just disappear into thin

8:13

air the second it gets cheap enough for

8:15

everyone to see clearly. And cheap AI is

8:18

running that exact same test right now

8:21

on you. So the only real question that

8:22

matters is the one I keep coming back to

8:25

lately. Is your AI use producing real

8:27

outcomes for you or is it just software

8:30

dressed up as progress? You already know

8:32

my field number on this by now if you go

8:33

to my website, seven out of 10 owners

8:35

who buy an AI system or pilot never use

8:37

that thing ever. So this cost collapse

8:40

does not fix that for you. It widens the

8:41

gap between the people who use AI for

8:43

real and the people who only collect

8:45

subscriptions. So pick the right side of

8:47

the gap, own your workflow and watch

8:49

your sales instead of your token count.

8:51

That's what's happening around in the

8:52

world I think. Anyway, that's my take on

8:54

it. If there's anything that you'd like

8:55

to see more from me, why don't you

8:57

comment down below and hype this video

8:59

if you don't mind because that's how

9:00

this channel runs. You ask, I build, we

9:02

all learn. See you in the next video.

Interactive Summary

The video discusses the shifting landscape of AI development, emphasizing a transition from solely relying on expensive 'frontier' models to leveraging cost-effective, open-source alternatives. The creator argues that businesses should stop focusing on benchmark scores and instead optimize for 'completed outcomes per dollar,' using powerful models only for complex architectural tasks while employing smaller, cheaper models for repetitive work. The video also warns against premature model tuning and stresses the importance of owning one's workflow to avoid risks associated with vendor lock-in and price volatility.

Suggested questions

3 ready-made prompts