HomeVideos

I'm Buying Every Share I Can (Investors Aren't Ready)

Now Playing

I'm Buying Every Share I Can (Investors Aren't Ready)

Transcript

456 segments

0:00

The single biggest argument against AI

0:02

stocks just died. Wall Street analysts

0:04

have been saying the same thing for

0:06

years now. AI is a bubble and the entire

0:08

tech sector will crash as soon as

0:10

spending slows down even a little. But

0:13

three major developments just put an end

0:15

to that argument and kicked off what

0:17

could be the next phase of the AI era.

0:19

My name is Alex and I spent 8 years as

0:21

an electrical engineer and AI researcher

0:23

at MIT. And I've never seen chip

0:26

companies move this fast. Let me show

0:28

you what's going on and how I'm

0:30

investing in it. Your time is valuable,

0:32

so let's get right into it. The biggest

0:34

case against AI has always been that

0:36

spending will eventually slow down, and

0:39

when it does, every stock that lives and

0:41

dies by that spending will go down with

0:43

it. It's only a matter of time.

0:46

Hyperscalers like Google, Amazon,

0:48

Microsoft, and Meta Platforms report

0:50

their capex budgets every quarter, and

0:52

every quarter, it's been growing like

0:54

crazy. In fact, their AI spending has

0:56

been compounding at over 60% per year

0:59

since 2023, and it's actually expected

1:02

to speed up. These four companies are

1:04

currently on course to spend over $700

1:07

billion on data centers this year alone,

1:10

compared to the $375 billion they spent

1:14

in 2025. That's a 90% increase

1:17

year-over-year. And it's coming at a

1:19

huge cost. Amazon's free cash flow went

1:22

from $18 billion to negative7.6 billion

1:26

over the last 12 months. But just a few

1:28

weeks ago, Amazon increased their capex

1:30

budget for 2026 from $200 billion to

1:34

$220 billion due to the price of memory

1:37

chips. That same week, Alphabet added

1:39

$15 billion to its capex budget while

1:42

reporting that their free cash flow

1:44

turned negative for the first time since

1:46

Google went public in 2004. and Meta's

1:49

free cash flows are down by 91%. You'd

1:52

think that spending would slow down once

1:54

the biggest companies driving it burn

1:56

through all their cash flows and then

1:58

some, but instead they're borrowing

2:00

money and issuing new shares. Amazon

2:03

sold $37 billion worth of bonds in

2:06

quarter 1. Alphabet issued over $65

2:09

billion worth of bonds so far this year

2:11

and announced an $85 billion equity

2:14

raise in June, the biggest raise in

2:16

American corporate history. debt covered

2:18

9% of hyperscaler capex in 2024. Today,

2:23

it covers 32%. It's important for

2:25

investors to understand why spending

2:27

isn't slowing down and why these

2:29

companies keep investing in AI at all

2:32

costs. Data center infrastructure is a

2:34

zero- sum game. There's only so much

2:37

land and grid connected power in the

2:39

first place, and every acre and gawatt

2:41

you don't get is one that your

2:43

competition does. A big new data center

2:45

can wait anywhere from 2 to 6 years for

2:48

a grid connection. And even when it's

2:50

all bought up, spending will just shift

2:52

to the next big bottleneck. Cooling,

2:55

compute density, network, and memory

2:57

speeds, all of which are zero sum games,

2:59

too. Since every single part of the AI

3:02

stack is currently supply constrained.

3:04

So the question isn't whether AI

3:06

spending will slow down, but where it

3:09

will shift to next. And three of the

3:11

biggest AI chip companies on Earth,

3:13

Nvidia, Cerabus, and AMD, all have very

3:16

different answers. Let's start with AMD.

3:18

AMD's answer is memory. An AI model is

3:21

made up of billions of numbers called

3:23

parameters. And the actual value of each

3:25

parameter comes from the patterns that

3:27

it learned during training. Those values

3:29

are called weights. And each weight

3:31

takes up about 16 bits or two bytes of

3:35

memory. So, a model with 400 billion

3:38

parameters needs around 800 GB of

3:40

memory. A single Nvidia H200 GPU has 141

3:45

GB of memory. So, you'd need about six

3:47

of them just to ask the model a single

3:49

question. Today, those weights live in

3:51

memory that sit outside of the

3:53

processor, and the processor spends most

3:56

of its time waiting for data to arrive.

3:58

That wait time is the bottleneck. On

4:01

August 6th, AMD announced plans to

4:03

acquire a company that fixes this by

4:06

etching the model directly into the

4:08

chip. The company is called Talis, a

4:10

chip startup founded in 2023, and it

4:13

etches AI model weights directly into

4:15

the metal layers of a chip, so they

4:17

don't have to be loaded from memory. The

4:19

benefits here are huge. Total throughput

4:22

goes way up and latency goes way down

4:24

when the processor stops having to wait

4:26

for data to come from memory. But the

4:28

trade-off is pretty huge, too. Each chip

4:31

is permanently dedicated to a single

4:33

specific AI model, exchanging

4:35

generalpurpose programmability for

4:38

maximum efficiency. If the acquisition

4:40

does go through, AMD will combine these

4:42

specialized chips with their own

4:44

instinct GPUs and rack scale setups like

4:47

Helios so that the GPUs can process

4:50

prompts and manage dynamic workloads.

4:52

While these new fixedweight chips handle

4:54

the memoryheavy decode phase of

4:56

inference, kind of like Nvidia's recent

4:58

partnership with Grock. But where

5:00

Grock's language processing units or

5:02

LPUs have 500 megabytes of ultraast

5:05

onchip memory, Talis's HC1 chips turn

5:09

the model directly into hardware. A

5:11

weight of one is a physically connected

5:13

wire and a zero is a broken path. So the

5:16

math is automatically executed as data

5:18

simply flows through the chip without

5:20

any traditional instruction cycles. And

5:23

speaking of data, I started getting way

5:25

more spam phone calls and texts right

5:27

around the end of the pandemic. And when

5:29

I asked my friends, they said the same

5:31

thing. So I decided to dig into it and

5:34

found that hundreds of online data

5:35

brokers were collecting and selling my

5:38

personal information. And they might be

5:40

selling yours, too. That's when I joined

5:42

Delete Me, the sponsor of this video,

5:44

and I've been with them ever since.

5:46

Delete Me is a hands-free subscription

5:48

service that will remove your personal

5:50

information from those online data

5:51

brokers. Every quarter, you get a

5:53

privacy report showing you what they've

5:55

done. I just got my 17th report, and

5:58

Delete Me reviewed over 60,000 listings

6:01

from me so far, and they removed my data

6:03

over 600 separate times. Not just my

6:06

name and address, but my wife's and my

6:09

families, too. But here's the best part.

6:11

They'll keep scanning these websites

6:13

even after they remove my information.

6:16

Talk about a total lifesaver. So, if you

6:18

care about your family's privacy and you

6:20

like saving money, you can get 20% off

6:23

any consumer plan with my code symbol 20

6:26

by going to joindeme.com/syol20

6:30

or with my link in the description. And

6:32

a big thank you to delete me for

6:33

supporting the channel and keeping my

6:35

family's data safe. All right, so

6:37

Talis's chip completely destroys the

6:39

memory wall, which is the big bottleneck

6:41

of moving data back and forth between a

6:43

processor and separate memory. Here's

6:45

what that means in practice. A typical

6:48

cloud AI provider can serve a model like

6:50

Llama 3.18B at around 140 tokens per

6:54

second per user. That same model can hit

6:57

17,000 tokens per second on Batalis

7:00

chip, making it around 120 times faster

7:04

than a traditional GPU. responses are

7:06

generated so fast that entire pages of

7:09

text appear instantly rather than

7:11

streaming in line by line. But here's

7:14

the big problem. If the chip is directly

7:16

tied to a single model and chips take

7:18

years to design, doesn't that mean that

7:21

the chip will be obsolete by the time it

7:23

leaves the production line? Well,

7:24

actually, Talus found a way to get

7:26

around that problem with a pretty clever

7:28

design. Instead of redesigning the whole

7:30

chip for every new model, Talus uses a

7:33

tiered manufacturing process that breaks

7:35

the chip into more than a 100 physical

7:37

layers of silicon and metal stacked on

7:39

top of each other. Everything underneath

7:41

the top two layers, the transistors, the

7:44

power distribution, the math units, and

7:46

the memory never has to change. Only the

7:49

two metal layers at the very top

7:51

translate a model specific weights into

7:53

a physical chip. And since 98% of the

7:56

chip stays the same between models, they

7:58

can be made ahead of time and stored

8:00

until a model is chosen for that chip.

8:02

Then once the two layers are designed

8:04

specifically for a new model, TSMC can

8:07

pull the base wafer off a shelf, add

8:09

those two layers, and ship the chip

8:10

within 60 days. This two-month

8:13

turnaround completely shifts the

8:15

economics around data centers. Instead

8:17

of buying generalpurpose GPUs or

8:19

spending many years and many more

8:21

billions of dollars designing custom

8:23

AS6, they can just buy a cluster of

8:25

chips hardwired for specific models, run

8:27

those chips for a year, and then

8:29

literally swap them out when the next

8:31

model drops. And not only is it up to

8:34

120 times faster than a traditional GPU,

8:36

it's also roughly 20 times cheaper since

8:39

it's made on older 6nanometer

8:41

technology. It has no high bandwidth

8:43

memory and it uses much simpler chip

8:46

packaging. Talis also says that their

8:48

chips cost less than a penny per million

8:50

tokens to run versus about 3.8 for an

8:54

Nvidia Blackwell running the same model.

8:56

So, when spending shifts from acquiring

8:58

land, buildings, and power, AMD is

9:01

betting that it'll shift into

9:02

hyperoptimized, lowcost chips that can

9:05

be easily swapped out as models keep

9:07

improving. But while AMD is cutting

9:09

costs, Cerebrris is cutting cables. A

9:12

modern AI cluster has thousands of

9:14

separate chips stitched together with

9:16

network cables and switches. Every time

9:19

data moves from one chip to another, it

9:21

costs time and power. And just like with

9:24

memory, chips can spend more time

9:26

waiting for the data to arrive over a

9:28

network than they spend processing it.

9:30

Think about what this actually means.

9:32

Chip makers spend tens of thousands of

9:34

dollars turning a big silicon wafer into

9:37

dozens of individual chips. Then they

9:40

spend billions of dollars connecting

9:41

those chips back together inside data

9:43

centers. Cerebras is betting their

9:45

entire company that this approach is

9:47

wrong. So they never cut the wafer in

9:49

the first place. Instead, they turn it

9:52

into one massive chip called the wafer

9:54

scale engine, or WSE for short. Their

9:57

current generation is the WSE3, and each

10:00

side is about 8 in. That's around 30

10:03

times more area than Nvidia's Blackwell

10:05

B200 chips. If Nvidia's chips are the

10:08

size of postage stamps, then Cerebras'

10:10

are the size of dinner plates. But, as

10:12

you know, size doesn't matter. It's all

10:15

about how you use it. Cerebras' chips

10:17

have 4 trillion transistors. That's 19

10:20

times more than Nvidia's B200's. But

10:23

since the chip is around 30 times

10:25

bigger, that means that Nvidia actually

10:27

packs over 50% more transistors into the

10:30

same area because their chips are made

10:32

on a more advanced process node by TSMC.

10:35

The wafer scale engine also has a

10:37

whopping 900,000 cores, four times more

10:40

than Blackwell, and also comes with a

10:42

whopping 44 GB of SRAMM, which is the

10:46

same kind of high-speed memory used in

10:48

the Gro LPUs, except, you know, 88 times

10:52

more of it. That memory can move data at

10:54

21 pabytes per second, which is about

10:57

2600 times the memory bandwidth of

11:00

Nvidia's Blackwell B200's. There are

11:02

over 18,000 titles on Netflix and their

11:05

uncompressed master archive is about

11:08

four pabytes. That means this Cerebras

11:10

chip can move Netflix's entire library

11:13

between cores five times every second.

11:16

That's a huge deal for AI inference

11:18

performance. And it's all because Nvidia

11:20

has to move their data between chips,

11:22

across cables, and through switches. All

11:25

of which add extra time to every

11:27

transfer. While Cerebrris simply moves

11:29

data across one massive chip. No hops,

11:32

no cables, only compute. Cerebras just

11:35

reported earnings for the first time

11:36

since going public and the numbers speak

11:39

for themselves. The remaining

11:40

performance obligations hit $25.4

11:43

billion. That's basically their backlog.

11:46

Cerebras is guiding for $880 to $890

11:49

million in revenue this year. So that

11:52

backlog is worth about 29 times

11:54

everything they expect to sell in 2026.

11:57

They expect to deliver on about 22% of

12:00

that backlog in the next 2 years and

12:02

another 43% in the 2 years after that.

12:05

That first chunk works out to about $5.6

12:08

billion by the middle of 2028 or roughly

12:11

$2.8 billion per year, which means their

12:14

revenue should roughly triple. Revenue

12:16

from hardware came in at $82 million,

12:19

which is up 17% from last year. On the

12:22

official books, hardware sales look like

12:24

they fell by 23%. But that's only

12:27

because of a $28 million charge for

12:29

stock warrants that Cerebras handed to

12:31

Open AAI. If you remove that charge,

12:34

their hardware sales grew. This is

12:36

exactly why I ignore headlines and dig

12:38

into the numbers myself. All right,

12:41

hardware sales are up, but only by 17%.

12:44

On the other hand, revenues from cloud

12:46

and services hit $127 million, which is

12:50

up $287%

12:52

year-over-year. So, the real money for

12:55

Cerebras is in selling access to their

12:57

machines, not selling the machines

12:59

themselves. And their gross margins

13:02

fell, but mostly because they can't

13:04

build capacity fast enough to serve all

13:06

the demand for their chips. So, they

13:08

actually end up renting their own

13:09

systems back from the cloud companies

13:11

that bought them to fill the gap. So,

13:14

while AMD and Talus are betting on model

13:16

specific chips, Cerebrris is betting on

13:19

wafer sized ones. But we can't talk

13:21

about the future of AI chips without

13:23

talking about Nvidia. And if you feel

13:26

I've earned it, consider hitting the

13:27

like button and subscribing to the

13:29

channel. That really helps and it lets

13:31

me know to make more content like this.

13:33

Thanks. Now, let's talk about Nvidia.

13:36

The question isn't whether AI spending

13:38

will slow down, it's where it will shift

13:40

to next. Nvidia's answer is to the chips

13:43

themselves as investable assets. Think

13:46

about how you'd build a skyscraper.

13:48

Nobody is paying cash. A developer

13:50

borrows most of the money because the

13:52

building holds its value. And the lender

13:54

knows that it can be sold if something

13:56

goes sideways. That's all an investable

13:58

asset is something you can borrow

14:00

against because everyone agrees it'll

14:03

still be worth something later. AI data

14:05

centers mostly get paid for out of a

14:07

company's free cash flows, which as I

14:09

just said earlier is starting to run

14:11

out. So Nvidia is trying to change where

14:13

that money comes from altogether. On

14:16

August 10th, Nvidia announced deals with

14:18

Apollo, Black Rockck, Blackstone,

14:20

Brookfield, Goldman Sachs, and KKR to

14:23

build financing platforms for over half

14:26

a trillion dollars worth of AI

14:27

infrastructure. None of that money is

14:29

Nvidia's. It's pension funds, sovereign

14:32

wealth funds, insurers, the slowest and

14:34

most conservative money on the planet.

14:36

Jensen Hong's big idea is that a GPU

14:39

makes good collateral because someone

14:41

else will always want it. And Nvidia's

14:43

software updates extend its useful

14:45

lifespan over time. For example, Nvidia

14:48

came out with an open- source software

14:49

package called Tensor RT LLM, which

14:52

doubled the inference performance of

14:54

large language models running on H100s.

14:57

That means every H100 already sitting in

14:59

a data center got roughly twice as good

15:02

at running large language models

15:03

overnight for free. And that software

15:06

also works on Nvidia's older Ampear and

15:08

Ada Loveace chips, too. And all the way

15:11

up through Blackwell. The part that I'm

15:13

not so sure about is what these older

15:15

chips will actually sell for on the

15:17

secondhand market. Cars lose a lot of

15:19

value based on their age and how they

15:21

compare to newer models, not just on how

15:23

fast they drive or their miles per

15:25

gallon. AI chips might depreciate the

15:28

same way. We just don't know yet. There

15:30

are also three big catches to Nvidia's

15:32

deals. First, none of them are binding,

15:35

at least not yet. Nobody has actually

15:37

committed a single dollar. Second,

15:39

Nvidia could quietly be on the hook for

15:41

the difference. An Nvidia blog post

15:43

published the next day said that Nvidia

15:46

may cover up to 25% of the losses if the

15:49

equipment ends up being worth less than

15:51

a loan assumed. And third, the loans can

15:54

last longer than the contracts to pay

15:56

them off. The same day as the

15:57

announcement, Coreeave closed a $2.6

16:00

billion loan against its GPUs. The loan

16:03

runs for about 5 years, while the

16:05

customer contracts renting those same

16:07

GPUs only last for three. So, if those

16:10

GPUs don't get rented again 3 years from

16:12

now, even though they'll be much older,

16:14

that could cause some serious issues

16:16

with paying back the 5-year loan. So,

16:19

the big question the market needs to

16:20

answer is what is the actual useful

16:23

lifespan of a GPU? Is it 3 years? Is it

16:26

five? That determines whether they're

16:28

worth investing in as an asset class at

16:31

all. And we might get an answer on

16:33

August 26th when Nvidia reports their

16:35

earnings. Either way, the biggest case

16:38

against AI has always been that spending

16:40

will eventually slow down. And when it

16:42

does, every stock that lives or dies by

16:44

that spending will go down with it. But

16:47

Amazon and Google both went free cash

16:49

flow negative and raised capital to

16:52

spend even more. So the question isn't

16:54

whether AI spending will slow down, but

16:57

where it shifts to next. And three of

16:59

the biggest AI chip companies all have

17:01

different answers. AMD is betting it'll

17:04

shift to hyperoptimized lowcost chips

17:06

that can easily be swapped out as models

17:08

keep improving. Cerebras is betting on

17:10

waferiz chips that cut out rack level

17:13

network cables and switches altogether.

17:15

And Nvidia is betting that GPUs will

17:17

turn into an asset class of their own

17:19

because somebody will always want to

17:21

rent them. Either way, I expect AI

17:24

spending to keep speeding up until the

17:26

world runs out of land and power for

17:28

data centers. And after that, spending

17:30

will just shift to filling them. Money

17:33

is no longer the constraint here. That's

17:35

why I think investing in AI is still a

17:37

great way to get rich without getting

17:39

lucky. And if you want to see even more

17:42

stocks I'm buying to get rich without

17:43

getting lucky, check out this video

17:45

next. Either way, thanks for watching

17:47

and until next time, this is Tickerol U.

17:50

My name is Alex, reminding you that the

17:53

best investment you can make is in you.

Interactive Summary

This video challenges the common bearish view that AI spending is a bubble set to crash. Instead, it argues that hyperscalers like Google, Meta, and Amazon are aggressively increasing their capital expenditures to secure critical AI infrastructure—a zero-sum game regarding land and power. The video explores how three major chip companies are positioning themselves for this sustained spending: AMD, by developing hyper-optimized, swap-able chips to bypass memory bottlenecks; Cerebras, by using massive 'wafer-scale' engines to eliminate network latency between chips; and Nvidia, by attempting to transform GPUs into a formal, financeable asset class.

Suggested questions

4 ready-made prompts