I'm Buying Every Share I Can (Investors Aren't Ready)
456 segments
The single biggest argument against AI
stocks just died. Wall Street analysts
have been saying the same thing for
years now. AI is a bubble and the entire
tech sector will crash as soon as
spending slows down even a little. But
three major developments just put an end
to that argument and kicked off what
could be the next phase of the AI era.
My name is Alex and I spent 8 years as
an electrical engineer and AI researcher
at MIT. And I've never seen chip
companies move this fast. Let me show
you what's going on and how I'm
investing in it. Your time is valuable,
so let's get right into it. The biggest
case against AI has always been that
spending will eventually slow down, and
when it does, every stock that lives and
dies by that spending will go down with
it. It's only a matter of time.
Hyperscalers like Google, Amazon,
Microsoft, and Meta Platforms report
their capex budgets every quarter, and
every quarter, it's been growing like
crazy. In fact, their AI spending has
been compounding at over 60% per year
since 2023, and it's actually expected
to speed up. These four companies are
currently on course to spend over $700
billion on data centers this year alone,
compared to the $375 billion they spent
in 2025. That's a 90% increase
year-over-year. And it's coming at a
huge cost. Amazon's free cash flow went
from $18 billion to negative7.6 billion
over the last 12 months. But just a few
weeks ago, Amazon increased their capex
budget for 2026 from $200 billion to
$220 billion due to the price of memory
chips. That same week, Alphabet added
$15 billion to its capex budget while
reporting that their free cash flow
turned negative for the first time since
Google went public in 2004. and Meta's
free cash flows are down by 91%. You'd
think that spending would slow down once
the biggest companies driving it burn
through all their cash flows and then
some, but instead they're borrowing
money and issuing new shares. Amazon
sold $37 billion worth of bonds in
quarter 1. Alphabet issued over $65
billion worth of bonds so far this year
and announced an $85 billion equity
raise in June, the biggest raise in
American corporate history. debt covered
9% of hyperscaler capex in 2024. Today,
it covers 32%. It's important for
investors to understand why spending
isn't slowing down and why these
companies keep investing in AI at all
costs. Data center infrastructure is a
zero- sum game. There's only so much
land and grid connected power in the
first place, and every acre and gawatt
you don't get is one that your
competition does. A big new data center
can wait anywhere from 2 to 6 years for
a grid connection. And even when it's
all bought up, spending will just shift
to the next big bottleneck. Cooling,
compute density, network, and memory
speeds, all of which are zero sum games,
too. Since every single part of the AI
stack is currently supply constrained.
So the question isn't whether AI
spending will slow down, but where it
will shift to next. And three of the
biggest AI chip companies on Earth,
Nvidia, Cerabus, and AMD, all have very
different answers. Let's start with AMD.
AMD's answer is memory. An AI model is
made up of billions of numbers called
parameters. And the actual value of each
parameter comes from the patterns that
it learned during training. Those values
are called weights. And each weight
takes up about 16 bits or two bytes of
memory. So, a model with 400 billion
parameters needs around 800 GB of
memory. A single Nvidia H200 GPU has 141
GB of memory. So, you'd need about six
of them just to ask the model a single
question. Today, those weights live in
memory that sit outside of the
processor, and the processor spends most
of its time waiting for data to arrive.
That wait time is the bottleneck. On
August 6th, AMD announced plans to
acquire a company that fixes this by
etching the model directly into the
chip. The company is called Talis, a
chip startup founded in 2023, and it
etches AI model weights directly into
the metal layers of a chip, so they
don't have to be loaded from memory. The
benefits here are huge. Total throughput
goes way up and latency goes way down
when the processor stops having to wait
for data to come from memory. But the
trade-off is pretty huge, too. Each chip
is permanently dedicated to a single
specific AI model, exchanging
generalpurpose programmability for
maximum efficiency. If the acquisition
does go through, AMD will combine these
specialized chips with their own
instinct GPUs and rack scale setups like
Helios so that the GPUs can process
prompts and manage dynamic workloads.
While these new fixedweight chips handle
the memoryheavy decode phase of
inference, kind of like Nvidia's recent
partnership with Grock. But where
Grock's language processing units or
LPUs have 500 megabytes of ultraast
onchip memory, Talis's HC1 chips turn
the model directly into hardware. A
weight of one is a physically connected
wire and a zero is a broken path. So the
math is automatically executed as data
simply flows through the chip without
any traditional instruction cycles. And
speaking of data, I started getting way
more spam phone calls and texts right
around the end of the pandemic. And when
I asked my friends, they said the same
thing. So I decided to dig into it and
found that hundreds of online data
brokers were collecting and selling my
personal information. And they might be
selling yours, too. That's when I joined
Delete Me, the sponsor of this video,
and I've been with them ever since.
Delete Me is a hands-free subscription
service that will remove your personal
information from those online data
brokers. Every quarter, you get a
privacy report showing you what they've
done. I just got my 17th report, and
Delete Me reviewed over 60,000 listings
from me so far, and they removed my data
over 600 separate times. Not just my
name and address, but my wife's and my
families, too. But here's the best part.
They'll keep scanning these websites
even after they remove my information.
Talk about a total lifesaver. So, if you
care about your family's privacy and you
like saving money, you can get 20% off
any consumer plan with my code symbol 20
by going to joindeme.com/syol20
or with my link in the description. And
a big thank you to delete me for
supporting the channel and keeping my
family's data safe. All right, so
Talis's chip completely destroys the
memory wall, which is the big bottleneck
of moving data back and forth between a
processor and separate memory. Here's
what that means in practice. A typical
cloud AI provider can serve a model like
Llama 3.18B at around 140 tokens per
second per user. That same model can hit
17,000 tokens per second on Batalis
chip, making it around 120 times faster
than a traditional GPU. responses are
generated so fast that entire pages of
text appear instantly rather than
streaming in line by line. But here's
the big problem. If the chip is directly
tied to a single model and chips take
years to design, doesn't that mean that
the chip will be obsolete by the time it
leaves the production line? Well,
actually, Talus found a way to get
around that problem with a pretty clever
design. Instead of redesigning the whole
chip for every new model, Talus uses a
tiered manufacturing process that breaks
the chip into more than a 100 physical
layers of silicon and metal stacked on
top of each other. Everything underneath
the top two layers, the transistors, the
power distribution, the math units, and
the memory never has to change. Only the
two metal layers at the very top
translate a model specific weights into
a physical chip. And since 98% of the
chip stays the same between models, they
can be made ahead of time and stored
until a model is chosen for that chip.
Then once the two layers are designed
specifically for a new model, TSMC can
pull the base wafer off a shelf, add
those two layers, and ship the chip
within 60 days. This two-month
turnaround completely shifts the
economics around data centers. Instead
of buying generalpurpose GPUs or
spending many years and many more
billions of dollars designing custom
AS6, they can just buy a cluster of
chips hardwired for specific models, run
those chips for a year, and then
literally swap them out when the next
model drops. And not only is it up to
120 times faster than a traditional GPU,
it's also roughly 20 times cheaper since
it's made on older 6nanometer
technology. It has no high bandwidth
memory and it uses much simpler chip
packaging. Talis also says that their
chips cost less than a penny per million
tokens to run versus about 3.8 for an
Nvidia Blackwell running the same model.
So, when spending shifts from acquiring
land, buildings, and power, AMD is
betting that it'll shift into
hyperoptimized, lowcost chips that can
be easily swapped out as models keep
improving. But while AMD is cutting
costs, Cerebrris is cutting cables. A
modern AI cluster has thousands of
separate chips stitched together with
network cables and switches. Every time
data moves from one chip to another, it
costs time and power. And just like with
memory, chips can spend more time
waiting for the data to arrive over a
network than they spend processing it.
Think about what this actually means.
Chip makers spend tens of thousands of
dollars turning a big silicon wafer into
dozens of individual chips. Then they
spend billions of dollars connecting
those chips back together inside data
centers. Cerebras is betting their
entire company that this approach is
wrong. So they never cut the wafer in
the first place. Instead, they turn it
into one massive chip called the wafer
scale engine, or WSE for short. Their
current generation is the WSE3, and each
side is about 8 in. That's around 30
times more area than Nvidia's Blackwell
B200 chips. If Nvidia's chips are the
size of postage stamps, then Cerebras'
are the size of dinner plates. But, as
you know, size doesn't matter. It's all
about how you use it. Cerebras' chips
have 4 trillion transistors. That's 19
times more than Nvidia's B200's. But
since the chip is around 30 times
bigger, that means that Nvidia actually
packs over 50% more transistors into the
same area because their chips are made
on a more advanced process node by TSMC.
The wafer scale engine also has a
whopping 900,000 cores, four times more
than Blackwell, and also comes with a
whopping 44 GB of SRAMM, which is the
same kind of high-speed memory used in
the Gro LPUs, except, you know, 88 times
more of it. That memory can move data at
21 pabytes per second, which is about
2600 times the memory bandwidth of
Nvidia's Blackwell B200's. There are
over 18,000 titles on Netflix and their
uncompressed master archive is about
four pabytes. That means this Cerebras
chip can move Netflix's entire library
between cores five times every second.
That's a huge deal for AI inference
performance. And it's all because Nvidia
has to move their data between chips,
across cables, and through switches. All
of which add extra time to every
transfer. While Cerebrris simply moves
data across one massive chip. No hops,
no cables, only compute. Cerebras just
reported earnings for the first time
since going public and the numbers speak
for themselves. The remaining
performance obligations hit $25.4
billion. That's basically their backlog.
Cerebras is guiding for $880 to $890
million in revenue this year. So that
backlog is worth about 29 times
everything they expect to sell in 2026.
They expect to deliver on about 22% of
that backlog in the next 2 years and
another 43% in the 2 years after that.
That first chunk works out to about $5.6
billion by the middle of 2028 or roughly
$2.8 billion per year, which means their
revenue should roughly triple. Revenue
from hardware came in at $82 million,
which is up 17% from last year. On the
official books, hardware sales look like
they fell by 23%. But that's only
because of a $28 million charge for
stock warrants that Cerebras handed to
Open AAI. If you remove that charge,
their hardware sales grew. This is
exactly why I ignore headlines and dig
into the numbers myself. All right,
hardware sales are up, but only by 17%.
On the other hand, revenues from cloud
and services hit $127 million, which is
up $287%
year-over-year. So, the real money for
Cerebras is in selling access to their
machines, not selling the machines
themselves. And their gross margins
fell, but mostly because they can't
build capacity fast enough to serve all
the demand for their chips. So, they
actually end up renting their own
systems back from the cloud companies
that bought them to fill the gap. So,
while AMD and Talus are betting on model
specific chips, Cerebrris is betting on
wafer sized ones. But we can't talk
about the future of AI chips without
talking about Nvidia. And if you feel
I've earned it, consider hitting the
like button and subscribing to the
channel. That really helps and it lets
me know to make more content like this.
Thanks. Now, let's talk about Nvidia.
The question isn't whether AI spending
will slow down, it's where it will shift
to next. Nvidia's answer is to the chips
themselves as investable assets. Think
about how you'd build a skyscraper.
Nobody is paying cash. A developer
borrows most of the money because the
building holds its value. And the lender
knows that it can be sold if something
goes sideways. That's all an investable
asset is something you can borrow
against because everyone agrees it'll
still be worth something later. AI data
centers mostly get paid for out of a
company's free cash flows, which as I
just said earlier is starting to run
out. So Nvidia is trying to change where
that money comes from altogether. On
August 10th, Nvidia announced deals with
Apollo, Black Rockck, Blackstone,
Brookfield, Goldman Sachs, and KKR to
build financing platforms for over half
a trillion dollars worth of AI
infrastructure. None of that money is
Nvidia's. It's pension funds, sovereign
wealth funds, insurers, the slowest and
most conservative money on the planet.
Jensen Hong's big idea is that a GPU
makes good collateral because someone
else will always want it. And Nvidia's
software updates extend its useful
lifespan over time. For example, Nvidia
came out with an open- source software
package called Tensor RT LLM, which
doubled the inference performance of
large language models running on H100s.
That means every H100 already sitting in
a data center got roughly twice as good
at running large language models
overnight for free. And that software
also works on Nvidia's older Ampear and
Ada Loveace chips, too. And all the way
up through Blackwell. The part that I'm
not so sure about is what these older
chips will actually sell for on the
secondhand market. Cars lose a lot of
value based on their age and how they
compare to newer models, not just on how
fast they drive or their miles per
gallon. AI chips might depreciate the
same way. We just don't know yet. There
are also three big catches to Nvidia's
deals. First, none of them are binding,
at least not yet. Nobody has actually
committed a single dollar. Second,
Nvidia could quietly be on the hook for
the difference. An Nvidia blog post
published the next day said that Nvidia
may cover up to 25% of the losses if the
equipment ends up being worth less than
a loan assumed. And third, the loans can
last longer than the contracts to pay
them off. The same day as the
announcement, Coreeave closed a $2.6
billion loan against its GPUs. The loan
runs for about 5 years, while the
customer contracts renting those same
GPUs only last for three. So, if those
GPUs don't get rented again 3 years from
now, even though they'll be much older,
that could cause some serious issues
with paying back the 5-year loan. So,
the big question the market needs to
answer is what is the actual useful
lifespan of a GPU? Is it 3 years? Is it
five? That determines whether they're
worth investing in as an asset class at
all. And we might get an answer on
August 26th when Nvidia reports their
earnings. Either way, the biggest case
against AI has always been that spending
will eventually slow down. And when it
does, every stock that lives or dies by
that spending will go down with it. But
Amazon and Google both went free cash
flow negative and raised capital to
spend even more. So the question isn't
whether AI spending will slow down, but
where it shifts to next. And three of
the biggest AI chip companies all have
different answers. AMD is betting it'll
shift to hyperoptimized lowcost chips
that can easily be swapped out as models
keep improving. Cerebras is betting on
waferiz chips that cut out rack level
network cables and switches altogether.
And Nvidia is betting that GPUs will
turn into an asset class of their own
because somebody will always want to
rent them. Either way, I expect AI
spending to keep speeding up until the
world runs out of land and power for
data centers. And after that, spending
will just shift to filling them. Money
is no longer the constraint here. That's
why I think investing in AI is still a
great way to get rich without getting
lucky. And if you want to see even more
stocks I'm buying to get rich without
getting lucky, check out this video
next. Either way, thanks for watching
and until next time, this is Tickerol U.
My name is Alex, reminding you that the
best investment you can make is in you.
Ask follow-up questions or revisit key timestamps.
This video challenges the common bearish view that AI spending is a bubble set to crash. Instead, it argues that hyperscalers like Google, Meta, and Amazon are aggressively increasing their capital expenditures to secure critical AI infrastructure—a zero-sum game regarding land and power. The video explores how three major chip companies are positioning themselves for this sustained spending: AMD, by developing hyper-optimized, swap-able chips to bypass memory bottlenecks; Cerebras, by using massive 'wafer-scale' engines to eliminate network latency between chips; and Nvidia, by attempting to transform GPUs into a formal, financeable asset class.
Videos recently processed by our community