Ep. 013 - AWS Margins Jump 10% While Azure and GCP Flatline (Tokenomics)
1387 segments
Hello everyone, welcome back to
SemiAnalysis
weekly, episode number 13, lucky number
13. We're here with Joey, Jeremy, and
Kristen. We're going to talk about an
article that we put out recently
uh called Anthropic growth and bedrock
mix drive AWS margins higher while peers
lag. That means we're going to talk
about everything Anthropic, including
the uh recent announcement of their
series H, the release of Opus 4.8, but
with the spin on focusing on
infrastructure, like how exactly they
serve these tokens, especially with
their partnership with AWS. Guys, uh
welcome to the show. Excited to talk
through this.
>> Thanks, Jordan.
>> Man.
>> All right, so let's uh let's dig in with
the article itself and and talk a little
bit about the backstory. Like, I think
uh a lot of people understand the
concept of tokens, understand what GPUs
are, um but and and certainly like cloud
providers and neo clouds, but not
everybody is getting their tokens from
the same place. So, can one of you guys
give me a backstory on bedrock? Like,
what is AWS? Uh what are they, you know,
in terms of what they're doing for
Anthropic and how they're serving tokens
with bedrock? Joey, we'll start with
you.
>> Yeah, yeah, yeah. Yeah, I'll go through
that. And I'll do it kind of for all the
clouds as well. Um so so if we really
look across all the hyperscalers, and
especially the big three, um when we
look at Amazon, Microsoft, and Google,
there's kind of two big breakouts, uh
maybe even three. There's a bit on the
software side. So, if you look at
Microsoft like GitHub, um Copilot, you
know, that's an area that's like an AI,
you know, software-as-a-service type
product. They have AI, you know,
infrastructure-as-a-service where
they're just renting out um these
accelerator chips. Um and then they kind
of have this token-as-a-service
business, you know, which is where
they'll, you know, essentially expose
these outside models or their own own
models um
to consumers to interact with.
Uh so it's a little bit of a different
business uh because instead of just
renting the underlying chip, you're
you're renting the underlying model, you
know, buying it through your cloud, you
know, your CSP account and kind of that
enterprise, you know, spending
agreement, you know, you have all the
same benefits around security,
availability zones. Um and able to buy,
you know, third-party models, you know,
and also some of these first-party
models, you know, through your cloud
provider of choice.
Um so at a high level,
um those are kind of the three big, you
know, buckets right now at the large
hyperscalers that we see and we're
seeing some big, you know, changes or
differences between them and how they've
gone about it strategically. Um and then
the two main things was Anthropic's
growth um really led, you know,
Anthropic's growth and then uh Amazon's,
you know, strategy around Bedrock has
really driven their margins higher
recently. Um
was really the takeaway from the
article. And this token as a as a
service is obviously much better
business for the hyperscalers than
infrastructure as a service, you know,
in the AI side.
>> Yeah, and the if you look at the big
three, they're kind of going in opposite
directions right now. Um just in terms
of their operating margin, right? That
was the biggest chart from the article.
I'll put it up on screen right now. But
Crystal, can you maybe explain like when
we look at AWS versus Google and versus
Microsoft and you try and break out the
cloud business, like is this all to
blame on Anthropic from our perspective
or like what's driving AWS to improve
their operating margins while
Microsoft's is declining and Google's
kind of stays flat here?
>> So I feel like a lot of it is that cloud
usage, right? Like a lot of people are
using cloud more and it's mostly routing
through AWS. And then like what Joey was
saying, the token as a service um
business model, they just have a better
margin on that as opposed to like the
infrastructure as a service model.
>> Makes sense, yeah.
>> Yeah, so I I I I think like if you if
you take a step back, like the the world
of clouds has really changed a lot in
2023 as you as you started seeing these
neo clouds, these GPU as a service
businesses because uh in the old days,
which is basically 2022 and before,
um
cloud cloud service providers, you
basically had three. Right, Amazon,
Google Cloud, Microsoft Azure, they had
amazing margins, easy returns on on
capital. Some new entrants like Oracle
were trying to get in, but right the
market was dominated by three players
that had an amazing business.
And now you get to GPU as a service and
what we first started to realize is that
the barriers to entry are much lower.
So, Jordan, you're probably, you know,
the best person to talk about this. You
probably know the the CEO of 150 Neo
Clouds or 200 Neo Clouds.
How many [laughter] now? 300? And anyway
>> over 200. Yeah, yeah.
>> [laughter]
>> It is pretty insane.
Um, now obviously some are better than
than others, but the point is the the
market has much lower barriers to entry.
Uh, the moat that uh, cloud used to be
doesn't really exist in the AI era
because it's really more about
infrastructure and especially the end
users, the whole point of cloud
computing was you make the IT much
easier. Um, folks don't need to have
such a big IT department internally.
They can just rent through the cloud.
Uh, it's super easy. Everything is well.
Uh, now the big end users, they want to
have much more control. Uh, you shift
from uh, platform as a service in the
old days to now bare metal. Uh, folks
want just just the metal. Uh, Open AI
wants to think that the way they like
it, Microsoft, Meta. Um, and so this is
basically the token as a service
business is basically the first case uh,
at scale where you see a an AI cloud
provider having a a business with a
different profile. Uh, right, because
uh, obviously as you have less of a moat
in the GPU as a service era, margins go
down. Uh, Oracle, best example, they
have an RPO of half a trillion dollars,
uh, which is a few quarters ago that was
as of as of Q1 is still bigger than
Amazon's.
>> [laughter]
>> Their backlog is bigger than Amazon's,
but no one gives them the credit because
people know it's much more risky, the
returns are not the same. Um,
and we've seen empirically that every
every single GPU as a service cloud has
faced struggles when they started to
ramp up their business. But our view as
a firm is we actually think that GPU as
a service business model is sound
business. Companies like Coreweave have
a sound business, but there's this lag
effect where as you ramp up and bring
more capacity online, there's these lags
that make your margins go down
temporarily because your asset base
depreciates because you have to pay data
center leases and so on and so forth.
And so it's really interesting to see
that Amazon is basically the first cloud
provider that in a time of unprecedented
capacity expansion,
over a gigawatt per quarter now,
they expand margins.
And that kind of tells you that if at a
period of accelerated capacity delivery
they can expand margins, you start to
think of okay, what's the stabilized
margin of this business?
And that's one of the points that we
make is the stabilized margins of token
as a service for Amazon are actually
extremely rich.
The return on capital is fundamentally
much more sound.
>> Yeah, let me throw this chart
up on screen actually from the from the
article which is the percent of revenue
that is going to Bedrock. Bedrock being
the token as a service at Amazon, right?
You can obviously see it ramp up quite a
bit at roughly the the time that we're
in right now, right? Q4 of last year and
first quarter of this year. And then if
we compare that to the chart I had on
previously where we're seeing their
operating margins improve in the first
quarter of this year, they're saying
that's due to it's Can you explain maybe
in more detail why it's unprecedented to
say that Amazon can bring on a gigawatt
per quarter and still improve operating
margins?
>> So
to to to understand this, you basically
have to go back to like
why are the pure bare metal providers
seeing their margins go down? So you
look at Coreweave, which is a pure play,
so they're the the cleanest example.
Oracle actually is kind of kind of the
same. When you look at this chart, it
essentially what this tells you is that
their stabilized business
does something like you know 25%
operating margins 30 to 40% gross
margins but the the whole issue is
stabilized stabilized means that your
GPU cluster is fully functional
you're getting the the monthly rent or
whatever rent from your customer
take or pay so it's a flat fee you know
exactly how much revenue you're going to
make often times it's a 5-year take or
pay contract again fixed rate
so you know your revenue you know your
cost everything is stable but before
getting there
the obviously have to set up the data
center which is a huge capital expense
up front
and then you have this whole process or
if you have the data center built but
you need to fill it with equipment that
takes a few months with some new stock
types of equipment like GB200 which is
super complicated we've seen that lag
get longer and longer which means that
this period of time where you depreciate
your assets you pay data center rent you
pay some labor you pay a whole bunch of
stuff that period of time goes up and
you don't make any revenue because your
cluster is not yet rolled out right and
so that's the whole dilemma that these
guys are facing is they know their
business model is stabilized structure
but they're they've been facing
challenges to wrap it up some a bit
contractual again GB200 some just more
structural because there is this lag
and Amazon is the same thing right like
like everyone else they have the same
profile they're bringing on a whole lot
of data centers
a whole lot of XPUs these XPUs in theory
they should take time to bring online
yet despite this you're seeing their
margins go up which you know pretty good
sign for them
>> Yeah but it and it's clearly different
than the others
in this space where the percentage of
their total AI revenue that they're
reporting
as being from token as a service as
opposed to other products is much higher
than Google and Azure like they are not
just bringing on capacity they're
successfully selling it into
the labs that are using it to serve
tokens, right? For these models.
Um
>> Yeah, and I guess the the chart you you
showed earlier, like what I forgot to
mention is obviously the margin buffer
that you have uh when you when you sell
tokens with your infrastructure as
opposed to having a 5-year take or pay
fixed contract with a capped upside.
Right? So, that's kind of the key one of
the key points of of of the article.
>> Yeah, yeah.
Guys, can you talk a little bit about
the workload mix as well? Like obviously
um any provider could conceptually do
this, but not everybody is doing it suc-
it's not like Azure doesn't have a token
as a service business or Google. In
fact, even Crusoe, CoreWeave, Nabia are
all trying to get into this business,
too, but they really they need a
customer and they need a customer
serving the right type of workload for
it to really result in a bunch of
growth, I would say.
>> Yeah, I I I would say, you know, for
Amazon specifically, you know, they
benefit from the biggest customer base.
Yeah, they've been doing this for
20-plus years now. At AWS,
um people are very comfortable buying
that, you know, through them. You know,
people even buy infrastructure software,
you know, through them, a provider like
MongoDB, like Snowflake. Um pretty large
marketplace business. So, customers are
really comfortable with the security,
uh you know, at this point um buying
through them, having a single bill for
all this. Um and I and I think this
comes back to you some of the Anthropic
news today, you also on the ARR number
of 47 billion. You know, we look at Q1
and even in the Q2 here,
you know, a lot of the Anthropic's
um you know, the business mix is much
much different. So, Amazon is benefiting
from picking, you know, being 80 to 90%
cloud in Bedrock, you know, versus
Microsoft is is heavily OpenAI,
obviously.
Um you know, Google has a lot of Gemini,
which doesn't benefit as much from a lot
of these agentic coding tasks.
Um
and when we look at kind of the OpenAI
or you know, and the coding percentage
that's really driving Anthropic, um you
know, you're talking probably like 10
billion a month net new ARR in the last,
you know, 3 months, you know, of March,
April, and May
um, together. And, you know, a lot of
that, probably 80% of that uh of
Anthropic's net new ARR is in this API
business. You know, if we go to, you
know, OpenAI, you know, 60% of that
business in Q1 was was consumer,
consumer subscriptions, really. Um, so
um, there's kind of a mix of factors,
but but like Amazon was in the right
place at the right time
uh with Anthropic. And then two, you
know, they were able to give with
Trainium 2, I think, pretty you know,
interesting deal structure for both
parties. And I think that was the big
thing, too. We we kind of mentioned in
the article.
Um, obviously, like their mix of
Trainium, you know, they could there's
an infrastructure as a as a service like
fee component that Anthropic pays for
this infrastructure, like Bedrock
infrastructure. Um, but then there's
some interesting like hurdles around
revenue share, you know, and then and
really margin share that happened. And
because, you know, I think Anthropic was
probably at, you know, 25 million of ARR
per
um, you know, mid-one off the top of my
head in Q1 versus like probably six back
in Q2 four, like, you know, the numbers
really made sense for both parties. And
they really, you know, both benefited.
>> So, yeah. Eric, uh Crystal, can you talk
a little bit about those forecasts you
guys were making like going into this
end of Q1 and forecast for Q2? You don't
necessarily get disclosures the same way
from Anthropic, but we've at this point
kind of been bang on with the
disclosures with the disclosed revenue
figures and the margin figures, right?
>> Mhm.
>> [clears throat]
>> I think for
us, it's a little easier to forecast
Anthropic than it was OpenAI just
because so much of Anthropic um, ARR
comes from like the API side. And
because we also use Anthropic and
there's a lot of data out there about
like how people are using all of the
different uh Claude models, like, it's
so much easier to predict the workflow
and um, kind of see what the token
consumption is going to look like and
they've kept token pricing relatively
stable these past two releases more or
less, right? Whereas like OpenAI
all a lot of their revenue comes from
like these subscriptions and you never
know like if consumers are going to
switch over to another one, so it's a
lot harder to quantify the number of
users that are using the subscription
when they can just cancel anytime versus
for like API.
>> Yeah. Can you explain a little bit about
the
release on 4.8 and 4.7 cuz
pricing's remained the same but fast
mode has changed and maybe there's some
other there's been some other changes in
terms of how they're doing pricing on
the API.
>> Yeah, they said pricing is the same,
right? For regular but then for fast
mode it's different. Another cool thing
that they said apparently was that it
doesn't hallucinate as much and they had
like a pretty cool bar chart showing
that it's like it's rate of like
hallucination is a lot lower for 4.8 and
4.7 supposedly. I haven't tested it out
yet and they said that it's super close
to like mid those previews, so hopefully
it'll stop just, you know, pulling
numbers out of thin air from when we're
using it for our analyses.
>> That'd be good. That'd be good if
numbers weren't pulled out of thin air.
Yeah.
>> Yeah.
>> Um
in terms of the the fast mode and
consumer subscriptions, you have
comments there on like
I don't know what we've learned over the
past few weeks or months from digging
into how Anthropic is running their
business on AWS and like what's uh
maybe what's the biggest percentage of
their
revenue or their margin contribution
across those different mixes like those
different types of workflows that people
could be consuming tokens on the API
for?
>> Yeah, I mean we don't like we have that
in the tokenomics model like and we're
doing a big study right now on on what
percentage geo coding market you know is
of of current you know token spend
especially on the API side.
Um you know, we've done a lot of work
too on just the consumer and to some on
the B2B data of where Anthropic has has
just been taking a ton of share of net
new customers year to date both in
consumer subscriptions and enterprise.
Um but I mean it's really difficult and
a big question among a lot of our
clients like how big is the the coding
market currently. Um, I think Anthropic
recently said at their financial
services day that financial services was
the second biggest vertical.
Um,
and we kind of know there's a pretty big
gap between coding and and financial
services in terms of especially in API
spend.
Just of how people, you know, use this,
you know, anecdotally. Um,
so it's, you know, and you know, too,
there's been a lot of like recent
comments too on token maxing it and
especially the Fortune 500s like how do
you budget for that? How do you blow
through that spend over time? How do you
measure ROI?
Um, so things will have to change.
You know, I think and then people will
put some processes in place, but I mean
I know guys, you know, like Jeremy get
insane ROI at the data center model and
his team and and him
give with that and and you know, maybe
there's less policing out there at some
other you organizations that just let
people go crazy.
Um,
but that's, you know, too, another
contributing factor I think what was
just, you know, the success of coding
and then too, Anthropic like, you know,
starting to win a lot of net new share
on just the subscription side in both
B2B and and consumer, which we saw was
really interesting in Q1.
>> Yeah, I guess three things have happened
since the last time we talked about this
topic on the podcast, um, which is first
of all they signed that massive deal
with SpaceX AI Cursor, which added a ton
of uh, Jeremy Wyatt laugh at over there.
>> This SpaceX AI Cursor, beautiful,
beautiful.
>> Yeah, not Cursor not part of it yet
because they're trying not to change
their S-1, I think, but anyway, the uh,
the SpaceX S-1 revealed the many
billions that they're spending
Anthropic is spending with SpaceX with
the
uh, clause to let them back out. And uh,
meaning
let XAI has the ability to reclaim these
GPUs if they want, but that should be
some contribution to Anthropic's total
revenue. In other words, if they were
constrained by compute
uh for their ability to grow the the
business on the consumer subscription
side and enforce rate limits or things
like that, those those should go away
pretty quickly. The second thing that
happened is obviously that they raised
their series H $65 billion in funding at
a $965 billion valuation. So, raising 65
billion at 900 billion and resulting in
965 post-money. I don't know why they
didn't round that up to a nice even
trillion, but uh
we'll see.
And then
>> Isn't that like almost double from
February number, their valuation, right?
>> Yeah, what was the February number? They
were at 400 something?
>> 380 or 400, somewhere around there.
>> I guess they need to have their
valuation track with their ARR growth.
>> I mean, it's not a crazy multiple, 20x
like versus, you know, the the software
bubble back in '21, like
some of those were like 80x like
Snowflake, Cloudflare.
It's not like super expensive, and then
too, like, you know, from a Wall Street
Journal article, you know, recently my
financials, you know, that we also like
had in our model,
was, you know, they're they're
profitable now. You know, when you when
you exclude stock-based compensation,
it's, you know, Anthropic's a a pretty
good business model and like they're
seeing a ton of operating leverage.
Um and then we also know from the
Information article we read today like
you know, OpenAI's not seeing that um as
well and
um it's very clear like what the better
business is right now and and you know,
what the better business model is.
Um there's obviously a lot of like
operating, you know, deleverage if if
you know, people start cutting how much
they're spending on on coding tokens,
but
>> Yeah.
>> Um
>> But, we're not seeing that right now.
We're not
>> We're not.
>> I mean, yeah, they're still growing at
10 billion net new ARR a month the last,
you know, 3 months, which is, you know,
pretty crazy.
Um they're probably going to do another
10 in June.
And
I mean, who knows, you know, who knows
where they end up at the end of the year
like, you know, you know, Mythos getting
released as well, like that probably you
know, that does help.
Um
I mean
there's not like there's no train that's
slowing right now. And they're in all
the right places.
Um yeah, they're like seeing the
benefits of that like just throughout
the entire business model.
>> Big question here.
X XAI deal or SpaceX deal with
Anthropic?
Is it bullish or bearish for the markets
overall?
>> Which market? Neo cloud market or
compute demand? Overall, you know,
everything's related, right? Stocks,
compute demand,
um
Nvidia, all of it.
>> I think it's super bullish. Neo clouds,
man. Everybody wants to be a neo cloud.
Neo cloud is the terminal business. Even
AI labs want to be neo clouds selling
their compute to whoever they choose to
right now.
>> Well, but
I
I agree with this because this this deal
is basically
um
one player that was supposed to be a
source of demand becomes a source of
supply. So, now suddenly there's more
competition in the supply market, which
is GPUs as a service, and the
off-takers, there's one less.
I don't know, man.
It's uh
>> Yeah, I I disagree with this. I think
Cursor's training plenty of models on
Colossus right now. I think they
wouldn't have that provision in the
contract to take back their GPUs if they
really were going to be no source of
demand in the future.
Um I think that there's lots of demand
to go around and the reason that deal's
happening is because Anthropic's demand
is so overwhelming right now that they
need to do these crazy things like buy
compute from their competitors in order
to be able to serve that demand. If they
had a different way to be able to serve
that demand, they would be doing it, I
assume.
>> I mean, on the other hand, if you're a
you know, if you're XAI
uh or Meta or whoever uh other ones of
these labs that are uh lagging, you're
seeing Anthropic, you're like, "Whoa,
bro, this is what I could do if I was at
a frontier."
All right. Like if if Grok suddenly was
at the frontier, they could be a hundred
billion dollar AR business. So, to some
extent, you would argue this should make
them more bullish.
>> I think they said one has about 24
trillion of enterprise AI applications
carved out as future market TAM, right?
So, no don't even need to be a hundred
billion AR. You can just take that
forward a few more quarters and take it
to
>> [laughter]
>> 24 trillion, no problem.
>> saying So, their TAM is 24 trillion, but
uh that that's for what? GenAI
applications?
>> I'm going to bring up that chart now
from the S-1. Yeah, it's it's like
you know, there's like space and Telco
and then there's a really big section
for enterprise AI applications.
>> Okay, enterprise AI applications. So,
that would basically be the TAM for
Grok, right? And so, they're saying
actually we give up on the 24 trillion
TAM, you know, I'd rather give my
compute to Anthropic. That's a better
use case than fighting for a 24 trillion
dollars TAM.
>> Well, I I think in some ways it's just a
matter of when you can spend that money.
Like or when you can spend that compute.
In other words, there's potentially
some serial nature to the development of
AI progress, where you have to wait for
you or all of your competitors to run a
bunch of experiments to figure out what
the optimal model architecture and data
set mix or just creation of data,
synthetic data, whatever it is, is done.
Because if we were to just take a lot of
these models that people are running
today and try and run them on hardware
three years ago,
like actually the the model
architectures would run really well. Um
all the innovations in like sparse city
and attention and like these are um
these are huge improvements over like
dense models from three years ago,
right?
>> So, they they could have used that
compute uh if you're saying, "Hey, may
maybe they had a bottleneck because they
weren't able to figure out like new
architectures that would have enabled
them to like use that compute
efficiently." Then they could have used
that compute to
to do research on these specific topics,
right? Like less training, more
research. Uh or they could have used it
to like give a whole lot of tokens to
their employees and then make them much
more productive at, you know, doing a
bunch of research tasks.
Um you know, now we know that uh we know
that AI can do pretty complex science
problems, right? The OpenAI math stuff.
Uh which I know a few things about,
right?
>> [laughter]
>> Uh cuz
cuz my dad does math for a for a living.
Um
so uh yeah, I mean, I don't know. That
does tell you that uh it's a odd
decision uh when you're kind of in the
fight and giving up when
uh you're seeing the strongest signals
we've ever seen that this time is real
and accelerating and uh actually, you
know, $10 billion of ARR per month,
right? So,
>> [laughter]
>> it's sad and it's pretty odd. Uh
Cuz like I I I I I feel like this is
opposite to to to Meta. Like my sense is
like to some extent AI is giving up on
the frontier race, whereas Meta is if
anything getting more pulled up because
of what they're seeing from Anthropic.
They're like, "Hell yeah, this is what I
bet on. I'm going to [ __ ] double down
after having double down, you know, so
many times already." So, uh Meta is the
one that is seeing Anthropic and being
like, "Hell yeah."
Right? Uh I want that.
>> Well, I Okay, I think there's two
dynamics that you're overlooking here a
little bit. Potentially, one is that
Meta has a huge cash-generating business
that they can use to fund all of this
and SpaceX really just doesn't have this
business that generates hundreds of
billions of dollars of free cash flow
that they can pour into compute for the
research bets. So, they have to do it
based on venture capital, which is
unfortunately finite when you're talking
about the scale of tens or hundreds of
billions of dollars and needs returns on
some timeline, whereas Meta can do it on
a longer timeline. The second thing is I
think that the
optionality of having access to compute
that you can then take back and pour
into something is actually quite
powerful, which is to say, if they're a
cash-generating neo cloud business that
can do some research on the side and
then in the future have some
breakthrough or have some distribution
mode with X or you know, something in
the
you know, Starlink or something in Tesla
relationship or something that just
means they can take advantage of it.
They should be in theory be able to then
pour that compute into that that thing
that's just not ready yet. And
I guess what I'm saying is that well, I
I really actually love Joey to to cover
a little bit about like the earnings
before training concept which is say
like other labs are spending a whole
bunch of money training models right now
that they need some return on and right
now they are getting them namely
Anthropic and and Open AI is getting a
return on these models. But Meta is
getting no return on their models
outside of the Rexis stuff. There's no
return on new spark for example.
XAI has very limited. They had like a
million subscribers to Supergrok or
something. It's quite quite different
than
approaching a billion, you know, miles
for for some of these consumer
applications, right? And so I think that
the like the the play to say well,
cursor is doing pretty well training on
Kimmy. Why don't we just let the open
source guys build us a model for the
next year and then we'll take our
compute back and go run a bunch with it
instead of spending a bunch right now
just to keep being in fourth or fifth
place is is
uh well, I I think it plays into the
>> I have huge disagree.
>> earnings before training argument,
right?
>> Massive disagree, but we've been
monopolizing the the speech for a bit.
So I'll let Joey take in, but uh
I have huge disagree here.
>> Well, you got to you got to explain now
after he says something.
>> [laughter]
>> Yeah, sure. Well, yeah, I mean like
look, it's it's pretty simple, man. Like
uh
what we're seeing right now is that
there are tremendous return returns to
training compute.
I think it's pretty clear.
And you basically want to make sure you
have more than others if you want to
stay in the race.
Open source versus frontier, I think
it's pretty clear that the gap is
expanding not closing which everyone was
saying last year open source is going to
catch up or the gap is going to close.
The gap is closing. China is getting
closer. No, that's not happening. The
frontier is like beating to a massive
extent the open source models as
demonstrated by Anthropic they are our
trend.
Right like I think anyone who does
production or clothes sees the
difference between Claude and No Kimi or
Deep Seek V4. Um
if you were to take a guess, would you
imagine that the gap is going to expand
or is going to narrow? I would assume
that it's going to expand because one
has much more compute than the other
which goes back to the fundamental point
which is that you know training compute
has tremendous returns and so not having
compute means that you're disadvantaged
relative to competitors.
>> I think we're agreeing about one thing
which is that training compute has
massive returns if you're in the first
place, but not necessarily if you're in
fifth place. Like
>> Not necessarily. And I I think the point
is okay, let's imagine this. Let's say
Meta goes like completely crazy next
year. They secure um
an amount of compute so they have let's
say they have five x more training
compute than Anthropic. I guess do you
think they catch up? And they're they're
they're really far behind, but if they
have five x more compute than Anthropic,
do you think they catch up?
>> No, probably.
>> No?
>> No.
>> Obviously they need talent as well, but
>> Yeah, please go ahead.
>> I think it I think they need the talent.
And so I do think there's a potential
for I think you need both compute and
the talent basically.
Um and so I I think that like the maybe
the these go hand in hand where when XAI
gives up all of their talent or close to
it with all the co-founders leaving and
then they give up all their compute, it
kind of goes hand in hand there.
>> Yeah, I know 100% agreed, but I think
the point is these two things tell you
that they're basically out of the race.
Then it's going to be incredibly tough
for them to
to like
you know come back and uh extract value
out of the $24 trillion and write AI
applications market.
>> Okay, that came up again. So, I'm going
to pull that up on the S1 and then Troy,
maybe you can get us away from this
argument and talk about the uh
Yeah, this stuff here. Here's their
Here's their TAM. It wasn't 24, it was
22.7 trillion dedicated enterprise
application.
Little down here is Starlink.
>> [laughter]
>> That is insane.
>> Anyway.
>> Um
>> Does the race even matter though? Cuz I
feel like at a certain point if Meta
gets so much more compute and their
model gets so much better, even if
they're like number three, number four,
number five, or whatever, it's good
enough for most people to use, right?
And it's probably good enough to be
replacing a lot of jobs already. That
you don't have to be number one or
number two to be winning.
>> No. No. No. No.
>> Uh No, I Okay. I I think this kind of
comes down to your perspective on how
they actually use the models. So, like I
don't know. We're We're saying number
three, four, five here just to define it
from my perspective, which others may be
disagree with, is like Anthropic's uh in
first place right now because I believe
coding is the only thing that matters. I
think they've been proven correct.
Coding is not coding, it's computer use
and everything a human can do with a
computer and AI can do with a computer.
And so, therefore, this is interface to
the computer, not coding. So,
Anthropic's in first, OpenAI's in
second. I put Cursor in third. My
experience using Composer is
significantly better than using Gemini
or using Muse Spark because it's like
you can't use it
um and uh from using the Grok models
from xAI. And so, I I In some ways, I I
think that they have both the
distribution and the model to be in
third place right now. And the question
is just how much compute does Cursor
need to stay in the race? And I think
that having the optionality to feed them
more compute in the future is super
compelling. Whereas, they can't use it
right now. And so, why have it on your
balance sheet if you can't actually use
it? Why not turn it into a
revenue-generating asset and use it
later once you have more distribution or
once you build out the training stock to
improve it to the point where you can
actually do these hero runs. But we'll
see because there's there's it's really
interesting that there's four labs, five
labs in the US testing this theory from
different angles. Some are stacking
compute and have a bunch of revenue,
some are stacking compute and have no
revenue and some actually have quite a
bit of revenue if you look at cursor and
don't have that much compute right now
on a relative basis.
And I'd like to see all three pursue it
that way because
I'm not sure what the right playbook is
or
who the winner will be.
Um but we'll see. I mean it's going to
be interesting for Anthropic to attempt
to defend their number one position
because that's not a position they've
been in before. They've only had to play
catch up.
Um and I think that's actually quite
hard. I think it's quite hard to retain
talent. I think it's quite hard to
um keep pressing a compute advantage. I
think it's it's quite hard to motivate
users and and consumers to keep
consuming more instead of getting, you
know, grass is always greener with some
new feature from some competitor. And uh
I think it's up to them to to maintain a
trillion-dollar market cap.
We'll see.
>> This is Yeah, Jordan, I think it
>> Jeremy, you want to go?
>> I said we've been talking a lot, man. I
want other people to talk, but I had a
response for Crystal, but you go first
and then
>> it goes into like the the earnings
before, you know, training interest and
taxes. But I think it's really
interesting like, you know, we think of
ETA like before training interest taxes
is like the, you know, the cash
operating profits that they you generate
for running inference. Um and if you
want to think of training and research
as CapEx, you go back to this
conversation like I think this is a big
investor question, like even corporate
strategy question is that return on
invested capital. Um and like right now
we know like Anthropic is obviously
having massive massive returns on the
invested capital they not only put in
these models, but also like just for
coding applications, these computer
applications specifically. And then, you
know, we're seeing more and more news of
other
um you know hyperscalers uh or not
hyperscalers at their labs um trying to
get into this coding you know market. I
think there was some Microsoft news this
morning on that. Uh you know mixed mixed
opinions I think here of of how
successful that might be. Um but
training more and more models for this
um because obviously people do want to
use frontier models to Jeremy's point.
We even saw like the meta token maxing
article like
you know all that spends external. Uh
token spends external because of
you know because to Jordan's like
computer and coding point.
Uh that's where like you know there's a
ton of obviously product market fit and
they're seeing their own ROI when they
use the product.
Um
but yeah, that's heavy. I I'm guessing
Jordan's still on the call. So Jeremy
I'll I'll send it back over
>> Yeah, yeah.
>> to you.
>> J- just wanted to respond to like
Crystal's point because
yeah, I mean it obviously depends like
how you think of the the market evolves,
but
I really like the micro framework that
Malcolm has which is that okay, I think
like you know 2030 2035 like uh what
type of uh tasks are going to drive the
bulk of the total addressable market,
you know the dollars that people are
actually spend on AI. And there's tasks
that are like uh have a certain amount
where they're good enough that's kind of
finite like translation maybe at some
point there's only time it doesn't make
any sense uh to spend more on
translation, but then there's like these
very open-ended task where the spending
is pretty much infinite. And then on
legal for example, you could assume that
if AI is really good at at legal, if you
want to make sure you beat your
competitor, you want to probably want to
spend more on AI than he does, have more
intelligence than he does uh because you
want to gather more evidence, you want
to think through many different ways of
uh you know coming up with a defense and
whatnot. Um scientific research research
is typically the very open-ended use
case. Um you know health care um for us
as analysts, we try to get insights from
gathering a lot of data, very
open-ended. Um I would assume these use
cases are going to drive much more
spending than this the finite sort of
it's good enough, right? And and and
really when you think of these
open-ended use cases, what what matters
to be able to do like what you want to
do uh
uh in the cheapest way. And if you're
only doing the cheapest way, you're
going to have to use frontier models.
It's the whole point we made we made
like Mythos is actually 2 first cheaper
than Opus.
The model in terms of token pricing is 6
times, I think, 5 times more expensive
than Opus.
But you know, if it's like 10 times
smarter, if it requires 10x less tokens
to answer a given task, then it's
actually way cheaper to complete that
task with the the model, right? And so I
think like at least that's that's kind
of the way I view it. And so
I I I think the bulk of the market is
going to concentrate to
the frontier models. I think if you're
three or four or five, you're not going
to get any dollars. And I think like if
you take step back and think, "What are
the signals that we've seen in 2026? Did
the AI market has it beaten or has it
missed versus expectations we had in
2025?" You know, massive beat. But then
what I think is super interesting is
what is the composition of this beat?
Did everyone beat or is it just a few
companies, right? And I I think that is
actually kind of interesting. On the
frontier side, it's basically one
company, it's just Anthropic, massive
beat.
Google is probably tracking behind to
some extent when you look at just the
Gemini adoption and how much people were
spending on Gemini. OpenAI is tracking
behind. Obviously xAI, Meta are nowhere
to be seen. You could have hoped that
they would have had something. They
don't really. To be fair, on the
open-source side, I think that it's also
been a beat. I think there's been some
good adoption, but the dollars uh spent
are still pretty small. So I think it's
an interesting composition of massive
beat where it's basically all driven by
one company, which kind of gives gives
you a sign that it's pretty much winner
takes all. And if you're state of the
art, you get the bulk of the volume, and
if you're not,
yeah, people don't spend on you. I know
Joey stopped playing. What do you think?
>> Can you repeat the last part of that?
>> Holy [ __ ] you didn't listen to my
beautiful prose. I'll jump off.
>> I heard that I heard the most of it.
>> I got to jump off and respond to this.
So, Jeremy, I I think you're this is
totally this totally makes sense, but
we're also seeing massive beats or
massive reported ARR numbers from
startups serving open-source models
right now. It's not just Anthropic
growing.
There's there's huge growth for
Fireworks that's tied to Cursor.
>> that big, man.
>> Really big.
>> They're not that big, man. That is the
thing. They're they're not that big.
Some encouraging signals. I think Cursor
is at what? 2 2 billion now?
Uh and they were at what? Maybe
1.something at the end of 2025. There's
still like really good growth.
Definitely you could say Cursor is a
beat. I don't know WinZO, I guess,
they're in Google, but they're worth to
be seen.
>> Yeah.
>> No, but I I I I I agree with you. Some
of the open-source guys have had a beat,
but in terms of like dollar amount,
still like not very meaningful.
>> Yeah, I mean the claim from Fireworks is
$315 million of ARR, right? Like that's
>> What is 300 million between between
friends?
>> [laughter]
>> Okay. Like okay.
In the past, a startup unicorn was
interesting when they were a
billion-dollar valuation. Now they start
to approach billion-dollars ARR, and you
go, "Ah,
ah, whatever. Fly on the wall. Like
>> Relative to the size of the market, you
know, which is like already above 100
billion dollars. It's like in the grand
scheme of things, not that big. Think of
the the positioning of the different
hyperscalers. I mean, we've talked about
we've talked about like, you know,
business models and whatnot,
but the beauty of the the economics 2.0
this
magic model is just so accurate,
is that it covers
>> Yeah,
I mean, you know, in 2 weeks we've
gotten some pretty good feedback so far
from some of that
you know, these you know, hyperscalers
customers, things like that. so um it's
been it's been solid, but I mean, I
think right now like just given like
Amazon continues to win.
Yeah, I think the things that benefit
them you know, around their customer
base, you know, benefit you know, Azure
you know, pretty similarly.
Especially as you know, you look at you
know, more and more you know, token as a
service type models that come on the
foundry.
And then obviously I think you know,
Google if they get a coding model you
know, right now you know, when we look
at what was formerly Vertex and is now
Gemini H enterprise platform.
You know, we we still think you know,
Gemini API is a pretty decent percentage
of that. So we're not benefiting a ton
from
you know, Claude and some of those
things, but like that's probably the
biggest thing right now. I mean, we
still see Amazon like Bedrock you token
as a service is a pretty you know, I
think by the end of the year it could be
the majority of the AI business
at Amazon and even though Amazon like
lags Google and or GCP and Azure like as
you know, their AI mix. Yeah,
at at AWS AI is a lot smaller percentage
of the business. But with Bedrock going
to the majority of the AI business and
then infrastructure as a service being
you know, 80 90% at Azure and GCP.
Um
you know, it's really really
advantageous. But then too to that
point, it's a very easy for Azure to
then come in you know, add Claude as a
model.
Um and kind of you know, this token as a
service business isn't that hard for you
know, these people with massive customer
bases to implement. That's much harder
as you go down to like Oracle.
Um
and then to the neo clouds like like to
implement this at the same level just
given they don't have you know, the
massive inertia customer bases.
Um
that you know, really benefit from like
the old school software distribution
enterprise software distribution like
you know, emotes.
Um
>> Yeah, let me throw
let me throw this chart on on screen
just to make the point Jeremy was making
and you're making right now which is
like bunch of rounding error for
anything but the top three.
Hyperscalers when it comes to the
inference play business, right?
Um yeah, it's uh
>> [laughter]
>> Yeah, very small portion of the market
that's going to ever to everybody else
bucket.
>> And it's also I mean that that kind of
goes to the to the same thing the same
point I was mentioning earlier with
regards to like the composition of the
market.
One of the important point he makes in
that article is that um
the key to being a successful token as a
service business
is that actually just to have
partnerships with the big labs and have
access to frontier models. It's pretty
simple, right? So, I guess it's uh the
big disadvantage that currently Nebius,
CoreWeave, Iron, all of these guys have
is they don't yet have the partnership
and they also don't have maybe the
capital to be able to deploy GPUs
without a 5-year contract. Um so, it's
kind of like as a as a function of the
way the market works that a company like
CoreWeave or, you know, Nebius have the
bulk of I mean maybe not Nebius
actually, but Iron have the bulk of
their business contracted over multiple
years. Um whereas Amazon, you know, is
free to be more spec
and better pick more demand that don't
have that 5-year off-take uh take or pay
locked in.
>> Yeah, well
>> And that I think too, Jordan, I think on
like who wins is is definitely the
custom silicon portion.
Um You know, around Tranium and you
know, TPUs at GCP that I think is a big
thing versus like in the tokenomics
model that and then when we look at like
brain accelerator in
um and some of the data center model
numbers
you know, it's really at Azure you know,
just as mostly in video and then you
look at that vertical integration at
um
AWS and GCP, like that's another big
advantage for them.
Uh that that I think, you know, probably
you know, especially on the margin side,
I should say. And then as we've seen a
lot of, you know, inference get more
efficient and you know, gross margins on
inference
you know, drastically improve at the
labs, you know, the two you know,
frontier labs, major frontier labs over
the last two years like that's
definitely been
you know, another like key consideration
when you think about who's going to win
you in this market.
>> Yeah. Yeah, I mean that it's from the
technical perspective, everything that
we've criticized all the chip startups
about and TPU and Trainium is always
about usability. But if the entire chip,
you know, a gigawatt of Trainium at
Rainier is all just serving tokens from
one to three models. Well, the customer,
the end user customer doesn't even
necessarily need to know that much about
what chip is running if they're only
buying tokens.
And certainly not the end users, like
the actual terminal user of the tokens,
you know, when I'm using Cloud Code, I
have no understanding of whether my
token is coming from a TPU or a Trainium
accelerator or a GPU. Doesn't make a
difference, right? Okay, guys, well,
really interesting stuff today. Anything
left unsaid on the the topic of
uh Bedrock, tokens as a service,
Anthropic's growth.
>> Hey, I think we got it all, Jordan.
Winners win, losers lose, and it was a
clear trend.
>> [laughter]
>> Winners win. Awesome. Okay. Well, thanks
guys for coming on the show. Good job
today.
Ask follow-up questions or revisit key timestamps.
This episode of SemiAnalysis explores the growth of Anthropic and the significant impact of Amazon's Bedrock platform on AWS's margins. The discussion highlights a shift in the cloud market from traditional infrastructure-as-a-service to a 'token-as-a-service' model, which offers superior margins for hyperscalers. The participants analyze why Anthropic is succeeding, the importance of coding and computer-use workflows, and the competitive dynamics between frontier AI labs, including the impact of compute scarcity and the strategic decisions of companies like xAI and Meta.
Videos recently processed by our community