HomeVideos

Ep. 013 - AWS Margins Jump 10% While Azure and GCP Flatline (Tokenomics)

Now Playing

Ep. 013 - AWS Margins Jump 10% While Azure and GCP Flatline (Tokenomics)

Transcript

1387 segments

0:04

Hello everyone, welcome back to

0:05

SemiAnalysis

0:07

weekly, episode number 13, lucky number

0:09

13. We're here with Joey, Jeremy, and

0:11

Kristen. We're going to talk about an

0:13

article that we put out recently

0:15

uh called Anthropic growth and bedrock

0:17

mix drive AWS margins higher while peers

0:20

lag. That means we're going to talk

0:21

about everything Anthropic, including

0:23

the uh recent announcement of their

0:25

series H, the release of Opus 4.8, but

0:28

with the spin on focusing on

0:29

infrastructure, like how exactly they

0:31

serve these tokens, especially with

0:33

their partnership with AWS. Guys, uh

0:35

welcome to the show. Excited to talk

0:36

through this.

0:38

>> Thanks, Jordan.

0:38

>> Man.

0:39

>> All right, so let's uh let's dig in with

0:41

the article itself and and talk a little

0:43

bit about the backstory. Like, I think

0:45

uh a lot of people understand the

0:46

concept of tokens, understand what GPUs

0:49

are, um but and and certainly like cloud

0:52

providers and neo clouds, but not

0:53

everybody is getting their tokens from

0:55

the same place. So, can one of you guys

0:57

give me a backstory on bedrock? Like,

0:59

what is AWS? Uh what are they, you know,

1:02

in terms of what they're doing for

1:03

Anthropic and how they're serving tokens

1:05

with bedrock? Joey, we'll start with

1:07

you.

1:07

>> Yeah, yeah, yeah. Yeah, I'll go through

1:08

that. And I'll do it kind of for all the

1:10

clouds as well. Um so so if we really

1:12

look across all the hyperscalers, and

1:14

especially the big three, um when we

1:16

look at Amazon, Microsoft, and Google,

1:18

there's kind of two big breakouts, uh

1:20

maybe even three. There's a bit on the

1:22

software side. So, if you look at

1:23

Microsoft like GitHub, um Copilot, you

1:27

know, that's an area that's like an AI,

1:29

you know, software-as-a-service type

1:30

product. They have AI, you know,

1:32

infrastructure-as-a-service where

1:33

they're just renting out um these

1:35

accelerator chips. Um and then they kind

1:37

of have this token-as-a-service

1:38

business, you know, which is where

1:40

they'll, you know, essentially expose

1:42

these outside models or their own own

1:44

models um

1:46

to consumers to interact with.

1:48

Uh so it's a little bit of a different

1:49

business uh because instead of just

1:51

renting the underlying chip, you're

1:52

you're renting the underlying model, you

1:54

know, buying it through your cloud, you

1:56

know, your CSP account and kind of that

1:58

enterprise, you know, spending

1:59

agreement, you know, you have all the

2:01

same benefits around security,

2:02

availability zones. Um and able to buy,

2:06

you know, third-party models, you know,

2:08

and also some of these first-party

2:09

models, you know, through your cloud

2:10

provider of choice.

2:12

Um so at a high level,

2:14

um those are kind of the three big, you

2:17

know, buckets right now at the large

2:19

hyperscalers that we see and we're

2:21

seeing some big, you know, changes or

2:23

differences between them and how they've

2:25

gone about it strategically. Um and then

2:28

the two main things was Anthropic's

2:29

growth um really led, you know,

2:32

Anthropic's growth and then uh Amazon's,

2:35

you know, strategy around Bedrock has

2:36

really driven their margins higher

2:38

recently. Um

2:39

was really the takeaway from the

2:40

article. And this token as a as a

2:42

service is obviously much better

2:44

business for the hyperscalers than

2:46

infrastructure as a service, you know,

2:47

in the AI side.

2:49

>> Yeah, and the if you look at the big

2:51

three, they're kind of going in opposite

2:53

directions right now. Um just in terms

2:55

of their operating margin, right? That

2:56

was the biggest chart from the article.

2:58

I'll put it up on screen right now. But

3:00

Crystal, can you maybe explain like when

3:01

we look at AWS versus Google and versus

3:03

Microsoft and you try and break out the

3:05

cloud business, like is this all to

3:07

blame on Anthropic from our perspective

3:10

or like what's driving AWS to improve

3:12

their operating margins while

3:13

Microsoft's is declining and Google's

3:15

kind of stays flat here?

3:16

>> So I feel like a lot of it is that cloud

3:19

usage, right? Like a lot of people are

3:20

using cloud more and it's mostly routing

3:22

through AWS. And then like what Joey was

3:24

saying, the token as a service um

3:27

business model, they just have a better

3:28

margin on that as opposed to like the

3:31

infrastructure as a service model.

3:32

>> Makes sense, yeah.

3:33

>> Yeah, so I I I I think like if you if

3:36

you take a step back, like the the world

3:39

of clouds has really changed a lot in

3:41

2023 as you as you started seeing these

3:43

neo clouds, these GPU as a service

3:45

businesses because uh in the old days,

3:48

which is basically 2022 and before,

3:51

um

3:51

cloud cloud service providers, you

3:53

basically had three. Right, Amazon,

3:55

Google Cloud, Microsoft Azure, they had

3:57

amazing margins, easy returns on on

3:58

capital. Some new entrants like Oracle

4:01

were trying to get in, but right the

4:02

market was dominated by three players

4:04

that had an amazing business.

4:06

And now you get to GPU as a service and

4:07

what we first started to realize is that

4:09

the barriers to entry are much lower.

4:12

So, Jordan, you're probably, you know,

4:13

the best person to talk about this. You

4:15

probably know the the CEO of 150 Neo

4:18

Clouds or 200 Neo Clouds.

4:21

How many [laughter] now? 300? And anyway

4:25

>> over 200. Yeah, yeah.

4:26

>> [laughter]

4:27

>> It is pretty insane.

4:28

Um, now obviously some are better than

4:30

than others, but the point is the the

4:32

market has much lower barriers to entry.

4:34

Uh, the moat that uh, cloud used to be

4:36

doesn't really exist in the AI era

4:38

because it's really more about

4:39

infrastructure and especially the end

4:41

users, the whole point of cloud

4:43

computing was you make the IT much

4:45

easier. Um, folks don't need to have

4:47

such a big IT department internally.

4:49

They can just rent through the cloud.

4:50

Uh, it's super easy. Everything is well.

4:52

Uh, now the big end users, they want to

4:54

have much more control. Uh, you shift

4:56

from uh, platform as a service in the

4:58

old days to now bare metal. Uh, folks

5:00

want just just the metal. Uh, Open AI

5:02

wants to think that the way they like

5:04

it, Microsoft, Meta. Um, and so this is

5:07

basically the token as a service

5:08

business is basically the first case uh,

5:11

at scale where you see a an AI cloud

5:14

provider having a a business with a

5:16

different profile. Uh, right, because

5:18

uh, obviously as you have less of a moat

5:20

in the GPU as a service era, margins go

5:22

down. Uh, Oracle, best example, they

5:25

have an RPO of half a trillion dollars,

5:29

uh, which is a few quarters ago that was

5:31

as of as of Q1 is still bigger than

5:32

Amazon's.

5:34

>> [laughter]

5:34

>> Their backlog is bigger than Amazon's,

5:35

but no one gives them the credit because

5:37

people know it's much more risky, the

5:39

returns are not the same. Um,

5:41

and we've seen empirically that every

5:43

every single GPU as a service cloud has

5:46

faced struggles when they started to

5:48

ramp up their business. But our view as

5:50

a firm is we actually think that GPU as

5:52

a service business model is sound

5:53

business. Companies like Coreweave have

5:55

a sound business, but there's this lag

5:57

effect where as you ramp up and bring

5:58

more capacity online, there's these lags

6:01

that make your margins go down

6:03

temporarily because your asset base

6:05

depreciates because you have to pay data

6:07

center leases and so on and so forth.

6:09

And so it's really interesting to see

6:10

that Amazon is basically the first cloud

6:12

provider that in a time of unprecedented

6:15

capacity expansion,

6:17

over a gigawatt per quarter now,

6:19

they expand margins.

6:21

And that kind of tells you that if at a

6:24

period of accelerated capacity delivery

6:27

they can expand margins, you start to

6:28

think of okay, what's the stabilized

6:30

margin of this business?

6:32

And that's one of the points that we

6:33

make is the stabilized margins of token

6:34

as a service for Amazon are actually

6:36

extremely rich.

6:38

The return on capital is fundamentally

6:40

much more sound.

6:41

>> Yeah, let me throw this chart

6:43

up on screen actually from the from the

6:46

article which is the percent of revenue

6:48

that is going to Bedrock. Bedrock being

6:50

the token as a service at Amazon, right?

6:53

You can obviously see it ramp up quite a

6:55

bit at roughly the the time that we're

6:59

in right now, right? Q4 of last year and

7:01

first quarter of this year. And then if

7:03

we compare that to the chart I had on

7:04

previously where we're seeing their

7:06

operating margins improve in the first

7:07

quarter of this year, they're saying

7:09

that's due to it's Can you explain maybe

7:12

in more detail why it's unprecedented to

7:14

say that Amazon can bring on a gigawatt

7:16

per quarter and still improve operating

7:18

margins?

7:19

>> So

7:20

to to to understand this, you basically

7:22

have to go back to like

7:24

why are the pure bare metal providers

7:28

seeing their margins go down? So you

7:30

look at Coreweave, which is a pure play,

7:32

so they're the the cleanest example.

7:34

Oracle actually is kind of kind of the

7:35

same. When you look at this chart, it

7:37

essentially what this tells you is that

7:39

their stabilized business

7:41

does something like you know 25%

7:44

operating margins 30 to 40% gross

7:46

margins but the the whole issue is

7:49

stabilized stabilized means that your

7:52

GPU cluster is fully functional

7:55

you're getting the the monthly rent or

7:57

whatever rent from your customer

7:59

take or pay so it's a flat fee you know

8:01

exactly how much revenue you're going to

8:02

make often times it's a 5-year take or

8:04

pay contract again fixed rate

8:07

so you know your revenue you know your

8:08

cost everything is stable but before

8:11

getting there

8:12

the obviously have to set up the data

8:13

center which is a huge capital expense

8:16

up front

8:17

and then you have this whole process or

8:19

if you have the data center built but

8:21

you need to fill it with equipment that

8:23

takes a few months with some new stock

8:26

types of equipment like GB200 which is

8:28

super complicated we've seen that lag

8:30

get longer and longer which means that

8:32

this period of time where you depreciate

8:35

your assets you pay data center rent you

8:37

pay some labor you pay a whole bunch of

8:39

stuff that period of time goes up and

8:41

you don't make any revenue because your

8:42

cluster is not yet rolled out right and

8:44

so that's the whole dilemma that these

8:46

guys are facing is they know their

8:47

business model is stabilized structure

8:49

but they're they've been facing

8:50

challenges to wrap it up some a bit

8:52

contractual again GB200 some just more

8:55

structural because there is this lag

8:57

and Amazon is the same thing right like

9:00

like everyone else they have the same

9:01

profile they're bringing on a whole lot

9:04

of data centers

9:05

a whole lot of XPUs these XPUs in theory

9:08

they should take time to bring online

9:10

yet despite this you're seeing their

9:12

margins go up which you know pretty good

9:14

sign for them

9:15

>> Yeah but it and it's clearly different

9:17

than the others

9:18

in this space where the percentage of

9:20

their total AI revenue that they're

9:22

reporting

9:24

as being from token as a service as

9:26

opposed to other products is much higher

9:28

than Google and Azure like they are not

9:30

just bringing on capacity they're

9:31

successfully selling it into

9:34

the labs that are using it to serve

9:36

tokens, right? For these models.

9:39

Um

9:39

>> Yeah, and I guess the the chart you you

9:41

showed earlier, like what I forgot to

9:43

mention is obviously the margin buffer

9:45

that you have uh when you when you sell

9:47

tokens with your infrastructure as

9:49

opposed to having a 5-year take or pay

9:51

fixed contract with a capped upside.

9:54

Right? So, that's kind of the key one of

9:56

the key points of of of the article.

9:58

>> Yeah, yeah.

9:58

Guys, can you talk a little bit about

10:00

the workload mix as well? Like obviously

10:02

um any provider could conceptually do

10:05

this, but not everybody is doing it suc-

10:08

it's not like Azure doesn't have a token

10:10

as a service business or Google. In

10:13

fact, even Crusoe, CoreWeave, Nabia are

10:15

all trying to get into this business,

10:17

too, but they really they need a

10:19

customer and they need a customer

10:20

serving the right type of workload for

10:22

it to really result in a bunch of

10:24

growth, I would say.

10:25

>> Yeah, I I I would say, you know, for

10:28

Amazon specifically, you know, they

10:29

benefit from the biggest customer base.

10:31

Yeah, they've been doing this for

10:33

20-plus years now. At AWS,

10:36

um people are very comfortable buying

10:38

that, you know, through them. You know,

10:39

people even buy infrastructure software,

10:41

you know, through them, a provider like

10:42

MongoDB, like Snowflake. Um pretty large

10:45

marketplace business. So, customers are

10:47

really comfortable with the security,

10:49

uh you know, at this point um buying

10:51

through them, having a single bill for

10:52

all this. Um and I and I think this

10:54

comes back to you some of the Anthropic

10:56

news today, you also on the ARR number

10:58

of 47 billion. You know, we look at Q1

11:00

and even in the Q2 here,

11:02

you know, a lot of the Anthropic's

11:04

um you know, the business mix is much

11:06

much different. So, Amazon is benefiting

11:08

from picking, you know, being 80 to 90%

11:11

cloud in Bedrock, you know, versus

11:13

Microsoft is is heavily OpenAI,

11:15

obviously.

11:16

Um you know, Google has a lot of Gemini,

11:19

which doesn't benefit as much from a lot

11:20

of these agentic coding tasks.

11:22

Um

11:24

and when we look at kind of the OpenAI

11:26

or you know, and the coding percentage

11:28

that's really driving Anthropic, um you

11:31

know, you're talking probably like 10

11:32

billion a month net new ARR in the last,

11:34

you know, 3 months, you know, of March,

11:37

April, and May

11:38

um, together. And, you know, a lot of

11:41

that, probably 80% of that uh of

11:44

Anthropic's net new ARR is in this API

11:46

business. You know, if we go to, you

11:48

know, OpenAI, you know, 60% of that

11:50

business in Q1 was was consumer,

11:52

consumer subscriptions, really. Um, so

11:56

um, there's kind of a mix of factors,

11:59

but but like Amazon was in the right

12:00

place at the right time

12:02

uh with Anthropic. And then two, you

12:04

know, they were able to give with

12:05

Trainium 2, I think, pretty you know,

12:07

interesting deal structure for both

12:08

parties. And I think that was the big

12:10

thing, too. We we kind of mentioned in

12:11

the article.

12:12

Um, obviously, like their mix of

12:14

Trainium, you know, they could there's

12:16

an infrastructure as a as a service like

12:18

fee component that Anthropic pays for

12:20

this infrastructure, like Bedrock

12:21

infrastructure. Um, but then there's

12:23

some interesting like hurdles around

12:25

revenue share, you know, and then and

12:27

really margin share that happened. And

12:29

because, you know, I think Anthropic was

12:31

probably at, you know, 25 million of ARR

12:34

per

12:35

um, you know, mid-one off the top of my

12:37

head in Q1 versus like probably six back

12:40

in Q2 four, like, you know, the numbers

12:42

really made sense for both parties. And

12:44

they really, you know, both benefited.

12:46

>> So, yeah. Eric, uh Crystal, can you talk

12:49

a little bit about those forecasts you

12:51

guys were making like going into this

12:53

end of Q1 and forecast for Q2? You don't

12:56

necessarily get disclosures the same way

12:58

from Anthropic, but we've at this point

13:00

kind of been bang on with the

13:01

disclosures with the disclosed revenue

13:04

figures and the margin figures, right?

13:05

>> Mhm.

13:06

>> [clears throat]

13:06

>> I think for

13:08

us, it's a little easier to forecast

13:10

Anthropic than it was OpenAI just

13:11

because so much of Anthropic um, ARR

13:14

comes from like the API side. And

13:16

because we also use Anthropic and

13:17

there's a lot of data out there about

13:19

like how people are using all of the

13:21

different uh Claude models, like, it's

13:23

so much easier to predict the workflow

13:25

and um, kind of see what the token

13:27

consumption is going to look like and

13:29

they've kept token pricing relatively

13:31

stable these past two releases more or

13:34

less, right? Whereas like OpenAI

13:37

all a lot of their revenue comes from

13:38

like these subscriptions and you never

13:40

know like if consumers are going to

13:42

switch over to another one, so it's a

13:43

lot harder to quantify the number of

13:45

users that are using the subscription

13:47

when they can just cancel anytime versus

13:48

for like API.

13:50

>> Yeah. Can you explain a little bit about

13:51

the

13:53

release on 4.8 and 4.7 cuz

13:56

pricing's remained the same but fast

13:57

mode has changed and maybe there's some

13:59

other there's been some other changes in

14:00

terms of how they're doing pricing on

14:02

the API.

14:03

>> Yeah, they said pricing is the same,

14:05

right? For regular but then for fast

14:07

mode it's different. Another cool thing

14:08

that they said apparently was that it

14:10

doesn't hallucinate as much and they had

14:12

like a pretty cool bar chart showing

14:13

that it's like it's rate of like

14:15

hallucination is a lot lower for 4.8 and

14:17

4.7 supposedly. I haven't tested it out

14:19

yet and they said that it's super close

14:21

to like mid those previews, so hopefully

14:23

it'll stop just, you know, pulling

14:24

numbers out of thin air from when we're

14:26

using it for our analyses.

14:28

>> That'd be good. That'd be good if

14:29

numbers weren't pulled out of thin air.

14:31

Yeah.

14:31

>> Yeah.

14:33

>> Um

14:34

in terms of the the fast mode and

14:37

consumer subscriptions, you have

14:38

comments there on like

14:40

I don't know what we've learned over the

14:41

past few weeks or months from digging

14:44

into how Anthropic is running their

14:45

business on AWS and like what's uh

14:49

maybe what's the biggest percentage of

14:50

their

14:51

revenue or their margin contribution

14:53

across those different mixes like those

14:55

different types of workflows that people

14:57

could be consuming tokens on the API

14:58

for?

14:59

>> Yeah, I mean we don't like we have that

15:00

in the tokenomics model like and we're

15:02

doing a big study right now on on what

15:03

percentage geo coding market you know is

15:05

of of current you know token spend

15:07

especially on the API side.

15:09

Um you know, we've done a lot of work

15:11

too on just the consumer and to some on

15:13

the B2B data of where Anthropic has has

15:16

just been taking a ton of share of net

15:17

new customers year to date both in

15:18

consumer subscriptions and enterprise.

15:21

Um but I mean it's really difficult and

15:23

a big question among a lot of our

15:24

clients like how big is the the coding

15:26

market currently. Um, I think Anthropic

15:28

recently said at their financial

15:30

services day that financial services was

15:32

the second biggest vertical.

15:34

Um,

15:35

and we kind of know there's a pretty big

15:36

gap between coding and and financial

15:38

services in terms of especially in API

15:39

spend.

15:41

Just of how people, you know, use this,

15:42

you know, anecdotally. Um,

15:45

so it's, you know, and you know, too,

15:47

there's been a lot of like recent

15:48

comments too on token maxing it and

15:50

especially the Fortune 500s like how do

15:52

you budget for that? How do you blow

15:54

through that spend over time? How do you

15:55

measure ROI?

15:57

Um, so things will have to change.

15:59

You know, I think and then people will

16:00

put some processes in place, but I mean

16:03

I know guys, you know, like Jeremy get

16:05

insane ROI at the data center model and

16:08

his team and and him

16:10

give with that and and you know, maybe

16:11

there's less policing out there at some

16:13

other you organizations that just let

16:14

people go crazy.

16:16

Um,

16:17

but that's, you know, too, another

16:19

contributing factor I think what was

16:20

just, you know, the success of coding

16:22

and then too, Anthropic like, you know,

16:24

starting to win a lot of net new share

16:26

on just the subscription side in both

16:28

B2B and and consumer, which we saw was

16:30

really interesting in Q1.

16:33

>> Yeah, I guess three things have happened

16:35

since the last time we talked about this

16:36

topic on the podcast, um, which is first

16:40

of all they signed that massive deal

16:41

with SpaceX AI Cursor, which added a ton

16:46

of uh, Jeremy Wyatt laugh at over there.

16:49

>> This SpaceX AI Cursor, beautiful,

16:52

beautiful.

16:52

>> Yeah, not Cursor not part of it yet

16:55

because they're trying not to change

16:56

their S-1, I think, but anyway, the uh,

17:00

the SpaceX S-1 revealed the many

17:02

billions that they're spending

17:05

Anthropic is spending with SpaceX with

17:07

the

17:08

uh, clause to let them back out. And uh,

17:11

meaning

17:12

let XAI has the ability to reclaim these

17:14

GPUs if they want, but that should be

17:17

some contribution to Anthropic's total

17:19

revenue. In other words, if they were

17:20

constrained by compute

17:22

uh for their ability to grow the the

17:25

business on the consumer subscription

17:27

side and enforce rate limits or things

17:29

like that, those those should go away

17:30

pretty quickly. The second thing that

17:32

happened is obviously that they raised

17:33

their series H $65 billion in funding at

17:36

a $965 billion valuation. So, raising 65

17:40

billion at 900 billion and resulting in

17:43

965 post-money. I don't know why they

17:45

didn't round that up to a nice even

17:46

trillion, but uh

17:49

we'll see.

17:51

And then

17:51

>> Isn't that like almost double from

17:53

February number, their valuation, right?

17:55

>> Yeah, what was the February number? They

17:57

were at 400 something?

17:58

>> 380 or 400, somewhere around there.

18:00

>> I guess they need to have their

18:02

valuation track with their ARR growth.

18:04

>> I mean, it's not a crazy multiple, 20x

18:06

like versus, you know, the the software

18:09

bubble back in '21, like

18:11

some of those were like 80x like

18:13

Snowflake, Cloudflare.

18:15

It's not like super expensive, and then

18:16

too, like, you know, from a Wall Street

18:18

Journal article, you know, recently my

18:19

financials, you know, that we also like

18:21

had in our model,

18:23

was, you know, they're they're

18:24

profitable now. You know, when you when

18:26

you exclude stock-based compensation,

18:28

it's, you know, Anthropic's a a pretty

18:30

good business model and like they're

18:32

seeing a ton of operating leverage.

18:34

Um and then we also know from the

18:35

Information article we read today like

18:37

you know, OpenAI's not seeing that um as

18:40

well and

18:41

um it's very clear like what the better

18:43

business is right now and and you know,

18:45

what the better business model is.

18:47

Um there's obviously a lot of like

18:48

operating, you know, deleverage if if

18:51

you know, people start cutting how much

18:53

they're spending on on coding tokens,

18:55

but

18:55

>> Yeah.

18:56

>> Um

18:56

>> But, we're not seeing that right now.

18:58

We're not

18:58

>> We're not.

18:59

>> I mean, yeah, they're still growing at

19:00

10 billion net new ARR a month the last,

19:03

you know, 3 months, which is, you know,

19:05

pretty crazy.

19:06

Um they're probably going to do another

19:07

10 in June.

19:09

And

19:11

I mean, who knows, you know, who knows

19:12

where they end up at the end of the year

19:13

like, you know, you know, Mythos getting

19:15

released as well, like that probably you

19:17

know, that does help.

19:18

Um

19:19

I mean

19:21

there's not like there's no train that's

19:23

slowing right now. And they're in all

19:24

the right places.

19:26

Um yeah, they're like seeing the

19:27

benefits of that like just throughout

19:29

the entire business model.

19:31

>> Big question here.

19:32

X XAI deal or SpaceX deal with

19:35

Anthropic?

19:37

Is it bullish or bearish for the markets

19:39

overall?

19:39

>> Which market? Neo cloud market or

19:42

compute demand? Overall, you know,

19:44

everything's related, right? Stocks,

19:46

compute demand,

19:48

um

19:48

Nvidia, all of it.

19:50

>> I think it's super bullish. Neo clouds,

19:52

man. Everybody wants to be a neo cloud.

19:53

Neo cloud is the terminal business. Even

19:56

AI labs want to be neo clouds selling

19:58

their compute to whoever they choose to

19:59

right now.

20:00

>> Well, but

20:03

I

20:03

I agree with this because this this deal

20:05

is basically

20:07

um

20:08

one player that was supposed to be a

20:10

source of demand becomes a source of

20:12

supply. So, now suddenly there's more

20:14

competition in the supply market, which

20:16

is GPUs as a service, and the

20:17

off-takers, there's one less.

20:20

I don't know, man.

20:22

It's uh

20:22

>> Yeah, I I disagree with this. I think

20:24

Cursor's training plenty of models on

20:25

Colossus right now. I think they

20:27

wouldn't have that provision in the

20:28

contract to take back their GPUs if they

20:30

really were going to be no source of

20:32

demand in the future.

20:34

Um I think that there's lots of demand

20:37

to go around and the reason that deal's

20:38

happening is because Anthropic's demand

20:40

is so overwhelming right now that they

20:42

need to do these crazy things like buy

20:44

compute from their competitors in order

20:46

to be able to serve that demand. If they

20:48

had a different way to be able to serve

20:50

that demand, they would be doing it, I

20:51

assume.

20:52

>> I mean, on the other hand, if you're a

20:54

you know, if you're XAI

20:55

uh or Meta or whoever uh other ones of

20:58

these labs that are uh lagging, you're

21:01

seeing Anthropic, you're like, "Whoa,

21:03

bro, this is what I could do if I was at

21:05

a frontier."

21:06

All right. Like if if Grok suddenly was

21:09

at the frontier, they could be a hundred

21:11

billion dollar AR business. So, to some

21:14

extent, you would argue this should make

21:16

them more bullish.

21:17

>> I think they said one has about 24

21:20

trillion of enterprise AI applications

21:22

carved out as future market TAM, right?

21:24

So, no don't even need to be a hundred

21:26

billion AR. You can just take that

21:28

forward a few more quarters and take it

21:29

to

21:30

>> [laughter]

21:31

>> 24 trillion, no problem.

21:33

>> saying So, their TAM is 24 trillion, but

21:36

uh that that's for what? GenAI

21:38

applications?

21:38

>> I'm going to bring up that chart now

21:40

from the S-1. Yeah, it's it's like

21:42

you know, there's like space and Telco

21:44

and then there's a really big section

21:46

for enterprise AI applications.

21:49

>> Okay, enterprise AI applications. So,

21:51

that would basically be the TAM for

21:52

Grok, right? And so, they're saying

21:54

actually we give up on the 24 trillion

21:57

TAM, you know, I'd rather give my

21:58

compute to Anthropic. That's a better

22:00

use case than fighting for a 24 trillion

22:02

dollars TAM.

22:03

>> Well, I I think in some ways it's just a

22:05

matter of when you can spend that money.

22:08

Like or when you can spend that compute.

22:11

In other words, there's potentially

22:14

some serial nature to the development of

22:16

AI progress, where you have to wait for

22:18

you or all of your competitors to run a

22:20

bunch of experiments to figure out what

22:22

the optimal model architecture and data

22:24

set mix or just creation of data,

22:26

synthetic data, whatever it is, is done.

22:29

Because if we were to just take a lot of

22:30

these models that people are running

22:31

today and try and run them on hardware

22:32

three years ago,

22:34

like actually the the model

22:36

architectures would run really well. Um

22:38

all the innovations in like sparse city

22:40

and attention and like these are um

22:44

these are huge improvements over like

22:45

dense models from three years ago,

22:47

right?

22:48

>> So, they they could have used that

22:49

compute uh if you're saying, "Hey, may

22:51

maybe they had a bottleneck because they

22:53

weren't able to figure out like new

22:55

architectures that would have enabled

22:57

them to like use that compute

22:58

efficiently." Then they could have used

23:00

that compute to

23:02

to do research on these specific topics,

23:04

right? Like less training, more

23:05

research. Uh or they could have used it

23:08

to like give a whole lot of tokens to

23:09

their employees and then make them much

23:12

more productive at, you know, doing a

23:13

bunch of research tasks.

23:15

Um you know, now we know that uh we know

23:17

that AI can do pretty complex science

23:20

problems, right? The OpenAI math stuff.

23:22

Uh which I know a few things about,

23:24

right?

23:24

>> [laughter]

23:25

>> Uh cuz

23:27

cuz my dad does math for a for a living.

23:29

Um

23:31

so uh yeah, I mean, I don't know. That

23:32

does tell you that uh it's a odd

23:34

decision uh when you're kind of in the

23:37

fight and giving up when

23:39

uh you're seeing the strongest signals

23:41

we've ever seen that this time is real

23:42

and accelerating and uh actually, you

23:45

know, $10 billion of ARR per month,

23:47

right? So,

23:48

>> [laughter]

23:49

>> it's sad and it's pretty odd. Uh

23:51

Cuz like I I I I I feel like this is

23:53

opposite to to to Meta. Like my sense is

23:56

like to some extent AI is giving up on

23:58

the frontier race, whereas Meta is if

24:01

anything getting more pulled up because

24:02

of what they're seeing from Anthropic.

24:04

They're like, "Hell yeah, this is what I

24:05

bet on. I'm going to [ __ ] double down

24:07

after having double down, you know, so

24:09

many times already." So, uh Meta is the

24:11

one that is seeing Anthropic and being

24:13

like, "Hell yeah."

24:14

Right? Uh I want that.

24:17

>> Well, I Okay, I think there's two

24:18

dynamics that you're overlooking here a

24:19

little bit. Potentially, one is that

24:21

Meta has a huge cash-generating business

24:23

that they can use to fund all of this

24:25

and SpaceX really just doesn't have this

24:26

business that generates hundreds of

24:28

billions of dollars of free cash flow

24:30

that they can pour into compute for the

24:31

research bets. So, they have to do it

24:33

based on venture capital, which is

24:35

unfortunately finite when you're talking

24:37

about the scale of tens or hundreds of

24:39

billions of dollars and needs returns on

24:41

some timeline, whereas Meta can do it on

24:43

a longer timeline. The second thing is I

24:45

think that the

24:46

optionality of having access to compute

24:49

that you can then take back and pour

24:50

into something is actually quite

24:52

powerful, which is to say, if they're a

24:54

cash-generating neo cloud business that

24:56

can do some research on the side and

24:57

then in the future have some

24:59

breakthrough or have some distribution

25:01

mode with X or you know, something in

25:03

the

25:04

you know, Starlink or something in Tesla

25:07

relationship or something that just

25:08

means they can take advantage of it.

25:09

They should be in theory be able to then

25:12

pour that compute into that that thing

25:14

that's just not ready yet. And

25:17

I guess what I'm saying is that well, I

25:19

I really actually love Joey to to cover

25:21

a little bit about like the earnings

25:22

before training concept which is say

25:25

like other labs are spending a whole

25:26

bunch of money training models right now

25:28

that they need some return on and right

25:30

now they are getting them namely

25:32

Anthropic and and Open AI is getting a

25:33

return on these models. But Meta is

25:35

getting no return on their models

25:37

outside of the Rexis stuff. There's no

25:39

return on new spark for example.

25:41

XAI has very limited. They had like a

25:43

million subscribers to Supergrok or

25:45

something. It's quite quite different

25:47

than

25:48

approaching a billion, you know, miles

25:50

for for some of these consumer

25:51

applications, right? And so I think that

25:53

the like the the play to say well,

25:56

cursor is doing pretty well training on

25:58

Kimmy. Why don't we just let the open

26:00

source guys build us a model for the

26:02

next year and then we'll take our

26:03

compute back and go run a bunch with it

26:05

instead of spending a bunch right now

26:06

just to keep being in fourth or fifth

26:08

place is is

26:10

uh well, I I think it plays into the

26:12

>> I have huge disagree.

26:13

>> earnings before training argument,

26:14

right?

26:15

>> Massive disagree, but we've been

26:17

monopolizing the the speech for a bit.

26:20

So I'll let Joey take in, but uh

26:23

I have huge disagree here.

26:24

>> Well, you got to you got to explain now

26:27

after he says something.

26:29

>> [laughter]

26:29

>> Yeah, sure. Well, yeah, I mean like

26:32

look, it's it's pretty simple, man. Like

26:34

uh

26:35

what we're seeing right now is that

26:37

there are tremendous return returns to

26:39

training compute.

26:41

I think it's pretty clear.

26:43

And you basically want to make sure you

26:45

have more than others if you want to

26:46

stay in the race.

26:48

Open source versus frontier, I think

26:49

it's pretty clear that the gap is

26:52

expanding not closing which everyone was

26:54

saying last year open source is going to

26:56

catch up or the gap is going to close.

26:58

The gap is closing. China is getting

26:59

closer. No, that's not happening. The

27:01

frontier is like beating to a massive

27:04

extent the open source models as

27:06

demonstrated by Anthropic they are our

27:08

trend.

27:09

Right like I think anyone who does

27:10

production or clothes sees the

27:13

difference between Claude and No Kimi or

27:16

Deep Seek V4. Um

27:18

if you were to take a guess, would you

27:20

imagine that the gap is going to expand

27:23

or is going to narrow? I would assume

27:24

that it's going to expand because one

27:27

has much more compute than the other

27:29

which goes back to the fundamental point

27:31

which is that you know training compute

27:32

has tremendous returns and so not having

27:35

compute means that you're disadvantaged

27:37

relative to competitors.

27:39

>> I think we're agreeing about one thing

27:41

which is that training compute has

27:43

massive returns if you're in the first

27:45

place, but not necessarily if you're in

27:48

fifth place. Like

27:49

>> Not necessarily. And I I think the point

27:52

is okay, let's imagine this. Let's say

27:54

Meta goes like completely crazy next

27:57

year. They secure um

27:59

an amount of compute so they have let's

28:01

say they have five x more training

28:02

compute than Anthropic. I guess do you

28:05

think they catch up? And they're they're

28:06

they're really far behind, but if they

28:08

have five x more compute than Anthropic,

28:09

do you think they catch up?

28:10

>> No, probably.

28:11

>> No?

28:12

>> No.

28:12

>> Obviously they need talent as well, but

28:14

>> Yeah, please go ahead.

28:15

>> I think it I think they need the talent.

28:16

And so I do think there's a potential

28:18

for I think you need both compute and

28:19

the talent basically.

28:21

Um and so I I think that like the maybe

28:24

the these go hand in hand where when XAI

28:27

gives up all of their talent or close to

28:29

it with all the co-founders leaving and

28:31

then they give up all their compute, it

28:33

kind of goes hand in hand there.

28:34

>> Yeah, I know 100% agreed, but I think

28:36

the point is these two things tell you

28:38

that they're basically out of the race.

28:41

Then it's going to be incredibly tough

28:42

for them to

28:45

to like

28:46

you know come back and uh extract value

28:48

out of the $24 trillion and write AI

28:51

applications market.

28:52

>> Okay, that came up again. So, I'm going

28:54

to pull that up on the S1 and then Troy,

28:56

maybe you can get us away from this

28:57

argument and talk about the uh

29:00

Yeah, this stuff here. Here's their

29:02

Here's their TAM. It wasn't 24, it was

29:04

22.7 trillion dedicated enterprise

29:06

application.

29:10

Little down here is Starlink.

29:12

>> [laughter]

29:13

>> That is insane.

29:14

>> Anyway.

29:14

>> Um

29:15

>> Does the race even matter though? Cuz I

29:17

feel like at a certain point if Meta

29:18

gets so much more compute and their

29:20

model gets so much better, even if

29:21

they're like number three, number four,

29:23

number five, or whatever, it's good

29:24

enough for most people to use, right?

29:26

And it's probably good enough to be

29:27

replacing a lot of jobs already. That

29:29

you don't have to be number one or

29:30

number two to be winning.

29:31

>> No. No. No. No.

29:36

>> Uh No, I Okay. I I think this kind of

29:39

comes down to your perspective on how

29:41

they actually use the models. So, like I

29:43

don't know. We're We're saying number

29:44

three, four, five here just to define it

29:46

from my perspective, which others may be

29:48

disagree with, is like Anthropic's uh in

29:51

first place right now because I believe

29:52

coding is the only thing that matters. I

29:54

think they've been proven correct.

29:56

Coding is not coding, it's computer use

29:57

and everything a human can do with a

29:59

computer and AI can do with a computer.

30:00

And so, therefore, this is interface to

30:02

the computer, not coding. So,

30:03

Anthropic's in first, OpenAI's in

30:05

second. I put Cursor in third. My

30:07

experience using Composer is

30:08

significantly better than using Gemini

30:10

or using Muse Spark because it's like

30:13

you can't use it

30:14

um and uh from using the Grok models

30:17

from xAI. And so, I I In some ways, I I

30:19

think that they have both the

30:20

distribution and the model to be in

30:24

third place right now. And the question

30:26

is just how much compute does Cursor

30:28

need to stay in the race? And I think

30:31

that having the optionality to feed them

30:32

more compute in the future is super

30:34

compelling. Whereas, they can't use it

30:36

right now. And so, why have it on your

30:37

balance sheet if you can't actually use

30:39

it? Why not turn it into a

30:40

revenue-generating asset and use it

30:42

later once you have more distribution or

30:45

once you build out the training stock to

30:46

improve it to the point where you can

30:47

actually do these hero runs. But we'll

30:49

see because there's there's it's really

30:51

interesting that there's four labs, five

30:53

labs in the US testing this theory from

30:55

different angles. Some are stacking

30:57

compute and have a bunch of revenue,

30:59

some are stacking compute and have no

31:00

revenue and some actually have quite a

31:02

bit of revenue if you look at cursor and

31:03

don't have that much compute right now

31:05

on a relative basis.

31:07

And I'd like to see all three pursue it

31:09

that way because

31:10

I'm not sure what the right playbook is

31:12

or

31:13

who the winner will be.

31:14

Um but we'll see. I mean it's going to

31:16

be interesting for Anthropic to attempt

31:18

to defend their number one position

31:20

because that's not a position they've

31:21

been in before. They've only had to play

31:23

catch up.

31:24

Um and I think that's actually quite

31:26

hard. I think it's quite hard to retain

31:27

talent. I think it's quite hard to

31:29

um keep pressing a compute advantage. I

31:32

think it's it's quite hard to motivate

31:33

users and and consumers to keep

31:35

consuming more instead of getting, you

31:37

know, grass is always greener with some

31:38

new feature from some competitor. And uh

31:41

I think it's up to them to to maintain a

31:42

trillion-dollar market cap.

31:44

We'll see.

31:45

>> This is Yeah, Jordan, I think it

31:47

>> Jeremy, you want to go?

31:49

>> I said we've been talking a lot, man. I

31:50

want other people to talk, but I had a

31:52

response for Crystal, but you go first

31:54

and then

31:54

>> it goes into like the the earnings

31:56

before, you know, training interest and

31:58

taxes. But I think it's really

31:59

interesting like, you know, we think of

32:01

ETA like before training interest taxes

32:03

is like the, you know, the cash

32:04

operating profits that they you generate

32:06

for running inference. Um and if you

32:10

want to think of training and research

32:11

as CapEx, you go back to this

32:12

conversation like I think this is a big

32:15

investor question, like even corporate

32:16

strategy question is that return on

32:18

invested capital. Um and like right now

32:21

we know like Anthropic is obviously

32:23

having massive massive returns on the

32:26

invested capital they not only put in

32:27

these models, but also like just for

32:29

coding applications, these computer

32:31

applications specifically. And then, you

32:34

know, we're seeing more and more news of

32:35

other

32:36

um you know hyperscalers uh or not

32:39

hyperscalers at their labs um trying to

32:42

get into this coding you know market. I

32:44

think there was some Microsoft news this

32:45

morning on that. Uh you know mixed mixed

32:47

opinions I think here of of how

32:49

successful that might be. Um but

32:51

training more and more models for this

32:53

um because obviously people do want to

32:55

use frontier models to Jeremy's point.

32:57

We even saw like the meta token maxing

32:59

article like

33:00

you know all that spends external. Uh

33:02

token spends external because of

33:05

you know because to Jordan's like

33:06

computer and coding point.

33:08

Uh that's where like you know there's a

33:10

ton of obviously product market fit and

33:12

they're seeing their own ROI when they

33:14

use the product.

33:15

Um

33:17

but yeah, that's heavy. I I'm guessing

33:19

Jordan's still on the call. So Jeremy

33:20

I'll I'll send it back over

33:22

>> Yeah, yeah.

33:22

>> to you.

33:23

>> J- just wanted to respond to like

33:25

Crystal's point because

33:27

yeah, I mean it obviously depends like

33:29

how you think of the the market evolves,

33:32

but

33:33

I really like the micro framework that

33:35

Malcolm has which is that okay, I think

33:37

like you know 2030 2035 like uh what

33:41

type of uh tasks are going to drive the

33:43

bulk of the total addressable market,

33:45

you know the dollars that people are

33:47

actually spend on AI. And there's tasks

33:49

that are like uh have a certain amount

33:51

where they're good enough that's kind of

33:53

finite like translation maybe at some

33:55

point there's only time it doesn't make

33:57

any sense uh to spend more on

33:58

translation, but then there's like these

34:00

very open-ended task where the spending

34:03

is pretty much infinite. And then on

34:05

legal for example, you could assume that

34:07

if AI is really good at at legal, if you

34:09

want to make sure you beat your

34:10

competitor, you want to probably want to

34:12

spend more on AI than he does, have more

34:14

intelligence than he does uh because you

34:16

want to gather more evidence, you want

34:17

to think through many different ways of

34:20

uh you know coming up with a defense and

34:21

whatnot. Um scientific research research

34:25

is typically the very open-ended use

34:26

case. Um you know health care um for us

34:29

as analysts, we try to get insights from

34:32

gathering a lot of data, very

34:33

open-ended. Um I would assume these use

34:35

cases are going to drive much more

34:37

spending than this the finite sort of

34:40

it's good enough, right? And and and

34:42

really when you think of these

34:43

open-ended use cases, what what matters

34:45

to be able to do like what you want to

34:46

do uh

34:48

uh in the cheapest way. And if you're

34:50

only doing the cheapest way, you're

34:51

going to have to use frontier models.

34:53

It's the whole point we made we made

34:54

like Mythos is actually 2 first cheaper

34:57

than Opus.

34:58

The model in terms of token pricing is 6

35:01

times, I think, 5 times more expensive

35:03

than Opus.

35:05

But you know, if it's like 10 times

35:07

smarter, if it requires 10x less tokens

35:10

to answer a given task, then it's

35:12

actually way cheaper to complete that

35:14

task with the the model, right? And so I

35:16

think like at least that's that's kind

35:18

of the way I view it. And so

35:19

I I I think the bulk of the market is

35:21

going to concentrate to

35:23

the frontier models. I think if you're

35:26

three or four or five, you're not going

35:27

to get any dollars. And I think like if

35:29

you take step back and think, "What are

35:30

the signals that we've seen in 2026? Did

35:33

the AI market has it beaten or has it

35:36

missed versus expectations we had in

35:38

2025?" You know, massive beat. But then

35:41

what I think is super interesting is

35:43

what is the composition of this beat?

35:45

Did everyone beat or is it just a few

35:47

companies, right? And I I think that is

35:49

actually kind of interesting. On the

35:51

frontier side, it's basically one

35:52

company, it's just Anthropic, massive

35:54

beat.

35:55

Google is probably tracking behind to

35:56

some extent when you look at just the

35:58

Gemini adoption and how much people were

36:00

spending on Gemini. OpenAI is tracking

36:02

behind. Obviously xAI, Meta are nowhere

36:04

to be seen. You could have hoped that

36:06

they would have had something. They

36:07

don't really. To be fair, on the

36:09

open-source side, I think that it's also

36:11

been a beat. I think there's been some

36:12

good adoption, but the dollars uh spent

36:14

are still pretty small. So I think it's

36:16

an interesting composition of massive

36:19

beat where it's basically all driven by

36:20

one company, which kind of gives gives

36:22

you a sign that it's pretty much winner

36:24

takes all. And if you're state of the

36:25

art, you get the bulk of the volume, and

36:27

if you're not,

36:29

yeah, people don't spend on you. I know

36:31

Joey stopped playing. What do you think?

36:33

>> Can you repeat the last part of that?

36:34

>> Holy [ __ ] you didn't listen to my

36:35

beautiful prose. I'll jump off.

36:37

>> I heard that I heard the most of it.

36:39

>> I got to jump off and respond to this.

36:40

So, Jeremy, I I think you're this is

36:42

totally this totally makes sense, but

36:44

we're also seeing massive beats or

36:48

massive reported ARR numbers from

36:50

startups serving open-source models

36:52

right now. It's not just Anthropic

36:54

growing.

36:55

There's there's huge growth for

36:56

Fireworks that's tied to Cursor.

36:58

>> that big, man.

36:59

>> Really big.

37:00

>> They're not that big, man. That is the

37:01

thing. They're they're not that big.

37:03

Some encouraging signals. I think Cursor

37:05

is at what? 2 2 billion now?

37:07

Uh and they were at what? Maybe

37:08

1.something at the end of 2025. There's

37:11

still like really good growth.

37:13

Definitely you could say Cursor is a

37:14

beat. I don't know WinZO, I guess,

37:16

they're in Google, but they're worth to

37:17

be seen.

37:18

>> Yeah.

37:18

>> No, but I I I I I agree with you. Some

37:20

of the open-source guys have had a beat,

37:21

but in terms of like dollar amount,

37:23

still like not very meaningful.

37:24

>> Yeah, I mean the claim from Fireworks is

37:26

$315 million of ARR, right? Like that's

37:30

>> What is 300 million between between

37:32

friends?

37:33

>> [laughter]

37:35

>> Okay. Like okay.

37:36

In the past, a startup unicorn was

37:39

interesting when they were a

37:41

billion-dollar valuation. Now they start

37:43

to approach billion-dollars ARR, and you

37:45

go, "Ah,

37:47

ah, whatever. Fly on the wall. Like

37:50

>> Relative to the size of the market, you

37:52

know, which is like already above 100

37:54

billion dollars. It's like in the grand

37:56

scheme of things, not that big. Think of

37:58

the the positioning of the different

38:00

hyperscalers. I mean, we've talked about

38:02

we've talked about like, you know,

38:03

business models and whatnot,

38:05

but the beauty of the the economics 2.0

38:08

this

38:09

magic model is just so accurate,

38:12

is that it covers

38:12

>> Yeah,

38:14

I mean, you know, in 2 weeks we've

38:16

gotten some pretty good feedback so far

38:17

from some of that

38:19

you know, these you know, hyperscalers

38:20

customers, things like that. so um it's

38:23

been it's been solid, but I mean, I

38:25

think right now like just given like

38:27

Amazon continues to win.

38:29

Yeah, I think the things that benefit

38:30

them you know, around their customer

38:31

base, you know, benefit you know, Azure

38:34

you know, pretty similarly.

38:36

Especially as you know, you look at you

38:38

know, more and more you know, token as a

38:40

service type models that come on the

38:41

foundry.

38:43

And then obviously I think you know,

38:44

Google if they get a coding model you

38:46

know, right now you know, when we look

38:47

at what was formerly Vertex and is now

38:49

Gemini H enterprise platform.

38:52

You know, we we still think you know,

38:53

Gemini API is a pretty decent percentage

38:55

of that. So we're not benefiting a ton

38:58

from

38:59

you know, Claude and some of those

39:00

things, but like that's probably the

39:01

biggest thing right now. I mean, we

39:02

still see Amazon like Bedrock you token

39:05

as a service is a pretty you know, I

39:08

think by the end of the year it could be

39:09

the majority of the AI business

39:11

at Amazon and even though Amazon like

39:13

lags Google and or GCP and Azure like as

39:18

you know, their AI mix. Yeah,

39:20

at at AWS AI is a lot smaller percentage

39:23

of the business. But with Bedrock going

39:25

to the majority of the AI business and

39:27

then infrastructure as a service being

39:29

you know, 80 90% at Azure and GCP.

39:32

Um

39:33

you know, it's really really

39:34

advantageous. But then too to that

39:37

point, it's a very easy for Azure to

39:39

then come in you know, add Claude as a

39:41

model.

39:42

Um and kind of you know, this token as a

39:45

service business isn't that hard for you

39:47

know, these people with massive customer

39:48

bases to implement. That's much harder

39:50

as you go down to like Oracle.

39:52

Um

39:53

and then to the neo clouds like like to

39:55

implement this at the same level just

39:56

given they don't have you know, the

39:58

massive inertia customer bases.

40:00

Um

40:01

that you know, really benefit from like

40:03

the old school software distribution

40:04

enterprise software distribution like

40:07

you know, emotes.

40:08

Um

40:08

>> Yeah, let me throw

40:10

let me throw this chart on on screen

40:12

just to make the point Jeremy was making

40:13

and you're making right now which is

40:14

like bunch of rounding error for

40:16

anything but the top three.

40:19

Hyperscalers when it comes to the

40:20

inference play business, right?

40:23

Um yeah, it's uh

40:25

>> [laughter]

40:26

>> Yeah, very small portion of the market

40:28

that's going to ever to everybody else

40:29

bucket.

40:30

>> And it's also I mean that that kind of

40:32

goes to the to the same thing the same

40:34

point I was mentioning earlier with

40:36

regards to like the composition of the

40:37

market.

40:38

One of the important point he makes in

40:40

that article is that um

40:43

the key to being a successful token as a

40:45

service business

40:47

is that actually just to have

40:48

partnerships with the big labs and have

40:50

access to frontier models. It's pretty

40:52

simple, right? So, I guess it's uh the

40:55

big disadvantage that currently Nebius,

40:57

CoreWeave, Iron, all of these guys have

40:59

is they don't yet have the partnership

41:02

and they also don't have maybe the

41:03

capital to be able to deploy GPUs

41:06

without a 5-year contract. Um so, it's

41:09

kind of like as a as a function of the

41:10

way the market works that a company like

41:12

CoreWeave or, you know, Nebius have the

41:14

bulk of I mean maybe not Nebius

41:15

actually, but Iron have the bulk of

41:17

their business contracted over multiple

41:19

years. Um whereas Amazon, you know, is

41:22

free to be more spec

41:25

and better pick more demand that don't

41:27

have that 5-year off-take uh take or pay

41:29

locked in.

41:29

>> Yeah, well

41:30

>> And that I think too, Jordan, I think on

41:31

like who wins is is definitely the

41:33

custom silicon portion.

41:34

Um You know, around Tranium and you

41:37

know, TPUs at GCP that I think is a big

41:39

thing versus like in the tokenomics

41:40

model that and then when we look at like

41:42

brain accelerator in

41:44

um and some of the data center model

41:45

numbers

41:46

you know, it's really at Azure you know,

41:47

just as mostly in video and then you

41:49

look at that vertical integration at

41:52

um

41:53

AWS and GCP, like that's another big

41:55

advantage for them.

41:57

Uh that that I think, you know, probably

41:59

you know, especially on the margin side,

42:01

I should say. And then as we've seen a

42:02

lot of, you know, inference get more

42:04

efficient and you know, gross margins on

42:05

inference

42:07

you know, drastically improve at the

42:08

labs, you know, the two you know,

42:10

frontier labs, major frontier labs over

42:12

the last two years like that's

42:14

definitely been

42:15

you know, another like key consideration

42:17

when you think about who's going to win

42:19

you in this market.

42:20

>> Yeah. Yeah, I mean that it's from the

42:23

technical perspective, everything that

42:24

we've criticized all the chip startups

42:26

about and TPU and Trainium is always

42:28

about usability. But if the entire chip,

42:31

you know, a gigawatt of Trainium at

42:34

Rainier is all just serving tokens from

42:38

one to three models. Well, the customer,

42:42

the end user customer doesn't even

42:43

necessarily need to know that much about

42:45

what chip is running if they're only

42:46

buying tokens.

42:48

And certainly not the end users, like

42:50

the actual terminal user of the tokens,

42:54

you know, when I'm using Cloud Code, I

42:56

have no understanding of whether my

42:58

token is coming from a TPU or a Trainium

43:01

accelerator or a GPU. Doesn't make a

43:03

difference, right? Okay, guys, well,

43:05

really interesting stuff today. Anything

43:07

left unsaid on the the topic of

43:11

uh Bedrock, tokens as a service,

43:15

Anthropic's growth.

43:16

>> Hey, I think we got it all, Jordan.

43:17

Winners win, losers lose, and it was a

43:20

clear trend.

43:21

>> [laughter]

43:23

>> Winners win. Awesome. Okay. Well, thanks

43:25

guys for coming on the show. Good job

43:26

today.

Interactive Summary

This episode of SemiAnalysis explores the growth of Anthropic and the significant impact of Amazon's Bedrock platform on AWS's margins. The discussion highlights a shift in the cloud market from traditional infrastructure-as-a-service to a 'token-as-a-service' model, which offers superior margins for hyperscalers. The participants analyze why Anthropic is succeeding, the importance of coding and computer-use workflows, and the competitive dynamics between frontier AI labs, including the impact of compute scarcity and the strategic decisions of companies like xAI and Meta.

Suggested questions

3 ready-made prompts