HomeVideos

Anthropic's Katelyn Lesse & Angela Jiang: Building an Ecosystem, not a Walled Garden

Now Playing

Anthropic's Katelyn Lesse & Angela Jiang: Building an Ecosystem, not a Walled Garden

Transcript

1650 segments

0:00

The last layer of abstraction on top of

0:01

this is probably the coordination layer.

0:03

So, you have knowledge, and you have

0:04

execution, and you have coordination.

0:05

And at the coordination layer, we're

0:07

beginning to think of these things

0:08

called like strategies, where basically

0:10

it's almost like a meta harness. The

0:12

true low-level harness is designed for

0:13

execution, but the next one is about,

0:15

okay, if tokens aren't really fungible,

0:17

and you need to give them different

0:18

jobs, like maybe some this token is

0:20

advising versus this token is executing,

0:22

you want to start composing these like

0:24

these kind of orchestrated strategies

0:25

that go together, and they should sit on

0:27

top of all these things because at the

0:28

end of the day you still need to execute

0:29

and the execution still needs to know

0:30

what to do. So, everything in theory

0:32

should kind of like ladder together. And

0:34

so, I think, you know, if you were to

0:35

look at our road map and the maybe kind

0:36

of project forward a little bit where

0:37

you kind of expect us to go, we'll move

0:39

more and more from the knowledge layer

0:41

to the execution layer and from the

0:42

execution layer to the kind of

0:44

coordination layer in terms of the

0:45

abstractions that you [music] can see us

0:46

put out.

0:53

>> [music]

1:04

>> Caitlin and Angela, thank you so much

1:06

for joining us today. Lauren and I are

1:07

thrilled to have you here. You are

1:09

responsible for building Anthropic's

1:10

platform, and so you're responsible for

1:13

building what I think is one of the most

1:14

important, if not the most important

1:16

developer platform in the world. And we

1:18

are really excited to interview you

1:20

today to understand more about what's

1:21

ahead. And so, maybe just to get

1:23

started, can you give us the context of,

1:25

you know, what is Anthropic platform and

1:26

where do you sit within Anthropic?

1:28

>> Yeah, so platform is both our externally

1:31

facing APIs, our developer platform that

1:33

people build on top of when they want to

1:34

build applications and system that

1:36

systems that access Claude's

1:38

intelligence, as well as internally, um

1:41

we run our product infrastructure, and

1:43

basically we're the layer that our apps

1:46

build on top of uh internally as well.

1:49

>> Awesome. What's your North Star as a

1:50

team?

1:51

>> It's a great question. We actually,

1:52

because we have both internal and

1:53

external, we actually kind of have like

1:55

two North Stars, which is probably like,

1:57

you know, you'd be like, "What? You

1:58

shouldn't have one North Star." But, um

2:00

no, we we can

2:01

>> planetary system.

2:01

>> Yes, exactly. There's separate solar

2:03

systems, so it's fine. Um but, uh on the

2:05

internal side, like we really want to

2:07

provide is like literally as much

2:08

leverage as possible for our internal

2:10

teams to be able to ship like AGI-pilled

2:14

products.

2:15

Um and we want them to be able to move

2:16

fast, be able to have reliable, like

2:18

great, like uh platform uh to be able to

2:21

build on top of. But, I think that key

2:22

bit about speed is like really

2:24

intentional for us, and we really really

2:26

care about that internally. Externally,

2:28

um we actually have a lot more like

2:29

complicated set of things. Um but, one

2:31

of the true Norths that we have there is

2:33

to be able to basically give any builder

2:35

the tools to be able to work with Claude

2:37

to build whatever they want to build.

2:39

And so, it's a bit of a broad statement,

2:41

but as a result of that, uh boils its

2:43

itself down into, you know, being

2:45

wherever that business is. Like, we

2:46

really care about like bringing our

2:48

platform really really close to that

2:49

business. This is why we spend a lot of

2:51

time with the hyperscalers, integrating

2:52

really closely uh directly with them,

2:54

like AWS, Google, so on and so forth. Um

2:57

and it is a lot of like primitives that

2:59

we end up creating. We want people to be

3:01

able to express what they think their

3:03

product should be. We want them to be

3:04

able to almost do like custom software

3:07

in their own way, you know, like in this

3:08

new world with AI, uh what used to be

3:11

probably economically impossible was

3:13

that last mile of of custom software.

3:15

Now, in theory, should be like very very

3:18

achievable. Um and we want to give them

3:19

all the tools and all the capabilities

3:21

to go and do that. And so, sometimes

3:22

that comes in form of primitives and

3:24

APIs and higher-order abstractions, and

3:26

sometimes that comes in the form of just

3:27

like

3:29

standards. Uh so, for example, like

3:30

skills and MCP, um those are things just

3:33

like Claude needs them to be useful, and

3:35

we can just give them out to the rest of

3:36

the ecosystem, work with everyone to

3:38

help you create those things, and get

3:40

the best out of Claude. So, um I would

3:42

say externally, you know, we really are

3:43

oriented around just helping you just be

3:45

able to build, but internally, that

3:47

orientation, while still existing, is

3:49

probably more, you know, sp- specified

3:51

towards speed and being able to move

3:53

really quickly.

3:54

>> How do you decide what goes into the

3:56

platform, what gets externalized, and

3:58

what doesn't to decide what products

4:00

should be available?

4:01

>> Yeah, I mean, we generally try to have a

4:02

philosophy that we try to be consistent

4:05

across the board. It's actually one of

4:06

the reasons why we do internal and

4:08

external.

4:10

There's plenty of other, you know,

4:11

platform businesses and constructs where

4:13

you actually like bifurcate these two

4:14

things.

4:15

For us, we kind of try to intentionally

4:16

keep it equal and then as a result we

4:18

try to hold this philosophy as much as

4:20

we can around like, you know, for any

4:22

builder internal or external, even

4:24

though if our internal builders might

4:24

have some slightly different

4:26

requirements in the same way any user

4:27

would have slightly different

4:27

requirements.

4:29

We want to have the same primitives that

4:30

are available to everyone. And one of

4:32

the maybe the overarching thesis for

4:33

that is that we've just seen like the

4:36

capabilities of these models just grow

4:37

and that's just exponential. It's really

4:39

hard to figure out like a long-lasting

4:41

form factor. I think two years ago we

4:42

were all like everything's chat and now

4:44

everyone's like forget chat. Now you

4:45

just like agents. And like there's going

4:46

to be another form factor, another form

4:48

factor.

4:49

And we kind of imagine that like

4:51

constantly evolving. And so the best way

4:53

for us to kind of enable that for

4:55

everyone and also ourselves is to

4:56

actually build a really robust platform

4:58

that gives people those kinds of like

4:59

tools to figure out what those form

5:01

factors are. And I don't think we by any

5:03

means feel like we're the only ones

5:04

capable of figuring out that form

5:05

factor, like not at all. In fact, the

5:07

more democratization we can do on that

5:08

and help people and allow people to

5:10

experiment, I think the the more those

5:12

form factors will actually kind of

5:13

naturally come out of the market.

5:15

>> Yeah, and I think within our team we've

5:17

we've had moments where we're

5:18

experimenting even with just like a

5:20

packaging up of our primitives in a

5:22

different sort of higher order way and

5:24

we've thought about, okay, cool, we've

5:26

solved this exact type of problem with

5:28

this product that we built into the

5:30

world. And so we can go and dog food it

5:32

for ourselves, but we never want to fall

5:33

into this trap of like we're over

5:35

indexed on the problem as it needs to be

5:37

solved for an internal user like because

5:40

exactly what Angela said, internal users

5:42

have very specific requirements,

5:43

external users have very specific

5:45

requirements. And so if you over index

5:46

on one or the other, you fall into a

5:48

trap. So, a lot of the time what we'll

5:50

do is dog food something internally at

5:52

the same time that we open up early

5:54

access of some sort with external

5:56

customers so that we can kind of get a

5:57

range of feedback and bring those things

5:59

back into the platform.

6:00

>> I'd love to talk about the higher levels

6:02

of abstraction that you discussed. So, I

6:04

guess at the base level, this is just,

6:05

you know, raw access to Claude, Opus, or

6:08

whatever tokens.

6:10

How do you think about the I guess the

6:11

layer cake of abstractions above that?

6:14

>> Yeah, if you look back, um so when I

6:16

joined Anthropic around a year ago, um

6:19

the platform was basically just the

6:21

Messages API. There was a Messages API,

6:23

um you know, we had come out with

6:25

standards like MCP. We obviously have

6:27

developer tooling around our SDKs and

6:29

our docs and our console and things like

6:31

this, but for the most part it was a

6:32

stateless API. Um and what's interesting

6:35

to Angela's point on foreign factors

6:37

evolving over time is we found a lot of

6:40

our customers solving the same problems

6:42

over and over again that we also were

6:43

solving over and over again around as

6:46

the models got better at running for

6:47

longer and working with more context at

6:49

a given time, you want to build agents

6:52

that can succeed in a kind of

6:54

long-running context and even a remote

6:56

context that doesn't necessarily have a

6:58

human in the loop. And so, we found that

7:01

we could piece together our primitives

7:03

and stand up all the same infrastructure

7:05

that we're finding ourselves standing up

7:06

internally to power our own products and

7:08

arrive at some higher-order abstractions

7:10

that let you do more agentic work out of

7:13

the box. And the problems that we're

7:14

solving for you are, you know,

7:16

infrastructure being kind of a hard

7:18

thing to deal with. Like, how do you

7:20

figure out spawning sandboxes that are

7:22

going to have the right governance and

7:23

security and like, you know, spin them

7:25

up and spin them down when you need to

7:27

or the storage around transcript

7:30

sessions so that you can resume a

7:31

session if you stop it and pick it back

7:33

up later. Um so, that infrastructure is

7:35

a big thing that we wanted to be able to

7:36

provide more of out of the box and we do

7:38

more of that today. Um and then the

7:40

second thing just being harnesses and

7:42

harness engineering. There's a lot of

7:44

thought and energy going into how do I

7:46

do my prompt caching and how do I manage

7:48

my context window as well as how do I

7:50

actually just get more intelligence out

7:52

of the model, um, and how do I manage my

7:54

costs and things like that. So,

7:57

we've kind of packaged up our primitives

7:58

a bit more in tune with the problems

8:01

that we found ourselves solving to

8:02

provide more of these things out of the

8:04

box for people so that they can, if

8:07

they're building systems for them

8:08

themselves internally, if they're

8:09

building products, they can just be more

8:11

focused on the problems that they want

8:12

to be solving and if they want to

8:14

offload some aspects of those problems

8:16

to us, they can. Um, and that's kind of

8:19

the ethos.

8:19

>> And are your customers generally

8:20

choosing to opt from the grab bag of

8:22

stuff that you offer or are they like

8:24

how how often are they opting into the

8:25

just the the managed agents offering, I

8:27

guess, just take care of it all for me?

8:29

>> Um, it varies by like the the user

8:32

group. So, like for I would say, you

8:34

know, like really AI native startups,

8:36

like the ones who are like tinkering and

8:38

like experimenting at a really low

8:39

layer, they're just going to go for the

8:40

primitives.

8:41

Um, and then for everyone else, uh,

8:43

these are kind of classic, like more

8:44

like enterprises or areas where it's

8:46

like the purpose of the startup or the

8:48

philosophy behind the startup isn't

8:49

necessarily to optimize on, um, some

8:51

kind of hill climbing piece, it's more

8:52

like stringing together a bunch of

8:53

workflows and, you know, providing

8:55

unique user value at, uh, to that user.

8:57

For those people, um, you know, it's

8:59

just kind of not their core competency.

9:01

It's not what they want to focus their

9:02

time and resources and they reach much

9:03

more for these kind of like

9:04

higher-ordered like package offerings.

9:07

>> What are some examples of the primitives

9:08

you've released at different layers in

9:10

the last few months? We've seen a few of

9:11

them. Would love to hear.

9:13

>> Yeah, I think maybe one framing I would

9:14

give, um, for some of the constructs

9:16

that Cayden was talking about is like,

9:18

and this is a bit of an

9:19

oversimplification, but effectively

9:21

there's approximately like three like

9:23

layers of this cake. At the very bottom

9:25

is just kind of like like knowledge. And

9:28

so at this layer, like in many ways it's

9:30

it's knowledge about the model, it's

9:31

knowledge about the things that the

9:32

model needs, and it's just like the

9:33

ability to know how to actually do

9:35

something with Claude is maybe the way

9:36

I'd phrase that. And so there, the

9:38

primitives that we have spent more and

9:40

more time on have been actually things

9:42

of the past. Cuz like we still evolve

9:44

them, but they tend to be a little bit

9:45

more baked. Like for example, there's

9:47

very specific shapes and parameters we

9:48

put on the messages API, and it's more

9:51

like trying to expressly like showcase

9:54

Claude's like design.

9:57

Like Claude the model's actual design,

9:59

the way it thinks, the way it respects

10:00

certain parameters, the way it kind of

10:02

like

10:03

will do tool calls. Like all of those

10:05

different pieces. And then we started to

10:07

standardizing like tools, and then we

10:08

started standardizing bits and pieces of

10:09

like context that you could put in at

10:11

different moments in time, which is

10:13

concretely like skills and like memory.

10:15

And so those are like the kind of like

10:16

knowledge layer type of abstractions

10:18

that we've put out over the past I guess

10:20

like year plus plus a bit.

10:22

Um the next layer of abstraction that

10:24

we've actually started to spend more and

10:26

more of our time on is like once you

10:27

kind of know stuff, you then need to

10:29

like execute. And so at the execution

10:31

layer, that level of abstraction is the

10:33

part that Kaylin was talking about

10:34

around like we're doing these like

10:35

higher order pieces, but like what are

10:36

we putting higher order there? It really

10:38

is because you're now getting Claude to

10:40

execute work. It's not just to know

10:42

something, right? I can give it a

10:43

question, it'll give me an answer. You

10:45

can push string a lot of that stuff

10:46

together.

10:47

But now if you need to execute, like do

10:49

work, give me the output, edit files in

10:51

a bunch of different systems, that

10:53

becomes a lot more complicated and

10:54

requires infrastructure to handle. And

10:56

so that layer is basically I would say a

10:58

low-level harness plus managed

10:59

infrastructure as like the set of

11:01

abstractions. Today we just like our

11:03

high-level product for that is called

11:04

Claude managed agents.

11:06

And so that's like a piece, but we

11:07

started to wrap more and more pieces in

11:08

that. Um I think there's going to be a

11:10

layer like on top of that. We have like

11:12

some inklings of it if we started to

11:13

build towards, but the last layer of

11:15

abstraction on top of this is probably

11:16

the coordination layer. So you have

11:18

knowledge, and you have execution, and

11:19

you have coordination. And at the

11:20

coordination layer

11:22

we've started to expose some of these in

11:23

ways that like aren't very obvious, but

11:25

we're beginning to think of these things

11:27

called like strategies, where basically

11:29

it's almost like a a harness. Right, the

11:31

harness, the true low-level harness is

11:33

designed for execution.

11:34

But the next one is about, okay, if

11:36

tokens aren't really fungible, and you

11:38

need to give them different jobs, like

11:39

maybe some this token is advising versus

11:41

this token is executing, this token is

11:43

dreaming versus this token's executing,

11:44

so on and so forth. You want to start

11:46

composing these like these kind of

11:48

orchestrated strategies that go

11:49

together, and they should sit on top of

11:51

all these things because at the end of

11:52

the day you still need to execute, and

11:53

the execution still needs to know what

11:54

to do. So, everything in theory should

11:56

kind of like lad- ladder together. And

11:58

so, I think, you know, if you were to

11:59

look at our road map and the maybe kind

12:00

of project forward a little bit where

12:01

you kind of expect us to go, we'll move

12:03

more and more from the knowledge layer

12:06

to the execution layer, and from the

12:07

execution layer to the kind of

12:08

coordination layer in terms of the

12:09

abstractions that you can see us put

12:11

out.

12:11

>> That's a really cool road map.

12:13

>> How do you think this all comes together

12:15

into a broader ecosystem beyond just the

12:17

things that you guys are building? How

12:19

do you help support people building

12:21

products on top of it, and how do you

12:22

help them get the most out of all these

12:24

pieces?

12:25

>> Yeah, I think this is like super top of

12:27

mind for us. Like, we really want to

12:28

find a way to be

12:30

to support as many people in doing this

12:32

as we can. I think we're still like

12:34

learning. Like, a lot of the industry

12:35

like has evolved, we've seen, you know,

12:37

a lot of different pieces um get spun up

12:40

and spun down. And I think the the

12:43

operative part for for Caitlin and I has

12:45

been in the category of like making

12:47

sure, at least at the base layer, that

12:49

we provide as many primitives across the

12:51

board as possible. So, you know, this

12:53

kind of like yeah, like knowledge,

12:55

execution, coordination layer, we want

12:56

to give all of that out to everyone so

12:58

that people can start to compose and

13:00

create on top of that. Um and that's

13:02

just from a I think pure builder kind of

13:03

point of view. Then there's a point of

13:05

view around like, how do you kind of

13:06

like plug in with us? Right? Like, we're

13:08

also building first-party products of

13:10

our own. We've also created some ways to

13:12

embed natively with us, like for

13:13

example, connectors which are built on

13:15

top of the MCP spec. Um and we try to be

13:18

more open about those types of things.

13:19

And we're starting to figure out like

13:21

what are the right bits and pieces, but

13:22

I mean, what we're really trying to do

13:24

is get to a place where, you know, in a

13:26

a company is able to get creative and

13:28

build on uh they can build whatever

13:30

products that they want. They can build

13:31

agents if they need to and then those

13:33

agents and those products could be

13:34

things that could plug into other

13:36

agents. Some of those agents could be

13:37

cloud agents, some of those agents could

13:38

be other people's agents.

13:40

But we want to be able to enable that

13:41

kind of like trans-actability across the

13:43

board. And then I think in order for all

13:45

of that to kind of ultimately be true,

13:47

there is a bit around like standard

13:49

setting. And I think there's the

13:51

traditional standard setting which is

13:52

around, you know, how do systems

13:53

interoperate? Um and that's uh you know,

13:56

things that you've kind of seen us do

13:56

with like skills and MCP.

13:59

But they're at again like the builder

14:00

layer. I think at a higher order layer,

14:01

there's also a bit around

14:02

interoperability and standard setting

14:04

around how do we all kind of like treat

14:06

safety together? And you know, we've

14:08

talked to a lot of these companies and

14:09

this is less from, you know,

14:11

philosophies aside, just more like no

14:13

one really wants to have technology

14:15

that's like, for example, like doing

14:17

negative things on on their service,

14:20

right? Uh so cyber I think is a great

14:21

example of this. You want to protect

14:23

your own systems from like negative

14:25

actors or bad actors. And so like these

14:28

kinds of like standard settings of like

14:30

how can we find ways to partner with

14:32

more and more people to be like, yeah,

14:34

we all kind of want to make sure our

14:36

critical infrastructure is good. We all

14:37

want to prevent like fraud or any of

14:40

those things from happening and how can

14:41

we work better with each of these

14:43

members? I think on the last layer,

14:44

we're still kind of like we're still

14:46

evolving and I think we're still very

14:47

much like trying to find ways that we

14:49

can be better and work with the rest of

14:50

the industry to bring people along and

14:53

and work with them.

14:54

Um but those are kind of like, you know,

14:55

the higher order primitives or pieces

14:57

that we wish to kind of like be in

14:59

place. Um so they can work with folks to

15:01

to ultimately solve this. I think if I

15:03

would like take a step back at the end

15:04

of the day on on all of these things,

15:06

um

15:06

you know, like this technology is so

15:08

transformative

15:10

and if it's a little bit like

15:11

electricity in the sense like before

15:13

electricity, there was just like, you

15:14

know, you had to like have a candle and

15:16

it was like you can only do so many

15:17

things.

15:18

Um

15:19

but with electricity, the reason why

15:21

it's such a transformative technology

15:22

for all of us and so greatly of a

15:24

utility is because you can actually like

15:26

wire it into everything. Everyone is

15:28

able to actually access it. We also have

15:29

like standards and ways to plug in and

15:31

do all the pieces that we need. And

15:33

that's not something that anybody can do

15:34

by themselves. They always have to work

15:36

with the ecosystem and work with

15:37

partners um to figure out a path

15:38

forward.

15:39

>> How do you think about the philosophy of

15:41

building an open ecosystem

15:43

uh versus a walled garden? And you know,

15:46

how do you think about what products are

15:47

really important for you to own first

15:49

party versus

15:50

where you're perfectly happy to plug

15:52

into other components of the ecosystem?

15:54

>> Yeah, there's So, maybe in using

15:57

Angela's kind of layered cake that we

15:59

talked about a little bit earlier,

16:00

you'll see that on some pieces of this

16:02

like execution for example, um what

16:05

we've done within something like Cloud

16:06

Managed Regions. And I think over time

16:08

you'll see us try to make this a little

16:09

bit more modular. We actually aren't

16:11

precious about you should run these

16:13

things on our infrastructure. Like it

16:15

should be sandboxes that we control or

16:17

it should be a storage layer that we

16:18

control. Um Well, we actually like for

16:21

example, we launched self-hosted

16:23

sandboxes and we partnered with Moodle

16:25

and Vercel and Cloudflare and a bunch of

16:27

other folks um even like Amazon's new

16:30

micro VMs um to have a first-class

16:33

offering where you can go plug any of

16:34

those things in. Um

16:37

We launched MCP tunnels so that you can

16:38

call out to your MCP servers that are

16:40

behind your firewall, right? And um be

16:43

able to punch through there. And so, for

16:45

some of these things, we you know, the

16:47

weather whether it runs on our

16:48

infrastructure versus somebody else's

16:50

infrastructure is actually not important

16:51

to us cuz the thing that's important to

16:53

us is more that the architecture of how

16:55

you put together these agents in a way

16:57

that will be powerful, in a way that

16:59

will be reliable and scalable. Um We

17:02

have strong opinions on that, and you

17:04

can kind of just conform to the

17:05

interfaces that we put out there and

17:07

plug those things in. Um And we think

17:09

that that generally is a thing that

17:11

works really well.

17:12

>> Yeah, I think on the the kind of like

17:14

verticals where we we build products, um

17:17

you know, I think we we kind of have

17:19

like two frames here. The first one is

17:21

we are always trying to figure out a

17:22

form factor, like an evolving form

17:24

factor. We, by the way, don't think form

17:26

factors are like static. It's like a

17:27

dynamic thing. So, what might be awesome

17:29

for 1 year's worth of AI development

17:31

will probably not be awesome for the

17:33

next year's worth. And we just kind of

17:34

try to have that mentality. We tell the

17:36

team uh just overall like around

17:38

Anthropic, everyone's always trying to

17:39

be like, "Is this AGI pilled enough?" Um

17:41

and then we always have this mentality

17:43

of like, you know, we build something,

17:44

it works. It was cool for a a year and

17:46

maybe it's not the right next thing. And

17:48

so, throw it away, try again. Um and we

17:50

we tell like platform users the same

17:52

thing. Um

17:53

We just think that's probably just like,

17:54

you know, attached to the technology.

17:56

But so, yeah, one one principle is like

17:57

trying to always constantly find this

17:58

new form factor. So, sometimes we'll

18:00

like launch products in certain areas to

18:01

try to showcase a new type of form

18:03

factor. Um it's not necessarily cuz we

18:05

think it's like the biggest TAM or the

18:07

most important thing to go after, but

18:09

sometimes like, "Okay, this is like

18:10

always been a really difficult thing and

18:11

people have always communicated this way

18:13

or tried something this way." And so,

18:15

can we show that maybe there's a

18:16

slightly different way? Um and because

18:18

the model capabilities are are so

18:20

advanced now, can we try to express it a

18:22

bit differently? Um

18:23

>> What's an example of that?

18:25

>> Yeah, you know, like uh

18:27

Claude design is a little bit of of that

18:29

way. I think depending on how you

18:30

squint, you might see it as like a way

18:32

that we kind of are going into design as

18:34

as as like, you know, one of the

18:35

verticals. But more often than not, it's

18:37

like if you take a look at what we're

18:39

trying to do with that product, there's

18:40

a couple of like decisions that were

18:42

made in there. The first one is that

18:43

like you can actually try to offload

18:45

more and more and more to Claude.

18:47

Um and so, it tries to be kind of

18:49

opinionated on like, you know, just just

18:50

like talk to it and like let it really

18:52

try to figure out. And yes, you can

18:53

still edit it and then do these kinds of

18:55

things, but kind of like discourage a

18:56

little of that and more just like let

18:58

just talk to Claude to go figure it out.

19:00

Um the second thing was it was really

19:02

trying to express that actually like

19:04

code is a is a a way to solve for things

19:07

that you wouldn't normally think would

19:08

be the way. So, a lot of people who have

19:10

built kind of generative um you know,

19:12

like slide decks or designs or whatever,

19:14

um, we'll pick uh the way of like they

19:16

have like some kind of design system,

19:18

you integrate against design system.

19:19

It's almost in the traditional like a

19:20

classic WYSIWYG style of designing

19:22

something. And with like cloud design,

19:25

it was like, "Okay, can we try to just

19:27

like use code purely, but have Cloud

19:30

generate that code, and would it like do

19:32

a good job?" And we found through some

19:34

experiments early on, it's like,

19:35

"Actually, it looks like it can kind of

19:37

do that. And how can we kind of showcase

19:38

that uh to the world?" So, that's like

19:41

an example. We have a lot of other

19:42

internal projects, and this kind of

19:43

falls in the category of like expressing

19:45

form factor. We'll all try it out

19:46

internally. It'll be super cool for like

19:48

2 weeks, and then we move on to the next

19:50

thing. Um, we never even ship the thing,

19:51

frankly. But yeah, we actually do a lot

19:53

of product experimentation in that area,

19:54

and that's like our labs team. And then

19:56

there's like the second category, which

19:57

is that we actually do look at TAM.

19:59

Like, we're a business, we do look at

20:00

TAM, we do look at areas that we think,

20:02

uh, you know, there'd be reasonable

20:04

agentic like operations that would

20:06

happen.

20:07

In those areas,

20:09

uh, we do tend to have an orientation

20:10

towards things that are more token

20:12

heavy. And by token heavy or token

20:13

hungry, maybe is the way I would say

20:14

that, is like what we mean is like, you

20:16

know, you for spending once you spend a,

20:19

like call it like one turn, you look at

20:21

the end of that turn, and you say like,

20:23

"Am I done, or am I actually so glad

20:25

that I did that thing, I want to do more

20:26

of that thing?" We like industries where

20:28

it's like the answer to that question,

20:30

you say, "I want to do more of that

20:31

thing." So, coding is obviously the one

20:32

that we all know. And the great thing

20:33

about coding is that what it's actually

20:35

doing is that like once you finish a

20:37

turn, you look at that, and you're like,

20:39

"That was incredible. I'm like unlocked.

20:41

I'm going to do like more. I'm going to

20:42

build more, I can do more." And there's

20:44

other services where it's like actually

20:45

when you finish that turn, you completed

20:47

the job, and you just move on. You know

20:48

what I mean? Um, and so we tend to like

20:51

go into the ones that are a bit more

20:52

like there's this kind of like iterative

20:53

flow, you're going to build more,

20:55

generate more together.

20:56

Um, and then the last angle that we kind

20:58

of take a look at is just sort of like,

20:59

you know, there's going to be certain

21:01

business functions that were like they

21:03

are the buyer that we like to go to. We

21:04

want to help them optimize their

21:06

workflows, help them create better

21:07

products there. And I think we've been

21:09

pretty transparent with some of the

21:10

verticalization like we've done like

21:11

finance, we've done like legal.

21:13

Um and we've tried to kind of like

21:15

narrow on into specific areas where we

21:17

feel like by having the right context

21:19

and the right tools and putting it

21:20

together in a good form factor is

21:22

probably useful um for us to to be able

21:24

to do. And in each of those areas, we do

21:28

we're trying to do a bit of like showing

21:29

the art of the possible across all the

21:31

different ways that you would accomplish

21:33

those outcomes. And so for you know,

21:35

like finance for example is a good one.

21:38

Um you know, we you could be a company

21:41

that solves problems in finance and you

21:43

could build directly on the messages API

21:45

and you can just get some tokens and you

21:46

can build everything else on top. Or you

21:49

could be someone who builds on Claude

21:50

managed agents, you can get a lot more

21:52

out of the box. Or you could say, "I'm

21:54

going to build a plugin that or like a

21:56

connector, right?" That's going to sit

21:58

within one of our products and within

22:00

those form factors. And we did recently,

22:02

we launched like Claude for financial

22:03

services is like, "Okay, cool. We've got

22:06

packages of skills and things like this

22:08

that you could choose to use within our

22:10

product, within other people's

22:11

products." We even launched like

22:12

cookbooks on, "Here's how you would use

22:14

Claude managed agents to go and do these

22:16

things." And so I think for us, it's all

22:19

kind of an experimentation around like,

22:21

you know, we provide people all these

22:23

different pieces and see kind of where

22:25

they run with it. And then sometimes we

22:26

put together products that are just

22:28

packaging of all of these things like

22:30

Claude Tag, I think, is a really good

22:32

example. Like we had been seeing people

22:34

in the industry go into like Shopify did

22:37

this with River, um Square Block

22:39

recently did this with Builder Bot. Um

22:41

there's like a few of these examples

22:43

where people said, "I'm going to I'm

22:44

going to build like an agentic platform

22:46

internal to my company and I'm going to

22:48

try to give it all the right context and

22:49

I'm going to make it accessible from

22:51

Slack or from various other um you know,

22:54

platforms that you'd want it to be

22:55

accessible at." And I think Claude Tag

22:58

was very much a

22:59

packaging of all those same things that

23:01

anybody could choose to build something

23:03

similar, but this is how we're kind of

23:06

like, well, this is how we're doing it

23:07

internally, and if you would like to

23:09

just kind of plug in and go, here's what

23:11

that looks like.

23:11

>> you think people misunderstood about

23:13

Claude Tag? Cuz there's all this like

23:14

ruckus about, oh my god, it's just a

23:16

Slack bot. Like, tell tell us [laughter]

23:17

what what the magic of Tag is.

23:19

>> Yeah, I think it's a great question. Um

23:21

and I I do think it actually showcases a

23:22

little bit of where maybe the future

23:24

could be going.

23:25

Um yeah, I think like the I think if you

23:27

look at products in the past, people are

23:28

like, oh, you really attached to like

23:30

the form or the the UI almost, right?

23:32

Like, it looks like this, so it's like

23:34

super cool.

23:35

Um and I think when you look at like

23:36

Tag, uh it like, yeah, like the way you

23:39

interact with it is that you like

23:40

literally tag it in Slack. Uh and so

23:42

yeah, that is like the interface. But

23:44

that's not really the important part.

23:46

The important part um is all the kind of

23:49

like context engineering and like

23:52

architecture that we put underneath the

23:54

hood so that Tag just works. It really

23:57

should just like just feel like a

23:58

coworker. Like a co- you know, if you go

24:00

to a company and you onboard to the the

24:02

coworker comes into your channel and

24:03

then you can chat with it. It's

24:05

proactive, it's figured out like what's

24:07

like useful, you could and um it just

24:09

gets stuff like done for you. And so, if

24:11

you think about, you know, especially

24:13

like non-technical audiences, this is

24:15

like it's a huge unlock. You just you

24:17

literally create a channel and then you

24:18

@ Claude or sometimes you don't even @

24:19

Claude and you're like, "Hey, I want to

24:20

be able to do this and do that, and I

24:22

can't figure out this, and how do I

24:23

actually like submit an expense report

24:25

again?" And traditionally, do you think

24:27

about how to solve that workflow, you

24:28

are going all over the place, and you're

24:30

talking to your manager, and you're

24:31

talking to your spin buddy, and it's

24:32

like really really complicated.

24:34

And uh today, now you just like go talk

24:36

to Claude Tag, and we do a lot of the

24:38

hard work on doing the context

24:40

engineering, the proactivity, a lot of

24:42

the harness pieces. I think Andre

24:44

Kaparthy said it really well, he's like,

24:45

it's an org-level harness. There's a lot

24:46

of like complexity baked into that. Like

24:48

Kaylee mentioned, like you can use our

24:50

APIs to go and construct that. You can

24:51

do a lot of the experimentation

24:52

yourself, obviously, but this is like an

24:54

opinionated take from Anthropic on like

24:56

how you can have this really awesome

24:58

always on kind of agent for your entire

25:01

entire company. And the bit that's like

25:04

futuristic I guess is like a lot of that

25:07

complexity is actually like it's like an

25:09

iceberg. It's like all the stuff

25:11

underneath it. That's actually becoming

25:13

the harder and harder and like useful

25:14

part that we're trying to like push

25:16

through. And I think we'll see more and

25:18

more like that kind of like tip bit

25:20

that's like outside in the water. It's

25:22

just like the interface can actually

25:24

constantly swap. Like today right like

25:26

Slack is a place where a lot of people

25:28

collaborate. A lot of business

25:29

collaborate. But also a lot of people

25:30

collaborate in Teams. And some people

25:31

collaborate by a WhatsApp group or they

25:34

text each other or they may some people

25:36

still email each other. And like those

25:38

could be the form factors that actually

25:39

completely you can imagine agents just

25:41

going there and being. And they're

25:42

almost taking up the same form factors

25:44

as humans have taken up. It was almost

25:46

like a very almost like boring take but

25:48

it's actually like I feel like the most

25:50

like forward one because you want the

25:52

agent and you want AI to basically be

25:55

like another person and it's helping

25:57

you. But like it's like you know very

25:58

intelligent can figure out all the

26:00

context and you can always have it to be

26:02

a really helpful assistant.

26:03

>> Totally. You talked about context and

26:05

then harnesses quite a bit and so your

26:08

team is just you know has such an

26:09

opinion and point of view on like what

26:11

it takes to build an exceptional agent.

26:12

I imagine a lot of that comes down to

26:14

the context engineering and the

26:15

harnesses.

26:15

>> Totally.

26:16

>> Maybe like what best practices of or

26:18

advice would you would you share with

26:20

people about what you need to get right

26:21

on the harness and what you need to get

26:22

right on the context?

26:24

>> Yeah, I think so it's interesting

26:26

because we've kind of talked about you

26:28

know we launched Claude managed agents

26:30

as this like very generic but high

26:33

performing harness because we've done

26:35

all the nitty-gritty work that's

26:37

actually like really boring and not

26:39

super interesting around how do you deal

26:41

with prompt caching? How do you deal

26:42

with context management? You like clear

26:44

old stuff out of the window. Sometimes

26:46

you like call tools programmatically so

26:48

you don't pull everything into the

26:49

context window and you can keep it

26:50

clean. There's a lot of those sort of

26:52

details on the lower level harness

26:54

layer.

26:55

Um

26:56

and I think honestly like best practices

26:58

are just stuff like

27:00

prompt caching, do it. Save a lot of

27:02

money and and token costs. Um obviously

27:05

like try to keep your context window

27:06

clear and then putting those things

27:08

together in

27:10

uh a harness that will be performant is

27:12

is you know sometimes specific to the

27:15

task that you're trying to accomplish,

27:17

right? And then of course evals. Um I'm

27:20

surprised we got this far into this

27:22

thing before one of us said the word

27:23

evals. But like you need evals um to

27:25

make sure that what you're trying to

27:26

accomplish is performant. Um

27:29

but I think we're we're starting to go

27:31

and Angela mentioned this a little bit

27:32

earlier is more of a concept of

27:34

strategies or meta harnesses. Because I

27:37

do think that yes, you can again make

27:39

this lower level harness that's going to

27:41

be performant and maybe that's

27:42

interesting for you to do yourself or

27:43

maybe not and you offload it to us. But

27:46

this concept that

27:48

you can take any given token and spend

27:49

that token on just executing or you

27:52

could take that same token and choose to

27:54

actually reflect on your past agentic

27:57

sessions and write learnings to memory

27:59

so that the next agent does a good job.

28:01

Or you could take that token and advise

28:03

with a bigger model so that a smaller

28:05

model can execute and do a better job.

28:07

Um or you can say execute execute and

28:09

then like a grader comes in and is like,

28:11

"Did you do a good job? No, you didn't.

28:13

Try again." right? And so I think the

28:15

the like interesting innovation is going

28:18

to come more at that higher level on

28:20

like the meta level, right? And I think

28:23

optimizing within those strategies is

28:25

something that our team is really

28:26

excited about and we're starting to do a

28:28

lot of work there. Um and I think a lot

28:30

of other people are starting to feel

28:31

really excited about this concept of

28:33

strategies and like the jobs you give to

28:34

tokens.

28:36

Because again like yes, there's best

28:37

practices on stuff like your prompt

28:39

caching and exactly how you clear stuff

28:41

out of your context window and how you

28:42

write your evals and like a lot of

28:43

things like this. But I I don't know

28:46

that there's necessarily so much juice

28:47

to squeeze in a lot of cases out of that

28:50

layer as compared to a layer higher than

28:52

that.

28:53

>> Yeah, and one of the reasons for that I

28:54

think is it has to do with the

28:55

generations of the the models. I if you

28:57

look like 2 years ago, a lot of the

28:59

harness was like a scaffold to kind of

29:01

like

29:02

tell the model to go from point A to

29:03

point B. And you had to like you really

29:05

had to like build in a lot you had to

29:07

practically build one wall here and one

29:08

wall here so that like the thing would

29:09

go in a straight line. And now the

29:11

models are actually very very steerable.

29:14

Um and so a lot of that steering you can

29:16

just put it in the prompt, right? You're

29:17

like go do go from point A to point B.

29:20

And the model like will go from point A

29:22

to point B. So a lot of if you have

29:24

harnesses um that are like designed to

29:27

kind of do that kind of like steering,

29:29

you can delete that part. Like that part

29:31

we actually frequently encourage like

29:32

you can delete part of those harnesses.

29:33

I think various people have said things

29:34

along those lines. And that's I think

29:36

what people oftentimes mean when they're

29:38

like you know the model will kind of

29:39

consume some of the scaffolding. And

29:40

like in that sense like for sure if your

29:42

scaffolding is telling it to go in

29:43

direction, um that it can just

29:45

intelligently figure out. Like that I

29:47

think will increasingly continue to to

29:48

be so.

29:49

But as a result of of this, what the

29:51

harness needs to start doing is more

29:53

allow it to run longer. And so that's

29:55

where like that execution bit tends to

29:57

be I think like it sounds like a maybe

29:59

somewhat silly point, but I do think it

30:00

results in a lot of differences. Because

30:02

because it you can go in the direction

30:04

that you tell it to go, you obviously

30:05

don't want it to stop at B. You don't be

30:06

like okay now go from B to C and then go

30:08

to F and then go to Z and then come back

30:10

to me on A. You know, something funky

30:12

like that. In order to be able to do a

30:14

lot of those things, the kinds of

30:16

harnesses that you do are less the

30:17

steering harness and it's more like

30:19

these kind of strategy harnesses that

30:20

Caitlyn's mentioning, which allows you

30:22

to operate at a slightly higher level of

30:24

thinking, which matches I think a lot of

30:26

the intelligence gains that we're

30:27

starting to see with the model.

30:27

>> Do you think task-specific harnesses

30:30

make sense? Or a vertical-specific or

30:32

task-specific harnesses?

30:33

>> I think people have different opinions

30:34

on this. Like our opinion is yes. I

30:36

don't think there's like a general

30:37

harness. I think there are some

30:39

capabilities that are obviously very

30:40

general and they tend to like be very

30:43

useful. Like coding is a capability that

30:45

like is very useful cuz you use it

30:47

across so many things and software as

30:50

you know it's just like eating so much

30:51

of of what is capable. So our ability to

30:53

like write software is therefore useful.

30:55

I think when you think about like very

30:57

very specific types of domains that were

30:59

going to require like a couple of pieces

31:01

of the harness to be sort of like

31:03

customized. One of that I do think is

31:06

how you choose to kind of like handle

31:08

sort of like errors between when you do

31:11

something and you hand something off to

31:12

the model.

31:13

So in like domains where you require

31:15

like an extreme level of verification,

31:17

that logic of how you handle that like

31:19

it again I think it sounds small but

31:21

like I totally understand why some

31:22

people feel like they really want to own

31:24

the harness because tweaking that last

31:25

bit will give you a ton of juice. And

31:27

especially domains like like legal and

31:29

finance where there's a lot of

31:30

consequences to you know not getting it

31:33

perfectly correct, like it's really

31:34

going to matter and that's going to be

31:35

the difference between your product and

31:36

someone else's product being the thing

31:38

that the user ultimately uses. And then

31:41

there are other domains for which like I

31:43

would say it's not going to matter as

31:45

much because you're able to compress it

31:47

into like a general model capability. So

31:49

the tweaks that I guess like you know

31:51

where where we feel like the domain

31:52

specificity is really going to matter is

31:54

the specific like verification logic

31:56

between the model

31:58

and your execution and then

32:00

I think it's going to be about like some

32:02

of these kind of like higher order

32:04

strategies on how well you're able to

32:07

actually like allocate your token

32:08

budget. I think the context is actually

32:12

a little like over done. Like yes you're

32:13

going to like throw in context and like

32:15

that's but any harness can actually

32:17

handle a lot of context and so that's

32:19

just more like you have the data and if

32:20

you have the data then obviously you're

32:21

you're uniquely qualified to do

32:23

something useful.

32:24

>> Yeah and I think when people say

32:25

harnesses they often mean a lot of

32:27

different things and I think this is why

32:29

in part there's so many different

32:31

opinions on this. Like you can think of

32:32

a harness is literally just like a loop

32:35

Um, like okay cool like user model, user

32:38

model tool, you know, like that sort of

32:40

thing. Um, then you could think of the

32:42

harness as also all of the tools that

32:45

are packaged up with the harness, right?

32:47

And and there's just like a lot of

32:48

different definitions of these things,

32:49

and I think the stuff that can be pretty

32:52

generic and like less interesting to own

32:55

and and deal with what I was kind of

32:56

saying earlier is like getting your

32:58

prompt caching right, right? Like, maybe

33:00

that is not the world's most interesting

33:02

thing. Choosing to clear out old tool

33:05

calls from the context window and and

33:07

things like that, right? Are like maybe

33:08

a little bit less interesting.

33:10

And you like go a layer higher into some

33:12

of the stuff Angela was talking about,

33:13

and then you get into like, okay, yeah,

33:15

these are things that I might want to

33:16

own and control. And so, it's

33:18

interesting with cloud managed agents,

33:20

like, the thing that we built today, we

33:22

call it higher order, but it's not

33:24

really like that high order in the sense

33:26

that you can choose to define all of the

33:29

tools that you want to bring in as

33:31

custom tools with the harness, right?

33:32

And like, we give you a lot of knobs to

33:35

control. You can define skills, you can

33:36

do your system prompts, you can do a

33:37

whole bunch of different things, MCP

33:39

servers and things like this. And I

33:41

think where, you know, we want to get to

33:43

is a point where you can literally just

33:45

tell an agent, here's the outcome I

33:46

want, and here's the budget that I want

33:47

to spend, like, ready, set, go. And you

33:49

maybe like don't think about any of

33:51

those things underneath. And so, I think

33:53

there's just a few different layers of

33:55

this, right? That for certain things,

33:57

like, you might want to sit at a

33:58

different layer of what you actually go

34:00

and control, um, and you can probably

34:02

get better outcomes within some of those

34:04

layers by doing a little bit more

34:06

optimization work.

34:07

>> Very cool. One of the things I'm curious

34:09

about, and then what that I love about

34:11

infrastructure and platform teams, is

34:12

that you get to see what the most

34:13

advanced users in the world are using

34:15

and learn from them. I'm curious, what

34:17

are some things that you're seeing and

34:18

learning from from the people building

34:19

on your platform?

34:21

>> There's some people who've been doing

34:22

some really funky ways of like handling

34:25

context.

34:26

Um, we ourselves explored this a lot.

34:28

This is actually like one of the reasons

34:29

why tag is like uh such a great product

34:31

is like there's a lot of really awesome

34:33

like context kind of engineering that

34:34

that's happening.

34:36

Um we've seen some teams be really

34:38

clever about like how they do that and

34:40

they are able to kind of think through

34:42

like, "Okay, if I have all these

34:43

contacts in a bunch of different places,

34:44

how can I proactively go reach out to

34:46

them? How can I try to generate enough

34:48

like um permissions across each of them

34:51

so and then feed that all into like an

34:53

agent?" And it's interesting that like

34:55

um I guess like this is kind of the

34:56

level of innovation that like we're

34:58

actually like very excited by. It

34:59

doesn't express itself as like a

35:00

completely different form factor, um but

35:03

what it actually does express itself as

35:04

is like maximally useful to users and

35:07

we've been seeing this more and more

35:08

with like inter- actually like internal

35:09

use cases

35:10

instead of like external ones. So like

35:12

companies who are becoming more AI

35:14

native basically, they're the ones we're

35:15

seeing increasingly more and more

35:17

innovation out of and so you know, we've

35:19

had like customers try to do this for

35:21

their like they've built their own like

35:22

custom SDLC kind of setup in very very

35:25

innovative ways. We've had uh ones who

35:27

do that for like in their entire back

35:29

office. And just like the kind of

35:30

nuances of how they like stream in

35:32

context, I think it's been like actually

35:34

really interesting in terms of like how

35:35

they've been putting together the

35:36

pieces. So that's been like one category

35:38

that's been like really really like

35:39

fascinating. Uh another category that's

35:42

been like really interesting has

35:44

actually been with companies that are

35:45

dealing with like really old-school

35:47

software. And so there's a lot of like

35:48

healthcare companies um that we kind of

35:49

engage with and you know, like they're

35:52

like the the systems I'm working with,

35:53

they don't even have APIs. Like that's

35:55

that's a a dream. Um and so you know,

35:58

how can they use computer use uh and

36:01

things like this to be able to start to

36:03

kind of automate and create more

36:04

connectivity with our systems. Um and

36:07

that area of innovation I think has been

36:08

really exciting. It's been really

36:09

interesting to see people try all sorts

36:11

of crazy stuff from like taking a laptop

36:12

and trying to like run a bunch of things

36:13

on it to auto-generate a bunch of things

36:15

that then their agents can go and use.

36:17

Um and this has actually been probably

36:18

like an area of um I think a lot of

36:20

innovation coming from a lot of our

36:22

customers that we want to find ways to

36:23

like support better and see like, "Okay,

36:25

maybe there are like how can we make

36:27

this easier for you? How can we help you

36:29

with some standardization? How can we

36:31

get it so that, you know, like you can

36:32

just have a spec and then Claude can

36:34

then respect it. And so it's much easier

36:36

for you to organically connect a lot of

36:38

these things. But yeah, maybe the the

36:39

general theme I would just give you is

36:40

like interestingly, a lot of the

36:42

innovation that's most exciting out

36:44

there right now has been uh this kind of

36:46

like context and connectivity layer,

36:48

which has been really fascinating.

36:49

>> Yeah, like a good one in that um we were

36:52

working with a customer who they built

36:54

some agents on Claude managed agents.

36:56

They also have some agents that they

36:57

built on other models and other

36:59

platforms and they've kind of optimized

37:01

each of these agents to be good at the

37:03

things that they want. They want these

37:04

agents to all be able to work well

37:06

together. Um and they kind of were like

37:08

wow, galaxy brain, like what if I expose

37:10

an MCP server on top of this agent so

37:12

that I can then go and like have this

37:14

other agent call a tool on that agent,

37:17

right? And and have these things just be

37:18

more modular and be able to work

37:20

together. And we're like We sat down

37:22

with them and worked through it and and

37:24

it worked perfectly. And it was pretty

37:25

cool. And so we're seeing a lot of again

37:27

that connectivity layer that I think is

37:29

one of the cooler areas where people are

37:31

innovating, but

37:33

outside of that, one thing that has been

37:34

cool is just seeing the shift in I guess

37:37

like industry trends of where we're

37:39

seeing a lot of our usage come from.

37:40

Like talked a lot about coding, like

37:42

coding as a category, like of course

37:44

absolutely exploded. There's so much

37:45

going on there. And we're starting to

37:47

see some of these emerging trends like

37:48

more recently.

37:50

Um we're starting to see manufacturing

37:52

really pick up as just a category where

37:55

people are building with AI and like one

37:56

of our PMs like getting on a flight to

37:58

Detroit to go like figure out with these

38:01

customers like what they need and what's

38:03

going on. And so I think we're going to

38:04

start to see a lot more just kind of

38:07

like outside of the box of what people

38:09

think about today sort of use cases,

38:11

which we're really excited about.

38:13

>> It seems like there's no there's a we

38:15

went through a token maxing moments of

38:16

history and now there's like the token

38:19

rationalization moments of history.

38:21

[laughter]

38:21

What are your thoughts on that and like

38:23

what what should companies be doing and

38:24

then how it how does the platform team

38:26

think about enabling that?

38:28

>> Yeah, I mean it it makes sense.

38:30

It makes sense from the high you start

38:32

to rationalize. I really like that

38:34

framing. And I think there's like a

38:36

couple things that that are like top of

38:37

mind for us on this front. I think like

38:39

again it makes sense and as these models

38:40

get more and more capable you're going

38:41

to hit like levels of intelligence

38:43

maxing that are like there that then you

38:45

want to do the next kind of dimension

38:47

and the next dimension after

38:48

intelligence will either be cost or it

38:50

will be speed.

38:52

And you just kind of like you know go

38:53

through that across all possible task

38:54

complexities in the distribution. And as

38:57

we kind of see that like happen you know

38:59

something that's like really top of mind

39:00

for us that we kind of try to spend some

39:01

time with users on is like what you

39:03

don't want to do is like stop AI usage.

39:06

Right like that's kind of the wrong

39:07

move. And we do actually see some of our

39:09

our customers do that. So often times

39:11

the way that AI spend has erupted inside

39:13

their company has been through some kind

39:14

of like shadow IT. You know like their

39:16

employees just like want to use it they

39:18

find a way they end up procuring it

39:19

themselves and before you know it like

39:21

half your org has like found some way to

39:23

have installed cloud code. And in that

39:25

world it is kind of hard to to manage

39:26

because these things are again like

39:27

they're very token hungry ultimately.

39:29

And so what we try to kind of encourage

39:30

our customers is like okay you don't

39:31

want to like stop the innovation. Like

39:33

if you are getting returns on top of

39:34

this you are shipping faster than ever

39:36

before you can like run more

39:37

operationally like efficient then those

39:40

are gains. And so the area that we

39:42

actually try to encourage people is like

39:44

if there is a way for you to kind of

39:46

construct again like a strategy that

39:48

allows you to design an architecture

39:50

that says like given a task assesses

39:52

level of complexity. I mean I'm

39:53

effectively

39:54

describing a router but like there are

39:56

ways to do this that are like I think a

39:57

bit better now. And so like this task

39:59

comes in has a certain level of

40:00

complexity for that level of complexity

40:02

like you can define some rules but for

40:03

the most part right if it's like a hard

40:05

task you should probably route that to

40:06

like a big super smart model. And if

40:09

it's not a hard task you can route that

40:10

to like cheaper models.

40:11

Designing that I think has a little bit

40:13

of like there's a lot of technical

40:14

complexity in that but it's like very

40:16

very doable. And we actually like,

40:17

encourage people to try those kinds of

40:19

things. I think ultimately

40:20

>> offer a router?

40:21

>> I I think within the Claude space, it

40:24

will like make sense. It's actually one

40:25

of the strategies we imagine like

40:26

designing, cuz the way that we are kind

40:28

of thinking a lot of these things is

40:29

like, it almost feels like every month

40:31

there's a new era of something.

40:33

Um, and if we just take a step back and

40:34

go, "Okay, and it they seems to be like

40:36

really fast." And so, what are the

40:37

different ways that are re-composable so

40:39

we can redesign very quickly for any new

40:42

whatever the cool thing is that month

40:44

kind of like

40:45

Um, and so this is like in that category

40:47

of things where we feel like we can

40:48

actually just like re-compose a lot of

40:49

our primitives and then design it. I

40:51

think the bit that we do feel really

40:53

strongly about on the model routing

40:55

front is like, we are designing our

40:58

platform for Claude and we want to make

40:59

sure that Claude is great at like

41:01

solving all these things. So, we'll like

41:03

restrict to that space

41:05

um, rather than you know, I don't think

41:06

we're that interested in saying like,

41:08

"Okay, and then you know, you should

41:09

route to a different model or whatever."

41:10

>> Makes sense.

41:11

>> Yeah, and and well, some of that too is

41:13

just like, I think we have a strong

41:14

belief that harnesses and and just like

41:17

the agentic layer should be tuned to the

41:19

model family that you use it with. And

41:21

so, I think there was a period where

41:23

people were kind of like, "Yeah, cool. I

41:24

can like build a harness and build an

41:25

agent and then just like plug in a

41:27

different model underneath." And they

41:29

were excited about routers from that

41:31

perspective. And I think we've started

41:33

to see um, like Vercel just did this

41:35

with harness agent for example. Like,

41:37

some of these players in the space like

41:38

come up a layer of abstraction and say,

41:40

"Actually like, plug in the whole

41:42

harness and the whole agent that's tied

41:44

to a model family." Which makes a lot of

41:45

sense. And so, what we could provide is

41:47

a little bit better, smarter like, how

41:49

do you mix and match the right models

41:50

within the model family underneath that

41:52

thing if that makes sense. But, yeah, on

41:54

the general question of token maxing,

41:56

costs, and these sorts of things,

41:58

I think we're just kind of going through

42:00

what feels like a normal, natural cycle

42:02

for companies in figuring out how to

42:04

make the best use of this technology and

42:07

run their businesses really well and

42:09

really effectively. And um, it's

42:11

interesting like before working at

42:12

Anthropic, I was at Stripe and we were

42:14

kind of in the very reasonable era of

42:16

like we paid a lot of attention to our

42:18

AWS bill. And so, you know, if someone

42:21

were to have built some background job

42:23

and they like didn't quite configure it

42:25

correctly and this thing's like burning

42:27

through like CPU or whatever it is,

42:29

right? Like at any given moment and

42:32

causing, you know, big increase in span

42:34

that's not actually worth it, right?

42:35

Like we have put in place the guardrails

42:38

to find that and then go ask that

42:40

engineer very nicely to please turn off

42:41

their background job that's not like

42:43

within the bounds of of what they should

42:45

be spending for the thing they're trying

42:46

to accomplish. I think those are the

42:48

things with AI that people are going to

42:50

start to go and figure out and I think

42:52

to Angela's point, I think that gets

42:54

dangerous is when you're kind of just

42:56

like, here's a cap and you're stuck

42:58

within your cap like ready, set, go. But

43:01

I do think that encouraging innovation,

43:03

encouraging people to, you know, create

43:06

really excellent outcomes with this

43:07

stuff and then coming in from the side

43:09

and looking and saying like, okay, well,

43:11

there are few different ways we probably

43:13

could have accomplished that outcome,

43:15

right? And one is like you take Opus and

43:18

you run it all night and you do

43:19

something crazy and another is maybe to

43:21

get a little bit smarter with the

43:23

strategies that you put together in

43:24

order to create that same outcome within

43:26

a lower cost and I think that's the like

43:29

next layer of thinking that everyone's

43:30

going to start to do.

43:31

>> Very cool. Is there anything that you

43:33

guys are excited about building over the

43:34

next two months that you can share hints

43:35

at what might come next?

43:37

>> Uh, yeah, I mean, I know we said this

43:38

word like 20 million times, so I

43:40

apologize, but like we really are trying

43:41

to build ways for you to compose

43:43

strategies.

43:44

Um, and so, uh, that is an area that

43:47

that we're like trying to move into that

43:48

kind of like, yeah, uh, coordination

43:50

layer of the abstraction. Um, and we

43:52

want to start at at this front because

43:53

the types of problems that we see people

43:55

building, they are at a layer where it's

43:58

like in order to get the most return on

44:00

this, you have to be a little clever

44:02

about like what is the nature of the

44:03

problem that you're solving. So, to give

44:05

you something like concrete, like if

44:06

when you try to solve for like let's say

44:08

you want to build an agent that's like

44:10

trying to

44:11

do bug hunting. And you could just send

44:14

one off to go and do that. And it's

44:16

going to give you a certain type of

44:17

return, a level of return possibility.

44:20

And then people kind of get stuck at

44:21

that and they're like, "Okay, my next

44:22

options are I can like

44:24

make a bigger I can just like swap the

44:26

model for a different probably bigger

44:28

model

44:29

or I could like let it run like longer.

44:31

And that's what pretty much like the

44:32

only two like levers that you have to

44:34

like try to make this like bug hunting

44:35

agent.

44:36

From a lot of experimentation, when we

44:37

do these kinds of things it's like

44:39

actually the thing like those two those

44:41

two things are still true. But you

44:42

actually have a third lever and tends to

44:44

actually do a lot more than you think it

44:45

does. Which is that actually if you were

44:47

to like best of end the thing, it would

44:49

like give you a lot more returns. But

44:50

like just to be just saying those words

44:52

are fine. There's plenty of papers and

44:54

people have published it to actually

44:55

build that thing and put it into

44:56

production so you can actually test it

44:58

on users and see the results for

44:59

yourself. That's like really really

45:01

freaking hard. And you end up building

45:02

all these like custom harnesses and so

45:04

on and so forth. So like you know, all

45:05

that stuff.

45:06

But we're seeing like this is where the

45:07

alpha is and it's hard. And so like in

45:10

the same very simple philosophy that we

45:12

talked about at the beginning, like if

45:13

it's like gives you the return that you

45:14

want and it's hard, we're going to go

45:15

try to just make it easy for you. So

45:17

then you can use it to then run the

45:19

experiments that you actually need to

45:20

run. It reminds me of when people were

45:21

talking about agent swarms a year ago.

45:23

It's some version of that.

45:25

>> Yeah. Has it been a whole year?

45:26

>> Yeah.

45:27

I know. We're finally there.

45:28

>> Yes.

45:29

Yeah. No, I think that that's like

45:31

that's a type of strategy. Exactly. In

45:33

the same way that you have like you

45:34

know, one big one that separates a bunch

45:35

that's another type of strategy. I think

45:37

people have thought about this maybe the

45:38

in the way of like human organization. I

45:41

guess it could be similar. But if you

45:42

take it to the kind of its end state,

45:43

it's actually more just like the token

45:44

has a job. And I think it's this job

45:46

piece that we're we're really indexed on

45:48

and

45:50

we see a lot of returns to you. And

45:52

that's the thing that we want to spend

45:53

time with users and the rest of the

45:55

ecosystem on on like how can we just

45:56

make that easier for folks to then

45:57

experiment? Like we can give you like

45:59

five jobs off the top of our head and

46:01

we'll probably "That's what we have

46:02

internally.

46:04

Um and if we give this out to the rest

46:05

of the ecosystem, there's probably going

46:06

to be like 100,000, 200,000, who knows

46:08

what other combinations of people could

46:09

put together."

46:10

>> Yeah. We want to be able to keep doing

46:11

this hill climbing on like, how do you

46:13

get the most value, the most

46:14

intelligence per dollar, and just put

46:16

that power in people's hands. But around

46:19

the edges of that, we have these

46:21

personas that have kind of just like

46:23

things they have to work through in

46:24

order to be able to like really deploy

46:26

AI either within their companies or

46:29

within their products. And um that's

46:31

like the sort of enterprise-ready

46:34

security and compliance controls and

46:36

things like this. But really even just

46:38

like making the platform more modular in

46:41

the right ways, like being able to plug

46:42

in different pieces of the solutions

46:44

that we're building, like I want to use

46:45

memory for this thing over here, right?

46:47

Or whatever else it is, and having a

46:50

truly excellent developer experience

46:51

around that, because we spend a lot of

46:53

time with enterprises who are like,

46:55

"Okay, I have this like walled garden,

46:58

and I need to figure out exactly how I

47:00

can plug these in." And so, we're We've

47:02

got a part of our team that's innovating

47:04

on things like strategies and jobs and

47:05

trying to help you maximize

47:06

intelligence. And they're like, "That's

47:07

really cool, but I can't actually use

47:09

any of that for XYZ reasons." So, I

47:11

think solving those problems is really,

47:13

really important to us. But then the

47:14

other persona is, you know, the like

47:16

weekend developer who's like, "I want to

47:18

go and build something useful for

47:20

myself, right?" And they're often doing

47:22

that on top of our platform and on top

47:24

of many other just pieces of developer

47:26

platforms in the community. And I think

47:28

for some of those folks, there's more

47:30

that we can do to be provide solutions

47:32

that are maybe more open or more

47:33

hackable or whatever it may be for those

47:35

folks to kind of just like go wild with

47:38

what we can offer them and have this

47:40

really excellent developer experience.

47:41

And so, I think there's a lot of stuff

47:43

that maybe I would put in the the

47:45

category of table stakes that I'm really

47:47

excited about, because I think those are

47:49

the things that then unlock getting

47:51

people to say, "Okay, yes, this thing

47:53

works for me, and now I can plug in on

47:55

some of the stuff that you guys are

47:56

doing that's really innovative and help

47:58

line me to get more intelligence and

47:59

save costs and things like that.

48:01

>> Wonderful.

48:02

>> Caitlin and Angela, I feel I mean you

48:04

were building one of the most important

48:06

developer platforms in the world and

48:08

talking to the two of you over time I

48:09

just feel really optimistic that that

48:12

platform is in very thoughtful

48:15

hands that that care about the

48:16

ecosystem. So thank you for taking the

48:18

time today to share what you're up to

48:19

and we look forward to what's ahead.

48:21

>> Thanks for having us.

48:22

>> Thank you guys.

48:33

>> [music]

48:49

[music]

Interactive Summary

In this discussion, Caitlin and Angela from Anthropic explore the evolution of their developer platform, emphasizing a transition from basic model access to a more robust, multi-layered architecture involving knowledge, execution, and coordination. They discuss their philosophy of dogfooding tools internally while ensuring developers have access to powerful, modular primitives. The conversation highlights their focus on "agentic" work, the introduction of higher-order abstractions like managed agents, and the emerging concept of "strategy harnesses" that allow for more complex, orchestrated AI behavior. They also address the shift from "token maxing" to "token rationalization" and how their platform aims to support building efficient, cost-effective, and highly capable agentic systems.

Suggested questions

4 ready-made prompts