HomeVideos

Thariq (Claude Code) @ Anthropic

Now Playing

Thariq (Claude Code) @ Anthropic

Transcript

798 segments

0:00

I I think it's just [music] more like

0:01

the highle idea of like you probably

0:03

have a lot of unknown unknowns, right?

0:04

And you're probably not being ambitious

0:06

enough. [music]

0:08

I think we sometimes oversimplify

0:10

agentic engineering where we're like

0:11

it's just [music] prompts, it's just

0:12

loops or whatever. And I think it's more

0:14

like no, we need to know what we're

0:15

doing. We need to like learn more and

0:17

then like be better at prompting [music]

0:19

and and make sure we're creating like

0:20

valuable work.

0:26

>> [music]

0:31

>> All right. Okay. [music] Cool. Hey guys,

0:32

uh my name is Stark. Is is the mic

0:35

working fine? Yeah, it's good. Uh cool.

0:38

Okay. Yeah. Um [applause] thank you.

0:41

Thank you. Yeah.

0:43

Um yeah, I uh work on the cloud code

0:45

team and uh yeah, I um I I've been

0:49

working on it for about a year now.

0:50

Before this, I was a a YC founder uh and

0:54

ran that company for about five years.

0:56

Um decided to get into AI. Actually,

0:57

Eric is here from Goodfire. Eric was

1:00

like my first sort of like AI gig, you

1:02

know. So, like I did some

1:03

interpretability work with with with

1:05

Goodfire. Uh eventually ended up on the

1:07

cloud code team. Um and uh yeah, so I

1:11

Okay, I I thought there were like going

1:12

to be seven people at this event. And so

1:14

I I was sort of like, you know, you have

1:17

to sort of like shape your content on on

1:19

on the audience. I I'm super excited to

1:21

talk about this. I I have a bunch of

1:23

ideas I want to talk about. Um and I

1:25

have a Yeah, I've had giving an AI

1:27

engineer keynote next week, but this is

1:30

sort of like mid formation of these

1:32

ideas. So there are like a few different

1:34

like things I'll I'll jump between. So

1:36

like please like, you know, bear with me

1:38

a little bit. I I think there's like a

1:40

good thread of thought, but you know,

1:42

like maybe everything won't be like the

1:44

decks won't be color aligned or

1:45

something, you know. So, um and I'm not

1:48

sure if this is the title I'll go with.

1:49

I'm I'm not sure, but okay, the the

1:53

primary thing I want to talk about is um

1:56

like we need to be better, you know? I

1:59

mean, like I I think like you know,

2:00

people tag me on Twitter sometimes,

2:02

they're like, "What are you spending all

2:03

the tokens on?" You know, like what what

2:04

what's happening? Where's the like

2:06

where's the growth? And I'm like, you're

2:07

right. You know what I mean? Like like

2:09

we like I think the goal for AI is to

2:12

like really meaningfully improve, you

2:15

know, like how uh like humanity, you

2:18

know, like like our GDP, right? And I

2:20

think that we haven't shown this yet. Uh

2:22

and so how why right I I think the thing

2:25

that we all obviously see is like

2:28

software is becoming super cheap, right?

2:30

Knowledge work is becoming cheap, right?

2:33

Um and this is like you know like

2:36

obvious to us but I think maybe less

2:38

obvious is that like generating value is

2:40

still really hard you know and like uh

2:42

building startups as I'm sure many of

2:44

you are doing here is extremely hard and

2:47

bu software and knowledge work are parts

2:50

of this right um but it's like not all

2:53

of it right and we haven't honestly

2:55

earned like our valuations we haven't

2:58

earned the like you know the amount of

2:59

money we're putting into AI yet I think

3:01

you know um because there's still more

3:03

value to create, right? And so, uh, why,

3:07

you know, and I I think like the the

3:09

thing I come back to that I've really

3:10

appreciated at Anthropic is like we have

3:12

this value that's like we don't

3:14

negotiate against ourselves, you know,

3:16

and so I remember being a CEO of a

3:19

company and being like, okay, what are

3:21

our priorities? Let's write down the

3:22

priorities and then let's figure out how

3:24

they trade off against each other. Um,

3:26

and it was very reasonable, but like,

3:28

uh, what if you were less reasonable,

3:30

you know? I mean, like what if you ask

3:31

reality to kind of show you what the

3:33

trade-offs are? And I feel like what I

3:35

really appreciate at Anthropic is we're

3:36

like, let's just do the thing, you know,

3:38

let's just do all of it. Force us to be

3:40

shown what the the trade-offs are,

3:42

right? And um I think with AI and Claude

3:46

more than more than anything, I think

3:47

it's like, you know, what are the

3:49

trade-offs really? It's harder to tell,

3:51

right? Like is good, fast, and cheap

3:53

still a trade-off? uh maybe not right

3:56

and so I think that this like the qu

3:58

like you know like how do we free

4:00

ourselves from the mental models of like

4:03

you know what our trade-offs are what

4:05

our ambitions are things like that um so

4:07

that we can ultimately deliver the

4:09

promise right of of AI so um this is

4:12

something that I've been thinking a lot

4:14

about and and you know I'm kind of like

4:17

okay why why you know why is this so

4:19

hard I think one of the reasons is that

4:21

like LMS are just really weird and

4:23

they're like a new thing and we need to

4:25

like figure this out, right? So, okay,

4:27

whatever. Um, yeah, I think it's like uh

4:30

I think this is like roughly the art of

4:32

like human agent interaction, you know,

4:34

and like it's a very I I think a lot of

4:38

the, you know, tech companies in like

4:40

the 2010s were built off human computer

4:42

interaction, right? Like UIs, like UX,

4:45

like incredible like patterns and like

4:47

now we have to have develop this like

4:48

new like technique, right? Human agent

4:51

interaction. Um, and I think like

4:53

talking to LLMs is like is a new skill,

4:55

right? It's like prompting or public,

4:57

sorry, it's like public speaking or

4:58

writing or any of these things. Uh, if

5:00

you're good at it, you can drive like a

5:02

ton more value than, you know, someone

5:05

who's not good at it, right? And and so

5:06

it's something that I think will

5:08

continue to be like high leverage for a

5:10

very long time. Um, and I think it's

5:13

important for us to I I think

5:14

acknowledge that, right? Like I don't

5:15

think the end goal is like uh you know,

5:17

you just put a sentence into cloud and

5:20

it like does the thing for you. I think

5:21

there's a lot of like depth here. Um and

5:24

uh it's because the models are like

5:27

grown not designed right so um you we

5:30

like are cultivating the data the RL

5:33

environments the like you know etc etc

5:36

like all of the things that like go into

5:38

pre-post mid training um but we don't

5:42

know what the end outcome will be until

5:43

we get the model right like it's like

5:45

you don't just decide like hey this

5:47

model is going to be you know 95% on

5:49

bench you grow it right and uh we you

5:53

grow with a lot of care. Uh but they're

5:55

organic things and like we don't exactly

5:58

know what they will be good at until

6:00

like we really try them, right? Uh we

6:02

have some guesses. Uh but these things

6:05

emerge like in spiky ways. And so I I

6:08

say that like models don't get smarter

6:10

in a straight line, right? They get

6:12

smarter in unexpected ways. Um I think

6:15

one of the like best examples is that

6:17

like uh you know if you were to go back

6:20

to like maybe Sonic 3.5 and you look at

6:22

cursor and you're like okay how does the

6:23

model like solve coding you're like okay

6:25

obviously the context window get really

6:27

really large like a 100 million token

6:30

context window and then like we just fit

6:32

the entire codebase in and just solves

6:33

it right like that's how you think the

6:35

model would get faster but but it didn't

6:37

happen like that right it like got

6:38

better by doing tool calling and by bash

6:41

and GP and and like how do you figure

6:43

that I don't know. You just have to you

6:44

have to do it, right? So, um I think

6:47

they're smarter than they think and

6:48

we're hobbling clawed is what I think a

6:50

lot about is that like part of the

6:52

reason that you know we haven't achieved

6:53

as much growth as we could is that we

6:55

like you know wow we are hobbling

6:58

clawed. Um so I'll give a case study

7:01

here uh that that I like. Uh Pokemon

7:04

ending in awe. There was this tweet

7:06

about like, you know, this font may be

7:09

too small, but basically like there was

7:11

like an AI hate tweet going on where I

7:14

was like, I can't believe chat GPD can't

7:16

answer Pokemon ending with AW. You know,

7:18

Cloud could answer it, but like we let's

7:20

not get into that. Um, [laughter]

7:22

but but well, it's uh why, right? So, uh

7:27

there are over a thousand Pokemon. Um,

7:30

exactly two of them end in awe. It's

7:32

Crocana and Dreadnaugh. I didn't know

7:34

that. I'm a big Pokemon fan, but like

7:36

you know most Pokemon people would not

7:39

know that a priority, right? Um and so

7:42

chatt found one. Um and so like how

7:45

would you solve this? Like like let's

7:47

say we obviously think the models are

7:48

smart enough. Why why couldn't it? And

7:50

how would it solve it? One idea is like

7:52

it's all in the weights. You just

7:53

remember the weights like you know like

7:55

obviously it knows every Pokemon. It

7:56

should just think fast enough or like it

7:59

should just remember or it should like

8:00

think you know and just think out all

8:02

the words. um or it searches the web,

8:05

right? And like all of these things like

8:07

it it will not give you the answer,

8:10

right? This is kind of empirical like is

8:12

there a version of like the architecture

8:14

that does just remember uh maybe you

8:16

know but like empirically this does not

8:18

happen but the models are obviously

8:20

smart enough. So what works is like you

8:23

ask cloud code you ask any like coding

8:26

any agent with a code exec tool right um

8:30

and cloud code will like get the list of

8:33

all Pokemon and GP for once that ended

8:36

aw right and so this is like one line

8:39

it'll just do it um and like if you were

8:42

you know like if if you're an average

8:44

user you just don't you're like this you

8:46

ask you know a model why what the

8:49

Pokemon ending in aw you are and it

8:51

doesn't and like you you just stop

8:53

there, right? Like you just don't have

8:55

the ability or you don't have the skill

8:57

of understanding claude well enough to

8:59

be like, oh, like how do I prompt it

9:00

better? You know, how do I like give it

9:02

the tools needed? Right? Um unhobling

9:04

it. So, but you can go further. You can

9:06

generate a web app that like you know uh

9:10

searches any like combination of reg x

9:13

for Pokemon, right? And like like

9:14

there's so much more abundance than than

9:17

you than you think or like than the

9:19

problem implies.

9:21

Um so yeah we call this capability

9:23

overhang right like we call like the

9:25

idea that like what the models can do

9:28

and like what we like you know are

9:30

utilizing them for are mismatched right

9:32

and I think definitely like you know

9:34

open 4.8 data has like incredible

9:36

capability overhang to me like fable

9:38

like you know just I feel bad about it

9:41

right so um I I think like uh yeah the

9:46

but but why is it hard I I think like um

9:49

one of the things to do is like let's

9:51

try and put ourselves in Claude's shoes

9:53

right and so the usual framing is like

9:55

it's just the model in the harness like

9:56

you just you figure out like you know

9:58

what the tools are and you put it in um

10:01

but like let's say you're claude like

10:03

it's more of a complicated relationship

10:04

between like the model, the harness, the

10:07

world and and you, right? So, um yeah,

10:11

it's like putting ourselves in its

10:13

shoes. So, let's say that you wanted to

10:16

like explain how the odds module works,

10:18

right? This is like a very this is maybe

10:20

even a good prompt that someone might

10:22

put in, right? Uh but in order to do

10:24

this, like it needs to know who you are,

10:26

right? Like do you know anything about

10:27

the codebase? Are you technical or not?

10:29

You know, like how technical are you? Uh

10:31

what level of depth do you want? Uh the

10:34

world, right? like how big is this

10:35

codebase? This actually has a big

10:37

implication, right? Like if you're doing

10:39

this in like a large legacy codebase,

10:42

you need to spend a lot of compute. If

10:43

you're doing it in a small codebase,

10:45

maybe you have no sub agents, you just

10:46

you just do it, right? And so this is

10:48

something that cloud needs to figure out

10:50

too.

10:52

Um and then like what other context can

10:55

it get, right? So can it look in git or

10:58

slack or like you know like is there

11:01

like more nuance around around the O

11:03

module? Why is the user asking me this?

11:06

Right? And like if like your boss came

11:08

up to you to like, hey, explain me what

11:10

the O module is. You might be like,

11:11

well, what for what, you know, I mean

11:13

like how how do I get more context so I

11:15

can answer this question in the way you

11:17

want, right? Uh but claude doesn't

11:19

really like get the for what, you know?

11:20

Um and so yeah, and of course like the

11:23

harness can help with some of this,

11:24

right? Like it can do memory, it can

11:26

like learn a little bit more about you.

11:28

Um, but there's still like a lot here

11:30

that like Claude, you know, like we have

11:33

some work to do. Uh, and and so like,

11:35

you know, one example of making this

11:37

prompt more explicit is like, uh, hey,

11:39

like I'm an experienced TypeScript

11:41

engineer. I have zero familiarity with

11:43

this module, but like it's kind of a

11:45

high compute problem, so use sub aents,

11:47

right? So, uh, this is one way that you

11:48

might get like a little bit more

11:50

precise. Um,

11:52

now, okay, let's see. I think like I

11:55

kind of want to switch tracks a little

11:59

bit here right now and I want to talk

12:00

about

12:03

um sorry this is one of those things

12:05

where I haven't figured exactly okay

12:07

cool yeah um yeah so another thing like

12:11

unhobbling claude putting thing putting

12:14

ourselves in claude shoes uh what are

12:17

like you know how do we discover what

12:18

claude can do right I think one of my

12:21

like uh I I think there's a lot of room

12:24

for surprise here. And so one of the

12:26

things I've been recently well surprised

12:28

about is video editing, right? So um

12:31

Claude like just edits videos that we

12:34

now, you know, like I shoot them with

12:36

the like a video agency and then uh it

12:39

end to end I don't use a video editor. I

12:41

just use cloud code and it generates

12:42

videos that you know look like uh look

12:45

like this, right? So it's um you know

12:49

showing me like it's generating the UI

12:52

as well. It's generated this across many

12:54

different cuts and so it's decided like

12:57

you know which are the best cuts of the

12:58

data. Um it's like cut out ums and

13:01

things like that, right? Um and it's,

13:03

you know, done some pretty like

13:04

impressive like UI, you know, like it's

13:06

uh done these overlays and things like

13:08

that. And uh yeah, I'd say like if you

13:10

were ask people, you know, hey, can

13:12

Claude edit videos like you know, they

13:16

they they wouldn't think it could,

13:18

right? Um and like how do you even like

13:20

like

13:22

what does the process or if you were to

13:23

ask Claude to edit a video, it probably

13:24

wouldn't be able to answer you in this

13:26

way, right? because it it just doesn't

13:28

you know know how to unhobble itself. So

13:31

um yeah how does that look like right

13:33

and how do you get to this? Um

13:38

I will so uh the high level is that it's

13:40

code generation right and so uh this is

13:44

like a representation of the folder I

13:47

got here and uh uh some of the previews.

13:52

Yeah. I mean like the the idea I want to

13:55

leave you with is like there's a bunch

13:56

of clips here. Oh, maybe I have it here.

13:59

Okay. Yeah. Uh there's a bunch of clips

14:01

that I'm given. Um, and this is

14:04

essentially the raw material of the like

14:07

video, you know, and then what I ask

14:09

cloud to do is I I give it a transcript

14:11

as well. This is the transcript I was

14:13

working from. Uh, so it has an idea of

14:15

like, you know, what I'm trying to say.

14:17

Um, and I ask it to transcribe it,

14:20

right? And so it does, uh, let's see, it

14:23

does a bunch of transcriptions.

14:26

Um

14:28

there's so much stuff here,

14:30

but this is like essentially what

14:32

knowledge work is increasingly becoming.

14:34

It's like just a folder of code and

14:36

scripts and data, you know, that is like

14:39

uh like controlled by an agent, right?

14:42

And so, um, yeah, roughly what it does

14:46

is, and the funny thing is I I asked it

14:49

to make this,

14:51

uh, artifact as well to show you guys to

14:54

sort of simplify the the process. Um,

14:57

but yeah, it makes a bunch of, uh,

15:00

transcriptions of the uh of the of each

15:04

video. It decides then which clips to

15:06

do, which map up to my transcript the

15:09

best, and then starts editing them,

15:12

clipping them together, and then making

15:14

UI to go along with it as well. So, it's

15:16

generated a bunch of UI here in using

15:19

React that will then compile compile

15:21

together into like the end video, right?

15:24

And uh it's also deciding which UI

15:26

elements to make given the like Figma

15:30

and you know, design system from our

15:32

team. Um,

15:35

more than that. So, I did all of this

15:36

the first pass. This was the first video

15:38

I did. Um, and this was the second. And

15:41

I think there's like a big difference

15:42

here is like the color. So, like one of

15:44

the things I had no idea about is color

15:46

grading, right? Color grading is I still

15:49

I know more about it now. Um, but like

15:52

you know when you get a video like turns

15:54

out they give you in in this like raw

15:55

format which is like sort of overexposed

15:58

and and there's actually a lot of art in

16:00

like oh how do you bring out the color

16:02

of a video right? Um and this was

16:05

something where uh this was a place

16:08

where I realized I had a lot of unknown

16:10

unknowns and uh okay let me come back

16:14

now I'm like sort of deciding how how I

16:17

want to like so yeah what does it mean

16:20

to be good at agentic engineering right

16:22

like I I think like how do you like sort

16:24

of uh how do you solve these problems

16:27

like how do you like uh yeah what's the

16:30

skill like and I think it comes down to

16:32

a lot of like unknown known, right? And

16:35

like with like there are known known

16:38

like what do you want? There are known

16:39

unknowns, what you haven't figured out

16:40

yet. Unknown unknowns, things like that

16:43

are obvious that you don't know yet. And

16:45

sorry, unknown knowns that are things

16:47

that are obvious but you haven't uh you

16:49

only recognize it when you see it. And

16:51

then unknown unknowns. And so in this

16:53

case, this was me being like, wait, I

16:55

have so many unknowns about color

16:57

grading, right? This is one of those

16:59

things where um in order for me to be an

17:01

better agentic engineer, I know a lot

17:03

about how ffmpeg works, how video

17:05

transcription works, how reotion works.

17:07

I know all these things and that's what

17:08

let me get far enough to edit this

17:11

video. Um but I didn't know enough about

17:13

unknown unknowns, right? Um or about

17:16

color grading in particular was like a

17:18

big unknown unknown for me. And so the

17:20

question is like how do you fix that,

17:22

right? And I think this is like a big

17:23

problem overall. Like we if our goal is

17:27

to ship better things faster to like,

17:29

you know, drive GDP, to make better

17:32

products, we have to get really good at

17:35

grounding down our unknown unknowns. Um,

17:37

and so what I did was I ended up asking

17:39

Claude to tell me about color grading.

17:42

And there is a version of this where

17:44

like sometimes you just ask Claude and

17:46

then you like sort of glaze over the

17:48

like the report, you know? This happens

17:50

a lot. I I think it's kind of like

17:52

education porn sort of. You're like,

17:53

"Oh, like yeah, this is, you know, it

17:55

looks cool." But I I I it generated this

17:57

and I honestly had no idea really what

17:59

color grading was from this. I couldn't

18:01

prompt it better. I couldn't like figure

18:03

out better. But I really tried to stick

18:05

through it. So I asked it to like sort

18:07

of show me create more visualizations. I

18:10

kept asking questions like, "Oh, like

18:12

hey, why you know like like what do

18:14

these different things do? Like what is

18:15

a vector scope?" like um uh yeah ask

18:20

giving me like some sort of like it

18:22

pulled literature elsewhere from you

18:25

know uh from color grading and put it in

18:27

here and then finally it made showed me

18:30

a bunch of examples of like okay this is

18:33

you know your old version this is a new

18:35

version uh ultimately color grading is

18:38

kind of like a shader it's like you know

18:39

you take in a pixel you output a

18:40

different pixel color um and this is

18:42

like a good mental model for me to build

18:44

I knew what a shader was Um, and the

18:47

thing that really got me here was like I

18:49

was able to build a visualization where

18:50

I could go over every pixel and see how

18:53

the value would change over time, you

18:55

know? And so like I think through this

18:57

like visualization, I was able to figure

18:59

out, okay, what does color grading do?

19:01

And then I could ultimately tell Claude

19:03

like, hey, I want something like this.

19:05

And like the insight I got here is that

19:08

you want to grade your skin differently

19:10

than uh the background because humans

19:13

have a larger lower dynamic range for a

19:16

skin color. Like it'll look weird if

19:19

you're like kind of purple, but the

19:20

background can be kind of purple. You

19:22

know what I mean? [laughter]

19:23

Yeah. Yeah. Yeah. And and in this

19:25

particular case, um my skin color and

19:28

the background were very close together.

19:30

And so there was like some, you know,

19:31

more interesting work for Claude to do,

19:33

but like I was able to like express this

19:35

problem and sort of figure out, okay,

19:37

what what is wrong with it, you know,

19:39

like what could be better? Um, and then

19:42

like fix it, right, through this like

19:44

process of like grounding down my

19:45

unknown unknowns. And uh, yeah, I think

19:48

that like there is so much work to do in

19:51

that case, right? And I think like some

19:52

of what we, you know, I think we

19:56

sometimes oversimplify agentic

19:57

engineering where we're like it's just

19:58

prompts, it's just loops or whatever.

20:00

And I think it's more like no, we need

20:02

to know what we're doing. We need to

20:03

like learn more and then like be better

20:05

at prompting and and make sure we're

20:06

creating like you know valuable work,

20:09

right? And in this case it this video

20:11

and would not have happened if I hadn't

20:14

been able to um you know to to edit it

20:17

myself. it would have like we actually

20:19

couldn't pay it someone to edit it fast

20:20

enough like we just couldn't. Um so yeah

20:25

there are some techniques on like how to

20:27

stay in the loop. Um

20:29

these are like you know I I'm writing

20:31

more on this like I think exploring with

20:33

Claude and brainstorming asking them to

20:36

interview like um we talked about in the

20:38

last talk building technical plans

20:40

explanation implementation notes

20:42

explainers um I'm not sure if I don't

20:45

think I want to go one by one over these

20:47

things. I think it's just more like the

20:48

highle idea of like you probably have a

20:51

lot of unknown unknowns, right? And

20:52

you're probably not being ambitious

20:53

enough uh kind of as a result almost and

20:56

like how do you sort of uh free

20:59

ourselves of this so that we can now

21:01

like unhobble Claude to do more more

21:03

useful work. Um so yeah that's uh that's

21:08

that's my talk. Yeah. [applause]

21:14

>> A question here. Yeah.

21:15

>> Claude tag.

21:17

>> Yes.

21:17

>> How much of your workflow has shifted

21:19

over to cloud tag?

21:20

>> Um quite a lot. Uh but I think like I

21:24

still do a lot of like exploration with

21:27

cloud tag for example. Um I like sort of

21:30

do a lot of thinking through it. So I

21:31

might ask it to generate like an

21:33

artifact or something and view it on my

21:35

phone. Um but most of the like in the

21:38

loop coding still happens with cloud

21:40

code. Um and uh it's more like you know

21:44

cloud tag is great at getting work

21:45

started. It's great at managing jobs and

21:47

things like that. Great at being

21:48

multiplayer. Yeah.

21:50

>> Is cloud tag still single player mode

21:52

within the organization or did it is it

21:54

multiplayer now?

21:55

>> Oh no, it's multiplayer by default for

21:56

everyone

21:57

>> by default but has it been adopted as

21:59

multiplayer like organizationally?

22:00

>> Um yeah I think so. Yeah. Yeah. We have

22:03

it in feedback channels and things like

22:05

that. people do together. Yeah.

22:07

>> Cool.

22:07

>> Yeah.

22:08

>> Uh could I get to know a little bit more

22:10

about your process for like empathizing

22:12

with the model and then designing the

22:13

environment around it because it's not

22:14

like pure in a human sense like

22:16

>> if you're leveraging pure reasoning then

22:17

you want like more composability or like

22:19

>> you know some level of common

22:20

denominator across every tool rather

22:22

than like adding on more and more.

22:24

>> So is there any other like intuitions

22:26

that you kind of learn through like

22:27

>> iterating with different tools? Yeah, I

22:30

I'm trying to figure out uh yeah, I I

22:33

think there's like I I overall think

22:35

this is an art kind of right now. Like I

22:37

I think that um one of the things uh

22:41

Okay, I do need to Oh, actually actually

22:43

no, I I skipped this part of the talk.

22:45

Maybe I can actually go back to this. Um

22:47

yeah, some some examples of how Claude

22:49

gets better over time. Um I do need to

22:52

take credit for this actually. Uh I did

22:54

the interview grill me thing like um I

22:57

you know I think Matt later made it into

22:59

a skill but I was the first one to like

23:01

sort of identify that. Um but ask user

23:03

question is a tool that I built in cloud

23:05

code. It helps cloud code ask you

23:07

questions and th this how you I've used

23:10

it has changed a lot over time where

23:12

like you know initially the model could

23:14

just call it once and then I tried being

23:16

like oh can you chain it together? Can

23:18

you interview me? um and then you know

23:20

would like put like 30 or 40 questions

23:22

together. Now I ask you to build HTML

23:25

reports you know and and sort of like uh

23:28

get much more in depth and then uh

23:29

select the answers from the HTML report.

23:32

Um and so the how you get through this

23:35

progression like Opus 4 could barely

23:37

call the ask user question tool. It was

23:38

like actually kind of hard. Um and now

23:40

it's like there's so much capability

23:42

overhang on top of it. Um, similarly

23:45

with like markdown and HTML like this is

23:48

something that you know I've talked

23:49

about a bunch before where you know

23:51

originally markdown uh was a way that

23:53

like maybe you know the model kept

23:55

itself on track and now then it became a

23:58

way of communicating to you and now like

24:00

it can create you know really rich HTML

24:02

reports and this is like the like the

24:04

model is getting smarter but in these

24:07

spiky ways right you need to switch

24:09

formats from markdown to HTML um in

24:12

order to unlock its capabilities. How do

24:14

you figure that out? I I don't have like

24:15

a science for you. You know what I mean?

24:17

Like I think that's like I think the

24:18

maybe trillion dollar question.

24:20

>> So just to clarify, there's like two

24:21

different ways. There's one versus

24:22

environment where like ideally it's just

24:24

like bash or super lenient where it's

24:26

like you have all these emerging

24:27

properties that come out of just like

24:28

more reasoning. Yeah.

24:29

>> And then the other is like it could only

24:31

work with the information it has. So

24:32

that's like more signal and like having

24:34

these or necessity these tools in the

24:36

first place.

24:37

>> Yeah. Like the context and things like

24:39

that. Yeah. I I mean um I I that seems

24:42

right but I need to think on it more

24:44

kind of if that's the right uh paradigm.

24:46

Yeah.

24:47

>> Yeah.

24:48

>> Um question. So to me a lot of agentic

24:51

work seems a little bit like the old

24:54

school unsexy waterfall like software

24:57

development right where you have

24:59

specifications then you have

25:00

implementation QA all of that right?

25:02

>> Yeah. So how would you say

25:07

one could go beyond that specifically

25:09

for agents because right now it's like

25:11

interview me that's basically

25:13

requirements and specifications right

25:15

like iterate that's QA

25:17

>> sure

25:17

>> so how do you think one can make it

25:19

better for agents

25:21

>> um what's the end goal you think like is

25:24

there like a problem you're seeing

25:25

>> build software right like um like for

25:27

example like I've been using it a lot

25:28

for software building

25:30

>> so a lot of times I have an idea of what

25:32

I wanted to build and I iterate, you

25:35

know, via different ways.

25:36

>> Sure.

25:37

>> So, but I still have somewhat of a

25:39

product in mind.

25:40

>> Yeah.

25:41

>> That I want to build. So, I kind of

25:42

follow this process, but I'm trying to

25:44

figure out how to do it in a better way.

25:46

So, that maybe because like to your

25:48

point, there's a capability overhead.

25:49

Like, what am I missing? What am I not

25:51

thinking? How should I think about

25:53

interacting with agents in a more

25:55

sophisticated way?

25:56

>> Yeah. So the the problem is that the

25:58

agent is not building exactly what you

26:00

want and like there's some mismatch

26:02

between you and the agent. Is that

26:02

right?

26:03

>> It's building what I want. It's just it

26:04

takes longer than I would want.

26:08

>> Um

26:10

yeah, I mean uh I think there is it kind

26:13

of depends on your particular flow. Like

26:15

I I think like one of the things I do a

26:16

lot is I build a lot of prototypes

26:18

first. And so these prototypes can be

26:20

really cheap. like you know you can sort

26:22

of like build a mockup in HTML so the

26:25

agent doesn't need to fully implement it

26:27

and so then you can sort of get a sense

26:29

of what you want. I often find that I

26:31

don't know what I want when I'm building

26:32

something and there's like a iterative

26:34

process of finding out what I want that

26:36

can happen like quite deep in the

26:38

implementation process and that's

26:40

usually what slows me down is like I'm

26:42

like 80% of the way there I'm like ah no

26:44

that's that's wrong you know so like how

26:46

do you like get that earlier and earlier

26:48

um and like what I like to do is I I

26:51

have this like iterative like spec

26:54

interview prototype uh see if I like the

26:57

prototype maybe even like make a

26:59

prototype PR are kind of all in along

27:00

the way. Um then I might reset and take

27:03

those learnings and and start again, you

27:05

know. Um because uh yeah, I just like I

27:09

want to build the most valuable thing I

27:10

can, you know. Yeah.

27:13

>> Yeah. Daisy.

27:14

>> Yes.

27:14

>> Do you still use plan mode?

27:16

>> No. Yeah. We need to do something about

27:19

this. Yeah. Yeah. Yeah. Daisy also works

27:21

on cloud code, so um Yeah. Yeah. Yeah.

27:24

Uh [laughter]

27:28

>> yes.

27:28

>> All right. Last question, Robbie.

27:30

>> Uh, how do you decide what to like have

27:32

Claude spend like take time to learn

27:34

from Claude? I feel like my primary form

27:35

of brain rot these days is like Claude

27:37

explaining stuff to me and like an hour

27:39

later I'm like, "Oh, I need remedial

27:41

matrix algebra to understand this." Like

27:43

just go home. I don't know.

27:44

>> Yeah. Yeah, I know what you mean. Like,

27:47

let's see. This is a good question.

27:50

Honestly, like I don't think there's

27:51

like an easy answer. I I think that like

27:53

um

27:55

you know, like Yeah. Like ultimately

27:57

it's like what is the goal that you're

27:58

doing? I I do think one thing about

28:00

agentic engineering is that it's so fun.

28:02

Like sometimes you can just be like just

28:03

do this and it's just fun and you're not

28:05

actually driving an outcome. I I'm not

28:07

saying that's bad but I'm saying like

28:08

you know like sometimes you do want to

28:10

be like okay what what is wrong about my

28:13

output? How do I get better? And and

28:15

probably the answer is that there are

28:17

unknown unknowns you have, you know,

28:18

like there's something that you're like

28:19

not able to express well enough, you

28:22

know, um or maybe you don't know enough

28:24

about the user or or whatever, right?

28:26

And like how do you like answer that?

28:27

And hopefully Claude can help you

28:29

answer. Um but yeah, you know, it's uh

28:34

yeah, you know, Claude, you can learn

28:36

about everything. Not you don't need to

28:38

learn about everything, but there are

28:39

particular things that you need to.

28:40

Yeah.

28:42

>> Cool. Thank you.

28:43

>> Thanks. [applause]

28:47

>> [music]

Interactive Summary

The speaker, a software engineer at Anthropic, discusses the evolving field of agentic engineering and human-agent interaction. They emphasize the importance of moving beyond simple prompting, encouraging users to overcome 'unknown unknowns' and better 'unhobble' AI models like Claude to maximize their potential. By integrating tools for code execution, iterative prototyping, and deep learning, engineers can achieve significantly higher productivity and generate more valuable results.

Suggested questions

3 ready-made prompts