Thariq (Claude Code) @ Anthropic
798 segments
I I think it's just [music] more like
the highle idea of like you probably
have a lot of unknown unknowns, right?
And you're probably not being ambitious
enough. [music]
I think we sometimes oversimplify
agentic engineering where we're like
it's just [music] prompts, it's just
loops or whatever. And I think it's more
like no, we need to know what we're
doing. We need to like learn more and
then like be better at prompting [music]
and and make sure we're creating like
valuable work.
>> [music]
>> All right. Okay. [music] Cool. Hey guys,
uh my name is Stark. Is is the mic
working fine? Yeah, it's good. Uh cool.
Okay. Yeah. Um [applause] thank you.
Thank you. Yeah.
Um yeah, I uh work on the cloud code
team and uh yeah, I um I I've been
working on it for about a year now.
Before this, I was a a YC founder uh and
ran that company for about five years.
Um decided to get into AI. Actually,
Eric is here from Goodfire. Eric was
like my first sort of like AI gig, you
know. So, like I did some
interpretability work with with with
Goodfire. Uh eventually ended up on the
cloud code team. Um and uh yeah, so I
Okay, I I thought there were like going
to be seven people at this event. And so
I I was sort of like, you know, you have
to sort of like shape your content on on
on the audience. I I'm super excited to
talk about this. I I have a bunch of
ideas I want to talk about. Um and I
have a Yeah, I've had giving an AI
engineer keynote next week, but this is
sort of like mid formation of these
ideas. So there are like a few different
like things I'll I'll jump between. So
like please like, you know, bear with me
a little bit. I I think there's like a
good thread of thought, but you know,
like maybe everything won't be like the
decks won't be color aligned or
something, you know. So, um and I'm not
sure if this is the title I'll go with.
I'm I'm not sure, but okay, the the
primary thing I want to talk about is um
like we need to be better, you know? I
mean, like I I think like you know,
people tag me on Twitter sometimes,
they're like, "What are you spending all
the tokens on?" You know, like what what
what's happening? Where's the like
where's the growth? And I'm like, you're
right. You know what I mean? Like like
we like I think the goal for AI is to
like really meaningfully improve, you
know, like how uh like humanity, you
know, like like our GDP, right? And I
think that we haven't shown this yet. Uh
and so how why right I I think the thing
that we all obviously see is like
software is becoming super cheap, right?
Knowledge work is becoming cheap, right?
Um and this is like you know like
obvious to us but I think maybe less
obvious is that like generating value is
still really hard you know and like uh
building startups as I'm sure many of
you are doing here is extremely hard and
bu software and knowledge work are parts
of this right um but it's like not all
of it right and we haven't honestly
earned like our valuations we haven't
earned the like you know the amount of
money we're putting into AI yet I think
you know um because there's still more
value to create, right? And so, uh, why,
you know, and I I think like the the
thing I come back to that I've really
appreciated at Anthropic is like we have
this value that's like we don't
negotiate against ourselves, you know,
and so I remember being a CEO of a
company and being like, okay, what are
our priorities? Let's write down the
priorities and then let's figure out how
they trade off against each other. Um,
and it was very reasonable, but like,
uh, what if you were less reasonable,
you know? I mean, like what if you ask
reality to kind of show you what the
trade-offs are? And I feel like what I
really appreciate at Anthropic is we're
like, let's just do the thing, you know,
let's just do all of it. Force us to be
shown what the the trade-offs are,
right? And um I think with AI and Claude
more than more than anything, I think
it's like, you know, what are the
trade-offs really? It's harder to tell,
right? Like is good, fast, and cheap
still a trade-off? uh maybe not right
and so I think that this like the qu
like you know like how do we free
ourselves from the mental models of like
you know what our trade-offs are what
our ambitions are things like that um so
that we can ultimately deliver the
promise right of of AI so um this is
something that I've been thinking a lot
about and and you know I'm kind of like
okay why why you know why is this so
hard I think one of the reasons is that
like LMS are just really weird and
they're like a new thing and we need to
like figure this out, right? So, okay,
whatever. Um, yeah, I think it's like uh
I think this is like roughly the art of
like human agent interaction, you know,
and like it's a very I I think a lot of
the, you know, tech companies in like
the 2010s were built off human computer
interaction, right? Like UIs, like UX,
like incredible like patterns and like
now we have to have develop this like
new like technique, right? Human agent
interaction. Um, and I think like
talking to LLMs is like is a new skill,
right? It's like prompting or public,
sorry, it's like public speaking or
writing or any of these things. Uh, if
you're good at it, you can drive like a
ton more value than, you know, someone
who's not good at it, right? And and so
it's something that I think will
continue to be like high leverage for a
very long time. Um, and I think it's
important for us to I I think
acknowledge that, right? Like I don't
think the end goal is like uh you know,
you just put a sentence into cloud and
it like does the thing for you. I think
there's a lot of like depth here. Um and
uh it's because the models are like
grown not designed right so um you we
like are cultivating the data the RL
environments the like you know etc etc
like all of the things that like go into
pre-post mid training um but we don't
know what the end outcome will be until
we get the model right like it's like
you don't just decide like hey this
model is going to be you know 95% on
bench you grow it right and uh we you
grow with a lot of care. Uh but they're
organic things and like we don't exactly
know what they will be good at until
like we really try them, right? Uh we
have some guesses. Uh but these things
emerge like in spiky ways. And so I I
say that like models don't get smarter
in a straight line, right? They get
smarter in unexpected ways. Um I think
one of the like best examples is that
like uh you know if you were to go back
to like maybe Sonic 3.5 and you look at
cursor and you're like okay how does the
model like solve coding you're like okay
obviously the context window get really
really large like a 100 million token
context window and then like we just fit
the entire codebase in and just solves
it right like that's how you think the
model would get faster but but it didn't
happen like that right it like got
better by doing tool calling and by bash
and GP and and like how do you figure
that I don't know. You just have to you
have to do it, right? So, um I think
they're smarter than they think and
we're hobbling clawed is what I think a
lot about is that like part of the
reason that you know we haven't achieved
as much growth as we could is that we
like you know wow we are hobbling
clawed. Um so I'll give a case study
here uh that that I like. Uh Pokemon
ending in awe. There was this tweet
about like, you know, this font may be
too small, but basically like there was
like an AI hate tweet going on where I
was like, I can't believe chat GPD can't
answer Pokemon ending with AW. You know,
Cloud could answer it, but like we let's
not get into that. Um, [laughter]
but but well, it's uh why, right? So, uh
there are over a thousand Pokemon. Um,
exactly two of them end in awe. It's
Crocana and Dreadnaugh. I didn't know
that. I'm a big Pokemon fan, but like
you know most Pokemon people would not
know that a priority, right? Um and so
chatt found one. Um and so like how
would you solve this? Like like let's
say we obviously think the models are
smart enough. Why why couldn't it? And
how would it solve it? One idea is like
it's all in the weights. You just
remember the weights like you know like
obviously it knows every Pokemon. It
should just think fast enough or like it
should just remember or it should like
think you know and just think out all
the words. um or it searches the web,
right? And like all of these things like
it it will not give you the answer,
right? This is kind of empirical like is
there a version of like the architecture
that does just remember uh maybe you
know but like empirically this does not
happen but the models are obviously
smart enough. So what works is like you
ask cloud code you ask any like coding
any agent with a code exec tool right um
and cloud code will like get the list of
all Pokemon and GP for once that ended
aw right and so this is like one line
it'll just do it um and like if you were
you know like if if you're an average
user you just don't you're like this you
ask you know a model why what the
Pokemon ending in aw you are and it
doesn't and like you you just stop
there, right? Like you just don't have
the ability or you don't have the skill
of understanding claude well enough to
be like, oh, like how do I prompt it
better? You know, how do I like give it
the tools needed? Right? Um unhobling
it. So, but you can go further. You can
generate a web app that like you know uh
searches any like combination of reg x
for Pokemon, right? And like like
there's so much more abundance than than
you than you think or like than the
problem implies.
Um so yeah we call this capability
overhang right like we call like the
idea that like what the models can do
and like what we like you know are
utilizing them for are mismatched right
and I think definitely like you know
open 4.8 data has like incredible
capability overhang to me like fable
like you know just I feel bad about it
right so um I I think like uh yeah the
but but why is it hard I I think like um
one of the things to do is like let's
try and put ourselves in Claude's shoes
right and so the usual framing is like
it's just the model in the harness like
you just you figure out like you know
what the tools are and you put it in um
but like let's say you're claude like
it's more of a complicated relationship
between like the model, the harness, the
world and and you, right? So, um yeah,
it's like putting ourselves in its
shoes. So, let's say that you wanted to
like explain how the odds module works,
right? This is like a very this is maybe
even a good prompt that someone might
put in, right? Uh but in order to do
this, like it needs to know who you are,
right? Like do you know anything about
the codebase? Are you technical or not?
You know, like how technical are you? Uh
what level of depth do you want? Uh the
world, right? like how big is this
codebase? This actually has a big
implication, right? Like if you're doing
this in like a large legacy codebase,
you need to spend a lot of compute. If
you're doing it in a small codebase,
maybe you have no sub agents, you just
you just do it, right? And so this is
something that cloud needs to figure out
too.
Um and then like what other context can
it get, right? So can it look in git or
slack or like you know like is there
like more nuance around around the O
module? Why is the user asking me this?
Right? And like if like your boss came
up to you to like, hey, explain me what
the O module is. You might be like,
well, what for what, you know, I mean
like how how do I get more context so I
can answer this question in the way you
want, right? Uh but claude doesn't
really like get the for what, you know?
Um and so yeah, and of course like the
harness can help with some of this,
right? Like it can do memory, it can
like learn a little bit more about you.
Um, but there's still like a lot here
that like Claude, you know, like we have
some work to do. Uh, and and so like,
you know, one example of making this
prompt more explicit is like, uh, hey,
like I'm an experienced TypeScript
engineer. I have zero familiarity with
this module, but like it's kind of a
high compute problem, so use sub aents,
right? So, uh, this is one way that you
might get like a little bit more
precise. Um,
now, okay, let's see. I think like I
kind of want to switch tracks a little
bit here right now and I want to talk
about
um sorry this is one of those things
where I haven't figured exactly okay
cool yeah um yeah so another thing like
unhobbling claude putting thing putting
ourselves in claude shoes uh what are
like you know how do we discover what
claude can do right I think one of my
like uh I I think there's a lot of room
for surprise here. And so one of the
things I've been recently well surprised
about is video editing, right? So um
Claude like just edits videos that we
now, you know, like I shoot them with
the like a video agency and then uh it
end to end I don't use a video editor. I
just use cloud code and it generates
videos that you know look like uh look
like this, right? So it's um you know
showing me like it's generating the UI
as well. It's generated this across many
different cuts and so it's decided like
you know which are the best cuts of the
data. Um it's like cut out ums and
things like that, right? Um and it's,
you know, done some pretty like
impressive like UI, you know, like it's
uh done these overlays and things like
that. And uh yeah, I'd say like if you
were ask people, you know, hey, can
Claude edit videos like you know, they
they they wouldn't think it could,
right? Um and like how do you even like
like
what does the process or if you were to
ask Claude to edit a video, it probably
wouldn't be able to answer you in this
way, right? because it it just doesn't
you know know how to unhobble itself. So
um yeah how does that look like right
and how do you get to this? Um
I will so uh the high level is that it's
code generation right and so uh this is
like a representation of the folder I
got here and uh uh some of the previews.
Yeah. I mean like the the idea I want to
leave you with is like there's a bunch
of clips here. Oh, maybe I have it here.
Okay. Yeah. Uh there's a bunch of clips
that I'm given. Um, and this is
essentially the raw material of the like
video, you know, and then what I ask
cloud to do is I I give it a transcript
as well. This is the transcript I was
working from. Uh, so it has an idea of
like, you know, what I'm trying to say.
Um, and I ask it to transcribe it,
right? And so it does, uh, let's see, it
does a bunch of transcriptions.
Um
there's so much stuff here,
but this is like essentially what
knowledge work is increasingly becoming.
It's like just a folder of code and
scripts and data, you know, that is like
uh like controlled by an agent, right?
And so, um, yeah, roughly what it does
is, and the funny thing is I I asked it
to make this,
uh, artifact as well to show you guys to
sort of simplify the the process. Um,
but yeah, it makes a bunch of, uh,
transcriptions of the uh of the of each
video. It decides then which clips to
do, which map up to my transcript the
best, and then starts editing them,
clipping them together, and then making
UI to go along with it as well. So, it's
generated a bunch of UI here in using
React that will then compile compile
together into like the end video, right?
And uh it's also deciding which UI
elements to make given the like Figma
and you know, design system from our
team. Um,
more than that. So, I did all of this
the first pass. This was the first video
I did. Um, and this was the second. And
I think there's like a big difference
here is like the color. So, like one of
the things I had no idea about is color
grading, right? Color grading is I still
I know more about it now. Um, but like
you know when you get a video like turns
out they give you in in this like raw
format which is like sort of overexposed
and and there's actually a lot of art in
like oh how do you bring out the color
of a video right? Um and this was
something where uh this was a place
where I realized I had a lot of unknown
unknowns and uh okay let me come back
now I'm like sort of deciding how how I
want to like so yeah what does it mean
to be good at agentic engineering right
like I I think like how do you like sort
of uh how do you solve these problems
like how do you like uh yeah what's the
skill like and I think it comes down to
a lot of like unknown known, right? And
like with like there are known known
like what do you want? There are known
unknowns, what you haven't figured out
yet. Unknown unknowns, things like that
are obvious that you don't know yet. And
sorry, unknown knowns that are things
that are obvious but you haven't uh you
only recognize it when you see it. And
then unknown unknowns. And so in this
case, this was me being like, wait, I
have so many unknowns about color
grading, right? This is one of those
things where um in order for me to be an
better agentic engineer, I know a lot
about how ffmpeg works, how video
transcription works, how reotion works.
I know all these things and that's what
let me get far enough to edit this
video. Um but I didn't know enough about
unknown unknowns, right? Um or about
color grading in particular was like a
big unknown unknown for me. And so the
question is like how do you fix that,
right? And I think this is like a big
problem overall. Like we if our goal is
to ship better things faster to like,
you know, drive GDP, to make better
products, we have to get really good at
grounding down our unknown unknowns. Um,
and so what I did was I ended up asking
Claude to tell me about color grading.
And there is a version of this where
like sometimes you just ask Claude and
then you like sort of glaze over the
like the report, you know? This happens
a lot. I I think it's kind of like
education porn sort of. You're like,
"Oh, like yeah, this is, you know, it
looks cool." But I I I it generated this
and I honestly had no idea really what
color grading was from this. I couldn't
prompt it better. I couldn't like figure
out better. But I really tried to stick
through it. So I asked it to like sort
of show me create more visualizations. I
kept asking questions like, "Oh, like
hey, why you know like like what do
these different things do? Like what is
a vector scope?" like um uh yeah ask
giving me like some sort of like it
pulled literature elsewhere from you
know uh from color grading and put it in
here and then finally it made showed me
a bunch of examples of like okay this is
you know your old version this is a new
version uh ultimately color grading is
kind of like a shader it's like you know
you take in a pixel you output a
different pixel color um and this is
like a good mental model for me to build
I knew what a shader was Um, and the
thing that really got me here was like I
was able to build a visualization where
I could go over every pixel and see how
the value would change over time, you
know? And so like I think through this
like visualization, I was able to figure
out, okay, what does color grading do?
And then I could ultimately tell Claude
like, hey, I want something like this.
And like the insight I got here is that
you want to grade your skin differently
than uh the background because humans
have a larger lower dynamic range for a
skin color. Like it'll look weird if
you're like kind of purple, but the
background can be kind of purple. You
know what I mean? [laughter]
Yeah. Yeah. Yeah. And and in this
particular case, um my skin color and
the background were very close together.
And so there was like some, you know,
more interesting work for Claude to do,
but like I was able to like express this
problem and sort of figure out, okay,
what what is wrong with it, you know,
like what could be better? Um, and then
like fix it, right, through this like
process of like grounding down my
unknown unknowns. And uh, yeah, I think
that like there is so much work to do in
that case, right? And I think like some
of what we, you know, I think we
sometimes oversimplify agentic
engineering where we're like it's just
prompts, it's just loops or whatever.
And I think it's more like no, we need
to know what we're doing. We need to
like learn more and then like be better
at prompting and and make sure we're
creating like you know valuable work,
right? And in this case it this video
and would not have happened if I hadn't
been able to um you know to to edit it
myself. it would have like we actually
couldn't pay it someone to edit it fast
enough like we just couldn't. Um so yeah
there are some techniques on like how to
stay in the loop. Um
these are like you know I I'm writing
more on this like I think exploring with
Claude and brainstorming asking them to
interview like um we talked about in the
last talk building technical plans
explanation implementation notes
explainers um I'm not sure if I don't
think I want to go one by one over these
things. I think it's just more like the
highle idea of like you probably have a
lot of unknown unknowns, right? And
you're probably not being ambitious
enough uh kind of as a result almost and
like how do you sort of uh free
ourselves of this so that we can now
like unhobble Claude to do more more
useful work. Um so yeah that's uh that's
that's my talk. Yeah. [applause]
>> A question here. Yeah.
>> Claude tag.
>> Yes.
>> How much of your workflow has shifted
over to cloud tag?
>> Um quite a lot. Uh but I think like I
still do a lot of like exploration with
cloud tag for example. Um I like sort of
do a lot of thinking through it. So I
might ask it to generate like an
artifact or something and view it on my
phone. Um but most of the like in the
loop coding still happens with cloud
code. Um and uh it's more like you know
cloud tag is great at getting work
started. It's great at managing jobs and
things like that. Great at being
multiplayer. Yeah.
>> Is cloud tag still single player mode
within the organization or did it is it
multiplayer now?
>> Oh no, it's multiplayer by default for
everyone
>> by default but has it been adopted as
multiplayer like organizationally?
>> Um yeah I think so. Yeah. Yeah. We have
it in feedback channels and things like
that. people do together. Yeah.
>> Cool.
>> Yeah.
>> Uh could I get to know a little bit more
about your process for like empathizing
with the model and then designing the
environment around it because it's not
like pure in a human sense like
>> if you're leveraging pure reasoning then
you want like more composability or like
>> you know some level of common
denominator across every tool rather
than like adding on more and more.
>> So is there any other like intuitions
that you kind of learn through like
>> iterating with different tools? Yeah, I
I'm trying to figure out uh yeah, I I
think there's like I I overall think
this is an art kind of right now. Like I
I think that um one of the things uh
Okay, I do need to Oh, actually actually
no, I I skipped this part of the talk.
Maybe I can actually go back to this. Um
yeah, some some examples of how Claude
gets better over time. Um I do need to
take credit for this actually. Uh I did
the interview grill me thing like um I
you know I think Matt later made it into
a skill but I was the first one to like
sort of identify that. Um but ask user
question is a tool that I built in cloud
code. It helps cloud code ask you
questions and th this how you I've used
it has changed a lot over time where
like you know initially the model could
just call it once and then I tried being
like oh can you chain it together? Can
you interview me? um and then you know
would like put like 30 or 40 questions
together. Now I ask you to build HTML
reports you know and and sort of like uh
get much more in depth and then uh
select the answers from the HTML report.
Um and so the how you get through this
progression like Opus 4 could barely
call the ask user question tool. It was
like actually kind of hard. Um and now
it's like there's so much capability
overhang on top of it. Um, similarly
with like markdown and HTML like this is
something that you know I've talked
about a bunch before where you know
originally markdown uh was a way that
like maybe you know the model kept
itself on track and now then it became a
way of communicating to you and now like
it can create you know really rich HTML
reports and this is like the like the
model is getting smarter but in these
spiky ways right you need to switch
formats from markdown to HTML um in
order to unlock its capabilities. How do
you figure that out? I I don't have like
a science for you. You know what I mean?
Like I think that's like I think the
maybe trillion dollar question.
>> So just to clarify, there's like two
different ways. There's one versus
environment where like ideally it's just
like bash or super lenient where it's
like you have all these emerging
properties that come out of just like
more reasoning. Yeah.
>> And then the other is like it could only
work with the information it has. So
that's like more signal and like having
these or necessity these tools in the
first place.
>> Yeah. Like the context and things like
that. Yeah. I I mean um I I that seems
right but I need to think on it more
kind of if that's the right uh paradigm.
Yeah.
>> Yeah.
>> Um question. So to me a lot of agentic
work seems a little bit like the old
school unsexy waterfall like software
development right where you have
specifications then you have
implementation QA all of that right?
>> Yeah. So how would you say
one could go beyond that specifically
for agents because right now it's like
interview me that's basically
requirements and specifications right
like iterate that's QA
>> sure
>> so how do you think one can make it
better for agents
>> um what's the end goal you think like is
there like a problem you're seeing
>> build software right like um like for
example like I've been using it a lot
for software building
>> so a lot of times I have an idea of what
I wanted to build and I iterate, you
know, via different ways.
>> Sure.
>> So, but I still have somewhat of a
product in mind.
>> Yeah.
>> That I want to build. So, I kind of
follow this process, but I'm trying to
figure out how to do it in a better way.
So, that maybe because like to your
point, there's a capability overhead.
Like, what am I missing? What am I not
thinking? How should I think about
interacting with agents in a more
sophisticated way?
>> Yeah. So the the problem is that the
agent is not building exactly what you
want and like there's some mismatch
between you and the agent. Is that
right?
>> It's building what I want. It's just it
takes longer than I would want.
>> Um
yeah, I mean uh I think there is it kind
of depends on your particular flow. Like
I I think like one of the things I do a
lot is I build a lot of prototypes
first. And so these prototypes can be
really cheap. like you know you can sort
of like build a mockup in HTML so the
agent doesn't need to fully implement it
and so then you can sort of get a sense
of what you want. I often find that I
don't know what I want when I'm building
something and there's like a iterative
process of finding out what I want that
can happen like quite deep in the
implementation process and that's
usually what slows me down is like I'm
like 80% of the way there I'm like ah no
that's that's wrong you know so like how
do you like get that earlier and earlier
um and like what I like to do is I I
have this like iterative like spec
interview prototype uh see if I like the
prototype maybe even like make a
prototype PR are kind of all in along
the way. Um then I might reset and take
those learnings and and start again, you
know. Um because uh yeah, I just like I
want to build the most valuable thing I
can, you know. Yeah.
>> Yeah. Daisy.
>> Yes.
>> Do you still use plan mode?
>> No. Yeah. We need to do something about
this. Yeah. Yeah. Yeah. Daisy also works
on cloud code, so um Yeah. Yeah. Yeah.
Uh [laughter]
>> yes.
>> All right. Last question, Robbie.
>> Uh, how do you decide what to like have
Claude spend like take time to learn
from Claude? I feel like my primary form
of brain rot these days is like Claude
explaining stuff to me and like an hour
later I'm like, "Oh, I need remedial
matrix algebra to understand this." Like
just go home. I don't know.
>> Yeah. Yeah, I know what you mean. Like,
let's see. This is a good question.
Honestly, like I don't think there's
like an easy answer. I I think that like
um
you know, like Yeah. Like ultimately
it's like what is the goal that you're
doing? I I do think one thing about
agentic engineering is that it's so fun.
Like sometimes you can just be like just
do this and it's just fun and you're not
actually driving an outcome. I I'm not
saying that's bad but I'm saying like
you know like sometimes you do want to
be like okay what what is wrong about my
output? How do I get better? And and
probably the answer is that there are
unknown unknowns you have, you know,
like there's something that you're like
not able to express well enough, you
know, um or maybe you don't know enough
about the user or or whatever, right?
And like how do you like answer that?
And hopefully Claude can help you
answer. Um but yeah, you know, it's uh
yeah, you know, Claude, you can learn
about everything. Not you don't need to
learn about everything, but there are
particular things that you need to.
Yeah.
>> Cool. Thank you.
>> Thanks. [applause]
>> [music]
Ask follow-up questions or revisit key timestamps.
The speaker, a software engineer at Anthropic, discusses the evolving field of agentic engineering and human-agent interaction. They emphasize the importance of moving beyond simple prompting, encouraging users to overcome 'unknown unknowns' and better 'unhobble' AI models like Claude to maximize their potential. By integrating tools for code execution, iterative prototyping, and deep learning, engineers can achieve significantly higher productivity and generate more valuable results.
Videos recently processed by our community