How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel
1503 segments
You want to externalize your taste and
your judgment before you start writing
evals.
>> No matter how many stuff I add and how
many evals I add, it cannot get it
perfect in one shot.
>> Do you have the bottom up evals? I think
this is where my question mark comes up
for me. [music] Bottom up is like data
driven. Top down is not about the data.
Claude is very very bad at coming up
with bottom up evals. That's all you.
The interface that it comes up with is
going to be so much better than me
looking at my data in say Google
spreadsheets [music] or something. We
joke that writing is the final boss for
LLMs. To like have it right like you in
a way that you are satisfied with is
extremely difficult.
Hey everyone, my guests today are Hambo
and Shrea, the instructors behind the
most popular AI eval course online. I
think they've taught over 4,500 students
now. So, they're here to share a really
practical example about how to use AI to
automate evals the right way without
doing a whole bunch of slop. And uh
yeah, really excited to have you both.
>> Really excited. Likewise. [laughter]
>> Great. Okay. So, so Haml, last time you
came on the podcast, I feel like I
learned a lot from you. I think some of
the lessons were manly review customer
traces and conversations to figure out
where things are going wrong. Use like
yes, no, past fail evals instead of
making up random scores. But maybe
before we get to the demo, can you guys
talk about at a high level about how
eval like last year?
>> So, the fundamentals still apply. You
still want to start everything by
looking at data and you want to do error
analysis and you want to have a
structured approach to looking at data
so that you can externalize your taste
and your judgment before you start
writing eval. The main thing that has
changed is getting agents to help you
look at the data in a very thoughtful
way. So agents have become a lot more
powerful over time and we have figured
out a workflow that allows you to kind
of have agents running in the background
while you look at data supporting you so
that you can have a lot more leverage in
that first stage of looking at data. And
that's what Treya is going to show today
is an example of that.
>> The other thing I'll add to that is that
um we we teach LLM judges a lot in our
course. LLM judges are a way to have an
LLM look at a trace and make a judgment
about a very specific failure mode. So
for example, uh is is this output too
long or is this output too short? Does
it follow a particular structure? Um LLM
judges have gotten very good at
evaluating these very well- definfined
criteria more accurately. So we are much
bigger fans I would say of LLM judges.
>> Got it. Since I have two ex experts
here, I'd love for you guys to take a
look at my pretty pretty dumb evals for
a skill that I built and maybe give me
some feedback. Basically, I built a
skill for podcast post-production and
part of the skill is it writes a bunch
of it if tried to find the top 10
takeaways from the podcast and then it
writes an initial draft as a newsletter
post and I built a bunch of evals for
for this or I as in claude a bunch of
evals for this. So uh so for example if
you go down here there's some evals
around like are the takeaways the right
length between 240 and 330 characters
and then there's also some like judgment
things is each takeaway understandable
on its own without watch the episode
that does it give plain useful advice so
so I guess there's kind of two types of
take two types of eos right one is just
like checking if the character lens is
correct and the other one is more like
judgment based but yeah do you guys have
any thoughts about this or like how how
can I make this better? Yeah, this is
awesome. I love that you have a real
demo for us to take a look at. There's
some things that I want to back up a
little bit on, which is what's in the
skill, right? If we're going to start
doing evals, it's very hard to do evalu.
And uh what was your I love that you
tried to make evals. So then kind of
what was your thought process on making
the evals? Was it like I'm just going to
ask Claude to make some emails for me or
I don't know you saw some outputs or you
you saw some things that you noticed
were like off. Um and then you wanted to
create emails around that or like tell
us a little bit more about that.
>> That's a good question. Yeah. So um let
me pull up my little blog post. Um so
basically I do these interviews and uh
there's a video and then the people who
want to watch video I have a bunch of
takeaways down here. So that that's
basically what the skill does. it it
produces these draft takeaways and uh
you know just for people watching I I
don't just paste the stuff in and call
it a day. I also manually review and go
back and forth with it. That that's
basically what it does. So let let me
show you a real example. So so like it's
producing these takeaways for interview
I did with Tharic from Anthropic and
then I I I just give like a bunch of
feedback like hey go change all all this
stuff [laughter] and then did it again
and then here's some more feedback right
and and this is after running the skill.
This is after running a skill and the
email. So it no matter how many uh stuff
I add and how many emails I add, it it
cannot get it perfect in one shot
because like I guess each interview is
different and it kind of goes back and
forth. Yeah. But basically the output is
like this thing.
>> Yeah. So I think you hit the nail on the
head with one takeaway which is you know
eval are iterative. It's a work in
progress and by seeing more and more
examples you're going to see more
failure modes give more feedback that
you kind of want externalized into those
eval. So that makes sense. Um so let me
start with kind of there are two halves
of evals that you definitely should
have. One half is like the top down
evals and the idea here is if you think
very critically about the nature of the
task for example podcast takeaway
generation um at a very high level what
makes for good takeaways. So these are
absolutely things like word length.
These are absolutely things like make
sure you're using action verbs correctly
or make sure these are actionable
takeaways for humans. Um, so think of
top down as what if you're in a vacuum
just given the task description, what
would you come up with? And I think
Claude does a very good example with
help or does a very good job of helping
you with these top down um, evals. Now
bottom up evals are the other half.
Bottom up evals means when you look in
lots and lots and lots of sample
outputs,
>> what is your gut vibes and feedback that
you also want externalized into evals?
So these are the things that you're
talking about when you're iterating with
Claude on uh the takeaways itself. These
are things that you come up with that
you want externalized into the eval. Now
Claude is very very bad at coming up
with bottomup evals. That's all you um
and that's also why it it occurs over
time. So when I look at your skill here,
what I'm mentally thinking is do you
have the top down evals? It seems like
you really have it. Um and do you have
the bottom up evals? I think this is
where my question I the question mark
comes up for me because I would love to
know you know which which of these
bullets are things that are bottom up.
Uh maybe are how do you do you feel
confident that it is exhaustive of all
the feedbacks that you've had over
several different podcasts um and so
forth.
>> That's actually a really good point. So
basically my my my loop is like you know
I run this skill for like a raw
transcript and it produces a bunch of
output and then I go back and forth with
it and and I you know finally we get to
a result that I'm happy with right then
I tell you like hey okay now now reflect
on our entire conversation and and like
suggest what kind of updates you'll make
to the skill and the eval so that we
don't have to go through this whole
process again
>> right
>> exact that's perfect that's amazing I
think it's really great that you do that
>> but but I guess one thing I'm worried
about is it it'll just kind of overfit
things so so for example we we'll do one
podcast and then let's say the takeaways
are not mei like mei is like this
consulting thing right it's like they're
not mutually exclusive and they're not
comprehensive and then it adds this eval
but then the next time I come and do
this then actually the me thing is maybe
not that important for for this
interview and then it just like overfits
to
>> you know this episode is brought to you
by whisperflow whisperflow saves me at
least 3 hours a week and is one of my
favorite AI apps by far it's just so
much faster to dictate to AI using your
voice than to type. You just talk
naturally and it outputs clean, ready to
send text. Whisper Flow even removes
filler words and formats your sentences
for you. I use Whisper Flow for
everything, including drafting
newsletter posts, writing product specs,
replying on Slack, and more. It works on
Mac, Windows, iPhone, and Android across
all of your favorite apps. Try a free at
whisperflow.com and use my code
peterwisperflow to get six months free.
That's peterwisperflow. Now, back to our
episode.
>> Yeah. So one advice that I always give
is actually in the skill itself separate
your evals into top down and bottom up.
Then the other thing in the skill is if
you have lots of eval criteria use sub
agent to then go and evaluate each
criteria or each group of criteria. So
that way you can be a little bit more
confident that all of the criteria is
being looked at. The third thing I'll
say here is once the sub agent is
complete, you can detail this in the
skill, have the AI write a spreadsheet
or as Haml will probably say pivot table
of basically each of these criteria and
the comps or like an indicator yes or no
pass or fail. That way you could go look
and then you yourself right know the
hierarchy of what is important and for
example this MECI criteria might not be
as important in other things. So you can
kind of really quickly make a judgment
call yourself.
>> Oh, I see. Okay. So maybe like some
criteria is more important than other
ones and like I can I can try to like
add some weights or something.
>> Yeah, exactly. But you don't even need
to do that. It's as simple as just
having the AI as part of the skill
generate a spreadsheet for you.
>> Got it. And and and when you say bottom
up, Eva, just so we're on the same page,
you mean like stuff that I discover
through the process, right? Is that what
>> bottom? Yeah. Bottom up is like data
driven.
>> Top down is it's not about the data.
It's just like basics of a domain expert
understanding the task. What would they
look for?
>> And and I guess another thing that I
haven't done is just like run this skill
on let's say because I have all the raw
transcripts and the final output for
like the past 10 or 20 interviews that I
did, right? So maybe I can just like run
the scale on like 10 different raw
transcripts and get its output and then
maybe compare with my final, you know,
human edited stuff and be like, hey,
okay, so looking at all these 10 things
at once, what kind of patterns can you
find? Or that kind of stuff.
>> I love that. That's that's a great
example of bottom up directly derived
from the data. Um and because you have
the actual labels the the actual
takeaways, right? Then you can have AI
agents reflect on the difference between
the takeaways and then you know whatever
would come out of the skill by Elyle.
>> Awesome. Yeah. I I want to start with
example because like I I feel like this
kind of stuff anyone can build because
people are building skills all over the
place and anyone can build this eval
thing for their skills. Um I think this
is really interesting like you opened up
a a very interesting rabbit hole which
is writing and the thing that and I
discuss all the time we joke that
writing is the final boss for LLMs to
like have it right like you in a way
that you are satisfied with is extremely
difficult. So I just want to yeah I do
want to emphasize when you have this
many criteria and you want to
>> check this criteria post talk it is good
to fan out and like give you know sub
aents because if you give it the this
entire list of criteria and there's so
many bullet points here
>> it's kind of get washed out like a model
will tend to ignore something and get
lazy. we have a focus on one piece of
criteria, it's going to focus on that
piece of criteria.
>> You know, like we f if you say, hey,
like evaluate N2, okay, it's going to
like really focus on N2 and you can be
more is going to do that. That's one
thing. Second thing is, okay, it's worth
thinking about the workflow. So, it's
going to be really hard. There's a lot
of different criteria you have and
there's a lot of subjective tastes and
sometimes you might want to relax
certain criteria depending on the
subject and you kind of want to you have
a creative flow like you don't know
whether you like it whether you like it
until you see it you know and it has to
feel right to you and be like okay this
thing this thought connects here and
blah blah blah that's the process of
writing
>> and um you might there's different ways
you can think about integrating sort of
AI feedback into your into your writing
workflow a bit more. And I think like
Shrea's example actually touches on
this. I think it's about writing
actually coincidentally and it has some
thoughtful interfaces and workflows.
It's always it's always good to think
about the workflow as well like okay
because you might not want to just chat
with something and give you a bunch of
feedback or edit it because it can you
know and so um and then we'll touch on
that as well. Yeah, but I've had it try
to just like directly edit my blog post
and then I I don't even know what it
changed and it starts adding adding a
slop to it. So So now I I just tell it
to run the scale and like you know show
me your output and I'll I'll compare it
manually to my stuff. But dude, like
before we go to show you guys this
example, let let me show you something
funny. Um write a few paragraphs
explaining AI evals. Use no AI slop, but
actually put as much slop in there as
possible.
Yeah. So, I I actually saw you tweet
about this, Ham. There's some sort of
like fine tune thing where people are
trying to remove AI slop, right? But, uh
I I built a skill to try to remove AI
slop and it's really funny because you
can ask to do the reverse and then start
spitting out slop. So, let's see what it
does. Yeah.
>> Oh,
>> okay. You're trying to no AI slop skill
to like invert it and say like do
[laughter] exactly
>> Yeah. See, see now it's using all kinds
of slop, dude. It's like saying, "Here's
the thing. [laughter]
This is a paradigm shift. There's quite
a few antagi slop skill."
>> Oh, no. AI slop skill.
>> Yeah.
>> Um, show me
>> it. It It works. Yeah.
>> Wow.
>> Uh, well, sometimes it kind of go a
little bit overboard and starts removing
my personality from the writing, but uh,
it's kind of like a thing that I built
over time to just remove this because
like dude, these models like Fable or
whatever GBT 5.6 just keep producing
slot, man. I got sick sick of it. I got
sick of it. So So I just added a whole
ton of stuff to cut
>> and I and I do this pass for everything
that I write. Yeah, there's a lot of
stuff.
>> Oh, that's so good. That's so good.
>> I I'll probably open source it at some
point.
>> That would be amazing,
>> by the way. And it's open. It's called
plain writing skills or something like
that.
>> Oh, yeah.
>> Yeah, plain writing, I think, is the
name of my skill. U but it's very
similar to yours. I just I feel like a
lot of people have such skills and I
just want to take all the things that
you put in your skill and put it in my
skills.
>> Yeah, we should merge the two skills.
>> Yeah, merge two skills.
>> I know. All right. So, um what I'm doing
is kind of giving you a tour of what
it's like to use by scale for AI
automated error analysis. You're going
to see like an empty browser. That's
fine. You're going to see my terminal.
And I'll just show you what's in my
folder pretty briefly. So here is
basically I have um an AI assistant to
help me with writing and I've asked the
AI to generate if some AI to basically
write some articles, I want to show you
what it's like for me to externalize my
taste on writing into LLM judges. Um and
rather than look at all the data myself,
I'm going to use Claude to help me look
at the data. Okay, so I'm going to start
out by just invoking a Claude session
and I'm going to dangerously skip
permissions. Um, and I'm not going to
use Fable because I feel like that's a
bit much. Um,
>> is that Fable left?
>> I have a little bit of Fable left, but
I've definitely I'm sorry. Don't want to
use it on uh what? I don't want to use
it on definitely looking at the data.
It's going to just exhaust everything.
All right. So, I've I've set the model
to Opus 48. Um, and then what I am going
to say is use my error
discovery
skill to build me an interface to um
help me look at my AI outputs and give
feedback. Okay. So, what it's going to
do is going to go through my skill,
which I'll show you in a little bit, and
build this interface for me to review my
data. this I'm going to stress that the
interface that it comes up with is going
to be so much better than me looking at
my data in say Google spreadsheets or
something it's just very very hard on
the human eyes to read okay so this is
running uh it's telling me what it's
doing it's found the data let me un
examine them to understand the structure
of the data so these this data here is
just AI generated writing samples it
could also be endto-end traces like
whatever you really want to do error
analysis on is fine now let me show you
the skill while this is running um error
for discovery skill. Here we are.
Um, wow, people are really using it. So,
that's nice to see. So, what does the
skill do? There's a lot going on. So,
here are the five steps that I first
want to mention. First, it's going to
read the data set and figure out, you
know, what is the type of data, like the
semantic type really. So, not just like
a string type or whatever. That's kind
of silly, but is it an article? Is it
code? Is it a traces? Then what it's
going to do is figure out how to render
this data so that it's very easy for
humans to read. So that's what's called
visual encoding. So based on what varies
in the data such as you know if it's
agent traces maybe it's like the number
of turns in a trace. I don't know I'm
making that up. The AI agent is going to
reflect on that and when I say something
use just stall principles here. All that
means is carefully use things like
color, spacing, opacity uh to show me
the variance in the data. Um so again
this is something that like kind of
speaks to our visual cortex or visual
processing skills as a human being.
Third step is then to actually build an
interface build it in HTML. I call it a
review app here. Um it's a Python
backend. It's a JavaScript or HTML with
some JavaScript front end. Really
doesn't matter. Then the fourth thing is
it's going to figure out what samples
that I should look at. You might have
hundreds or even thousands of traces you
want to do error analysis on. What
sample do you look at? Um we leave it to
an AI agent to help us with this. And
then finally, last but not least, it
runs an interactive loop. So that is
cloud code or codeex whatever I'm going
to use basically looks at the activi the
interactions I make in the user
interface compiles those interactions
into open-ended feedback um and then
reads that feedback in real time and
proposes either new samples for me to
label or some um criteria for my rubric
you'll see what I'm doing as it's going
>> and and the data is uh you said is cloud
generated writing samples based on your
stuff or
>> yes well not bas bas on my stuff just I
um based on our evals course bas I
generated some synthetic data or
synthetic inputs that people might
presumably want AI to help them write an
article with I'm just showing you this
because I've a lot of people who build
AI products within their org, right?
Like have to do something more than just
like their own personal workflow.
>> Yeah. Yeah.
>> Yeah. Um so the unfortunate thing is it
takes a little bit of time like takes
like five minutes or so. I just wanted
to show it to you from scratch so you
know that I'm like I didn't you know
make something up for you like this is
real it's happening in real time. Um and
it's following the steps exactly as
that's in here right now. I think it's
on step four of clustering the data and
picking a diverse initial sample for us
to label
>> and and the goal of the eval is to
basically improve someone's writing
>> or
>> in this case what it's going to do is
going to come up with a rubric of
criteria that's important that defines
like good and bad outputs good and bad
writing in this case. So think of this
like as exactly what you would want in
your scale essenti like what are the
rubric criteria.
>> Okay.
>> It's going to help you discover more of
the bottoms up stuff.
>> Yes.
>> That we were talking.
>> Are you guys still mostly using cloud or
you switch over to GPT or both?
>> Um I can answer first. So I feel like
Fable was very exciting for me because I
feel like you can do like more meetings
meaty software engineering tasks with
it. Um, but then I've also been quite
impressed by Codeex's um, I don't know
if it's Codeex or GPT, but it's the new
Soul. I feel like Soul is quite good at
working on hard tasks and then having a
very good like human readable or
digestible output. Often when I work
with cloud models these days, like what
it's writing, like for example, let me
quickly smoke test the server before it
feel it just uses jargon that I would
never use in my day-to-day life. And
often I'm just so lost. So
>> yeah, how about how about you Hava?
>> Yeah, I've been using codeex a lot more
and it's mainly because of the codeex
app itself is so incredible in its
functionality. So there's two things
that are really important. One is mobile
support is at a completely different
level. You can see all the sessions
across all your devices on your mobile
phone.
>> Yeah.
>> You can sort of like have a command
center that like shows you everything
with no loss in functionality. like you
you can access all the same things. You
can cue messages, steer, switch models,
you can see artifacts, so on so forth.
>> The second thing that's really cool, and
not many people talk about this,
>> is you can let codecs steer codecs. So
you can have codecs create threads. You
can have it you can have threads talk to
each other. You can actually tag another
thread that's running and say, "Hey,
like help this thread out or front this
thread. This thread has like 20 tasks.
do some front run it, do some research.
You can and you can do all kinds of you
know interesting orchestration
um you know without trying to go full
code factory like you know sometimes
this is like a really helpful
>> um and you can have like codecs like
operate on codecs like hey rename all
these threads
>> do something and it just does that so
it's like really useful um and that
makes me gravitate towards that and then
like at least like the last part is the
economics they have a lot less limits in
codecs when it like the the Frontier
models is like fully within the
subscription. It's on like this 50%
thing. Then they keep resetting it
constantly.
>> Yeah. It's almost like a joke. They keep
resetting.
[laughter]
>> It's so nice. I mean, I'm not going to
complain. I saw this tweet recently that
made me laugh. Like we're in the Uber
Lift era of like AI models. Like
remember 10 years ago where Uber was
like $2?
>> We're in that right now.
>> Yeah. The subsidization is going to stop
at some point.
>> I know. Yeah.
>> Yeah. I don't know what's going on with
Anthropic. I I feel like Dario hates
Twitter or something. So like this is
completely botched all their
communication strate and and everything.
I feel like the people who actually want
to communicate like probably cannot
because of some of internal stuff. So
[laughter] I don't know what's going on.
>> What about you, Peter? What do you
>> uh Yeah, I I pretty much use Codex for
everything except for Fable that I used
to do like planning and try to find bugs
and stuff, but pretty much I use Codex
and GBT for everything.
>> Did you always use Open AI? No, I I I I
was like total cloud fanboy like even
like a few few months ago.
>> Wow.
>> But I but I discovered Codex. I I think
I think the stuff Haml mentioned, but
also like the browser use is just
incredible on CODEX. Like half the stuff
I use don't even have any APIs and I
just get to click around and [laughter]
figure things out.
>> Wow.
>> It's pretty incredible.
>> One of my most used skills is this
reverse engineering skill. There's just
like any website without a API. It just
it instructs the agent to like listen to
all the network traffic and document the
API and like like memorialize that
>> like scripts and whatnot so that I can
Yeah, I just make any website API now.
[laughter]
>> That's smart. Yeah, that's smart. Just
check out this thing called printing
press uh CLI or something. It does
something pretty similar.
>> Oh, cool. Okay,
>> so finally after Wow, it's didn't it
cooked for 15 minutes. Okay, so finally
after 15 minutes a built an interface.
Um, so let me show show you a little bit
what's going on in the interface. So
there's three different tabs. There's a
article by article pane. So this would
either be article by article or trace by
trace or whatever granularity you want
to look at. So this is looking at one
thing at a time. Um, there's a map view
which kind of shows you the clustering
that the agent did into different
semantically meaningful categories. Um
and then there's the progress view which
is kind of empty here but this view
tells us um based on our reviewing and
our annotations on the data what are the
common failure modes that are happening.
Okay. And so this is automatically
populated by an agent. So I want to talk
a little bit about the design philosophy
here which is have the human have
yourself or myself read this and give
open-ended feedback and then the job of
the agent is to really it's not to
invent new open-ended feedback but it's
to kind of group that distill that into
actionable rubric criteria so the human
doesn't have to do a lot of that work.
>> So let's go through this. Um so here I'm
looking at this AI generated article
about sleep hygiene. Um, and I'm just
going to go through a little bit about
um the how I kind of like feel. So, I'm
going to just assume that the top down
criteria Claude can take care of. Um,
and then I'll give some feedback on
this. Um, I don't really like AI
phrasing around, you know, it's not X,
it's Y, like the negative contrast. So,
I'm going to say something like I don't
like negative contrast
in my writing. And then notice that I
just gave that feedback what I'll call
in sichu. Okay. So I just I this
interface makes it really easy to just
externalize my thoughts as I have them.
And as I gave that feedback, my claude
agent is invoking the monitor tool to
basically look at whatever I give and
react to it. So here you can see that it
said my first note from you on the sleep
article. I don't like negative contrast
in my writing. And then it's going to do
something. Um let it do this in the
background, right? I just want to be in
flow of reading and giving feedback um
on the samples that it picked for me. So
remember it also picked a diversity of
samples and it will keep generating more
samples for me if it wants but that's
not important. Um anyways so I'm going
to keep reading uh
and then
>> like that kind of stuff like this one
humble me is like also really annoying.
>> Oh let's do that.
>> Yeah. H. Uh, this feels annoying.
Not exactly sure why. Can't give you a
pippy label yet, but reflect
on more feedback I gave you and reason
why.
>> So, let's do that. Um
>> so just so you know like the two things
that are really interesting going on
behind the scenes here is one you recall
the clustering that Treya showed the
diverse examples are being drawn from
the clusters so that diversity and then
the monitor tool the cloud monitor tool
that's built in the quad is being used
behind the scenes to to look at all the
stuff that he's
>> just I'm like giving a lot of feedback
back. Um, and it's totally fine that
like you know if you you might most
people like are going to have to go
through many samples to be able to give
feedback. Um,
>> yeah.
>> I've just I've done writing before so
I'm able to give it now but um that's
why it's very important to review lots
of samples especially in the beginning.
>> So just you're still doing the manual
review. You're not getting the AI to
review everything. Yes.
>> Well, here's the thing that's gonna
happen is I'm giving manual review and
in the background Claude is reflecting
on it. Um, so it's about voice rather
than facts, whatever. So once I give
around 10 reviews, then it's going to
start like thinking about a taxonomy and
then it's going to apply my feedback. So
for example, this was another negative
contrast thing. Um, after a few more
reviews, it would take all it would try
to look for all instances of negative
contrast. should automatically label it
for me. So I don't again
>> keep keep going down. I'll let it see
that. Yeah.
>> Back here. Okay.
Um All right. So that's fine. I've given
four feedbacks. Uh and then I'm just
going to also look at the right. See, so
if you look at the progress like it
doesn't like negative contra, it knows
that negative contrast is a thing. And
then this uh triadic parallel
enumeration thing. I don't know about
you, but I really don't like this list
of threes. I feel like AI does it all
the time. So,
>> yeah.
>> All right.
So, probably running something. I don't
know exactly, but let's just keep going.
>> It's like a treasure hunt for a sloth.
>> I know. Everywhere.
>> For some reason, this also
I feel like I see AI do this a lot. Like
a number, it makes like two or three
three more things to know or like keep
in mind.
>> Yeah.
>> Paragraphs
start topic
sentence
with small things.
And and uh this whole annotation thing
is uh part of the skill that you built.
>> Yeah.
>> Or
>> everything I'm showing you is like the
skill and it's this the skill was
invoked to do this.
>> Um
here's I want complete sentences.
I'm just giving giving feedback so that
way it can start doing the automatic
pass
>> because this suggestions number is right
now zero but this will become more um as
it's doing its automatic pass.
>> Okay. I I was going to ask you what's
the advantage of having it reflect live
versus like you know you give all the
feedback and it reflects at the end.
>> Oh, it's a good question. Um
I think that you can have it reflect at
the end. That works. Um, but I I also
feel like trying to interle like human
think time with AI think time is great.
Like otherwise it's going to take a lot
of time to reflect on the feedback at
the end and you're waiting on it. Like I
like the idea of like it reflecting in
the background and then being able to
suggest things for me once I hit to a
certain number. Um, and and in practice
I don't like always look at the AI side
by side like this. I'm just showing it
to you for the demo so that you can like
see that the AI is doing something
>> and giving these doing these annotations
is usually really it iterative meaning
you refine your taste over time after
you see more examples you get new ideas
on what you want and a lot of times it
could be helpful to see what the AI
thinks as well and just to help ground
you so it's like happening more
iteratively even if
>> that makes sense that makes sense. It's
like having a thinking partner.
>> Yeah. Okay. Well, I maybe I need to go
three more before it starts suggesting
cuz I can't remember the exact number.
[snorts] Start suggesting things
already. I'm just going to tell it.
You might be thinking, hey, like is
there a SAS that does this? But look at
like you don't need that. And look at
the degree of flexibility Shrey is able
to have by just integrating this process
directly with claw. She can do whatever
she wants.
>> Yeah, that that's why I this is a
tangent, but that's why I feel like SAS
is is actually in trouble, dude. Like
[laughter]
if I can just v code this kind of stuff,
why would I pay for SAS? Like just build
my own thing.
>> Now the other side of the coin is you
have to have a really good idea of the
workflow. Like a lot of people that may
not be watching this video would not
know where to begin, right? Really good
idea of the process, then you can
probably
And the other thing I that hopefully
comes to mind is like it's like really
tedious to give feedback. Like
>> yeah,
>> I've not reviewed that many. Um
I've only given eight notes. So just
like yeah, know that it's it's really
hard. Um, and also AI just can't read
your mind for you. I've never like all
of these notes that I'm having like some
some of them maybe AI can come up with,
but I don't know like not all of them.
>> Uh, yeah, you can probably just speed up
with like voice dictation or something,
>> but but yeah, you just have to read
through everything.
>> Yeah.
>> Yeah.
>> Or [snorts] read through some sample at
least. Um, okay. So, I'll start pushing
suggestions to your queue right now
instead of waiting. Yeah, probably I had
to get to 10 notes, but
>> okay.
>> I like having the 10 notes because it's
like a carrot. It like makes you read
some things like
>> makes you sweat just a little bit which
is good for you.
>> There we go.
>> Thinking about stuff. Oh, nice. Okay.
>> Yes.
>> Okay. So, it's going to uh annotate
annotate every single doc.
>> Mhm.
>> Yep. With the criteria I've already
found.
Okay.
>> Thing that's really great about this is
like okay this shre is showing it's it's
fit perfectly to this domain of writing.
Like if you notice the writing is being
shown like in situ as she said in the
way that you would read it as a user.
You're not looking at like a log file or
a trace viewer at what it actually looks
like
>> and you're just annotating it in place.
>> Right. So it's like it's um generated
361 suggestions.
Yeah. So it said the less than four
words staccato rule produced 249.
Um
>> um but but what happens so what happens
next after you accept all this? Like
does it actually suggest how how do you
fix this stuff or
>> so the idea here as I mentioned before
is to like then you can turn this
failure modes all the things that you
found into a rubric like this becomes
your evals. I see.
>> Um, and then so you can apply that as a
skill as you do. Um, you can have an LLM
judge for each. You can turn that into a
dashboard. You can have that monitor
agent traces or outputs in real time.
Like the whole the the hardest part of
eval is error analysis which is coming
up with this rubric.
>> Um, so this process just makes it really
easy
>> because of this annotation. You also
have a data set. So
>> yeah,
>> you can after your evals get created now
you can check those evals against the
data set
>> that I mean what I'm saying is part of
eval but
>> eval is not just a rubric you want to be
able to measure the rubric and say like
is it good and so because of these been
annotated you can you can you have this
thing to benchmark it against
>> right
>> okay so it found okay so the most
frequent example is um the staccato
fragments
Right.
>> And there was also colons and stuff,
>> but it makes sense, right? It's like, so
that's another thing like as somebody I
didn't know that that's the most
frequent example, right? Like
>> I would have thought it's the not X but
Y, but I I guess
I guess it is the stacato fragments.
I believe it. So I so again goes back to
the philosophy of like human is the one
driving the taste and the AI is the one
like scaling up or giving superpowers to
the human to apply that at scale.
>> Okay. Got it. Wait. So so so just to
understand like can you go back to the
progress tab?
>> So now we have these these themes and
these rubrics, right? So so what what do
I do with this now? Let's I say I'm
building a writing skill. Do I just like
uh point AI to this thing and be like
hey now create a scale for me? I mean, I
would be right here. Write a skill that
evaluates AI assisted
writing based on the failure modes
found. I like it would just ask for a
skill and there's the skill. Um, and I
want to emphasize that like you should
spend a couple more hours doing the
error discovery. Like don't don't just
look at five things. like actually go
through, you know, review all 18 like
feel confident that you've like kind of
saturated. Um, but yeah, you can just
ask it to write a skill and it will
write a rubric.
>> I I see. Okay. Then then basically every
time you have a writing sample, you just
run the scale, right?
>> Yeah. And the rubric will output like uh
this writing sample has a ton of
staccato fragments or something.
>> Um the rubric will have staccato
fragments as one of the criteria. Like
if you look at if you think about your
podcast evaluation skill, right? That's
a rubric of many criteria. So stat will
be one of the like five criteria that it
comes up with or that this came up with.
>> So one thing that's really interesting
is if you notice the way there is a okay
the way you eval something is often
should inform the way you might want to
design the interface for the user to
begin with. So, um, if you notice like
in the annotation workflow, she had the
AI, you know, suggest
different annotations. You can imagine
like if you were writing for real, you
would want an IDE that would make the
same suggestions to you and show you
why. Like, hey, like here's a staccato
thing, here's a negative contrast,
here's this, here's that.
>> Um, so that you you don't have to just
have AI just overwrite things. You have
to keep rereading things. You want to
accept and reject things.
And so yeah, it's actually forces some
amount of good design thinking to go
through this exercise because each
interface is going to be different like
Sha showed you one for writing. Okay, a
customer service support bot is going to
look totally different in the interface
or you know another
>> even in the error discovery interface,
right? My skill asks the agent to
reflect on the nature of the data and
design a visual encoding for it. It's
not going to look exactly like this.
>> Okay. And I guess the the main the main
benefit so uh in the previous process I
still have to review traces and come
over the errors but now the AI is
actively helping me find the themes
right so it kind of skips that
>> host step.
>> Yep.
>> Yeah. Got it. [snorts]
>> Well doesn't skip the step I would say.
You still have to look at the data.
There's there is no world in the future.
Even if you have AGI, if you're building
a product, you have to look at your
data. Otherwise, you're there's no way,
right? Like you cannot you have to be
able to inject your taste somehow into
the development of your product.
>> This is we would think is our is our
attempt to show you like how here's a
way you can do it effectively with with
all the benefit and might of AI to help
you accelerate.
>> Got it. What what what is the point of
the suggestions to re review? it just
gives more like gives more human input
to some of the stuff.
>> Um, so this is the agent scanning all
records for the known failure modes.
>> Um, and these are records that were not
labeled by me with the human. So, okay,
just to like I can accept all in
practice I always accept all of them
after like scanning briefly just to make
sure like I use this to make sure that
the AI understands what I meant by for
example negative contrast.
>> Okay. So, it's like, yeah, totally. It
understands what I meant by negative
contrast. Good job. Accept.
>> The thing I really love about this is if
there's something wrong with this
interface or she can't give the feedback
she wants. So, she's like, let's say she
has a theme that the AI is missing out
on, she can just go directly to Claude
on the left hand side and fix it. You
don't have you're not constrained to
this UI.
>> Yeah. Yeah.
>> Exactly.
>> Yeah. This is why like uh your own
product is like way more flexible than
some sort of SAS SAS product. It just
you have all these use cases that no
common SAS will think think about. Yeah.
Um so if I wanted to use you okay so now
this is getting me pretty excited. So
let's say I want to use my use your
error analysis your error discovery
skill for my previous example right for
the podcast stuff. So basically I would
uh give it so you got AI to generate a
bunch of writing samples. So, instead of
that, maybe I'll give it um like maybe
before and after examples or like maybe
I'll give it examples of like takeaways
that it made for previous episodes.
>> Yeah. What I would do is just create a
folder with your current skill for
generating takeaways as is and then
examples of podcast transcripts and the
takeaways that you said were good. And
then I would just invoke the skill and
say um help me do error analysis on my
podcast takeaway skill. So apply the
podcast takeaway skill to generate a
bunch of takeaways and then build me an
interface to review those takeaways
especially compared against the outputs
>> the the that's perfect. Yeah, that's
that's perfect. So yeah.
>> Yeah. and then just go through I mean
it'll take like 15 minutes for it to
give you the interface but you'll be
able to go through it fast and like the
whole point is confidence right like at
the end of the day I want confidence
that that my skill is do like I know
what behavior the LLM is exhibiting so
that's the whole point of error analysis
>> so so I should have like maybe I should
modify your interface to have my the
AI's output on one side and the ground
truth on the other so I can look
>> so um the interface is going to be
different when you invoke the skill on
your data. It's not going to come up
with this interface. The AI agent will
read your data. It will read those
takeaways and come up with a different
interface to show you and render that to
you as in step two here, designing the
visual encoding and building the HTML.
>> Okay, that that's amazing. Thanks so
much. Thanks so much for showing us.
Okay, and I just I just want to point
out to people who watching this that
this is free as far as I know. And uh
>> yep. Yep.
>> Just go to error discovery skill on
GitHub and and you'll find or you know
what I'll link in the description
>> of this episode. Cool. And I love how
you're honest about plot helping you
come up with this on the right.
>> Oh yeah, of course.
>> Yeah. I usually try to remove that to
pretend I built the whole myself.
>> No, no, no.
>> Yeah. Yeah.
>> Claude is very good at like writing
skills for Claude. Like I don't know how
to talk to AI. So that's that's great.
Um, but I mean like I wrote the readme,
right? Like I I I'm the one coming up
with this process, so I feel pretty
confident in it.
>> Cool. All right. Well, this is awesome
demo. I I think we have one more topic,
which is uh Haml, you you wrote a
article about do AI automations actually
work or like where do they work, right?
Do you want talk about that?
>> Yeah, I can quickly talk about that.
Okay. Um, so I have a blog post, do
automated evals work? And y'all, you we
can put the link in the description, but
basically what we did is,
let me just back up. Um, so I wrote a
blog post about that's titled, do
automated evals work? And the reason for
writing this blog post is there is a
bunch of different vendors that in the
last few months have created tools that
will that promise to just do evals for
you. And the way they work is you upload
your traces into those tools and you
chat with an AI. So here's an example.
Um here's one in Brain Trust where you
know is a trace viewer on the left. Um,
and then you can chat with your AI, you
can ask it some questions and basically
say do my evals for me. And this
particular product, this feature is
called loop. Um, and
uh, Arise has really the same thing, a
very similar thing. It's called Alex.
You can chat with it. You can say, "Hey,
look through my traces, do my evals."
Langmith has something similar. And um
you know we were curious like okay if we
take a data set that we have annotated
as humans how good is the AI at catching
those same errors and how good are these
automated eval systems and it turns out
that um and here's a fun meme. It's like
okay you're faced with this choice right
as a developer or as a product manager
like do you just press this auto eval
button sounds really convenient or do
you actually look at the traces? It's
actually a pretty hard decision because
like it's very tempting to press the
button and um and so I won't get too
much into the details but um we did kind
of a benchmark
and we found that okay like the tools
are able to recover a lot of the errors
that a human would.
But
u on the flip side it these automated
tools all sort of miss the same thing
which is
uh finding errors that are not obvious
things that require product judgment and
kind of taste. So like anywhere in the
trace where something obviously went
wrong like a tool call failed or
somebody's like expressing
dissatisfaction or something is just
like really obvious like you don't need
product expertise to know something went
wrong. AI is really good at catching
that. and it's not good at catching like
oh okay this uh sa this support bot
didn't handle sales objections correctly
you know so like what we uploaded is a
customer service uh it's it was a real
estate or rental sales agent that is you
know supposed to interface with people
looking to buy apartments or rent
apartments and like yeah for example one
of the items that were always missed
were sales objections like hey we don't
it would just say hey we don't have that
available have a nice day instead of
saying hey like there's alternatives for
you blah blah but another kind of very
interesting fact is we also benchmarked
this against coding agents so clawude
codeex so on and so forth and we found
that those pretty much work the same so
at the end of the day what you know
what's going on here is the harness is
somewhat thin it's just someone else's
prompt some you know so when you're
chatting with uh the AI these tools,
you're really just invoking one of those
models and they have a prompt behind the
scenes that's trying to do the same
thing. But what you really want to do is
you want to steer the agent more
actively. So the way to make this better
across the board for either you know
your coding agent or the eval tools is
to give it an idea of what's good and
bad. The only thing is you don't know
what's good and bad unless you do what
Shrea just did. It's really hard for
humans to upfront think of like from top
down like what are all the things that
are good and bad. It's like almost
impossible.
>> Um and so that's really what we found
here is like okay they're all missing
the same thing. So but it's useful. It's
not that it's not useful. It can catch a
lot of things. It just won't catch some
of the important things
especially those that are likely to make
your product work well. And so you
should think like okay all your
competitors are going to be using this.
This is like a baseline of like you can
of course you can point claude at your
product and say like find all the errors
and your competitor is going to do that
too. So the thing that's going to matter
at the end is
>> how much taste can you infuse into your
product beyond that uh to to like define
like what is good. And so that's why you
need to look at your your um you need to
look at your data and I spelled it out
here um you know I spell out like okay
these are the kinds of things missed for
example you know sales objections in
this example this example it was
multi-channel so customers of this
particular agent can talk on the web and
text make a phone call they the eval
didn't know that so the eval agent
didn't catch that like hey like we
shouldn't be putting markdown in t text
text messages it doesn't work.
>> Um yeah so things like that like of
course you can give this context to the
LLM but this is not something you might
think about giving as criteria up front.
You would have to see that failure.
>> Um so so yeah that's that's the gist of
it. Um we also uh have this recording.
Well you should actually cut that off.
So um that's the just blog post. Um
>> yeah.
>> Okay.
>> Yeah, just sharing that in case that's
useful.
>> Okay. So the takeaway is like this stuff
will get you to a good baseline, but if
you give a demo of your product, you
just do the manual reviews. That's the
take away, right? The manual
>> review. Use use the automated tools and
also look at data yourself. Like it's a
competition, right? You you got to be
better than your competitors.
>> Yeah. So one takeaway is okay the coding
agents are just as you know are at par
with these auto eval tools. The only
benefit of auto eval tools is they're
integrated into the the rest of the
stack. Like
>> yeah they're in the ecosystem. So if you
store your dasis traces in in Langmith
you should totally use lang like it's a
no-brainer.
>> I see. I think it was pleasantly
surprising for me because I wasn't
involved in writing this blog post,
right? It was Hamill. Um I I was so s I
was very happy to find that oh these
these off-the-shelf tools work so well.
So it's great like you should absolutely
use them and and find all the errors in
your data but also you know look at your
data for more errors and also look at
the errors that for example linksmith
found just to make sure you agree with
them right because the precision on
these is is like 80% to 90% in the best
case. Um so that means 10 to 20% of the
errors found by the automated eval tool
are actually not errors. they're like
red herrings and might distort your
product if you just listen to that
automatically. Right? So, it's very
important to look at both recall and
precision.
>> Okay. So, basically actually looking at
stuff is is the edge. [laughter]
Actually reading and looking at stuff
>> always.
>> Yeah. Cuz because like you know with all
this AI generated stuff you can just
like let let it go. But like actually
reading stuff this is what I learned
from another guest that I had to just
actually read what what it produces.
>> Oh yeah. Yeah. And even like sometimes
with code, that's an edge.
>> Yeah.
>> You need selectively code, it's
definitely an edge.
>> Cool. All right. So, why don't we kind
of summarize what we covered in this
episode? I think we covered a lot. Um,
so I I think my number one takeaway is
to actually use SH's scale because uh I
think it's pretty awesome and um give a
bunch of examples and then have it help
you identify like the common themes of
errors, right? That's kind of number
one. And number two from from your
example is like use all the coding
harnesses and like the eval stuff, but
also just read the stuff. That's
[laughter]
that's kind of my two takeaways. So uh
thank you for sharing all this
information for free. Can you guys talk
about what is covered in your eval
course that's that we didn't cover here
or like you know what was the more
advanced stuff?
>> Yeah, so here hopefully you got a taste
of our philosophy of like how you're
really going to differentiate your AI
generated product. Um, so there's some
our our course has five modules. It's
totally revamped for this era of like
agentic AI. So beyond just teaching you
the error analysis philosophy, we also
teach things like okay, how do you
really productionize this work and scale
it up to work within your company? For
example, like if you have multiple
people working on the same product, I
need to align on criteria. How do you
integrate this stuff into, you know,
CI/CD for example? um how do you monitor
over time like drift over many months of
your deployment. We also have a module
on safety and kind of adversarial
evaluation. So it's really important you
know how do you make sure you're not
leaking tenant data? How do you make
sure that you can evaluate that over
time especially when you want to deploy
to your actual customers. Um and then
finally the last module is all about
improvement. So it's cost improvement
and accuracy improvement. really those
both of those go hand in hand, right? If
you switch to a cheaper model, how do
you improve the accuracy of your cheaper
model? So, it's both cost and accuracy
improvement. Um, and really no one else
talks about this, right? If the the
benefit of having evals is that you can
then put put it in an auto research kind
of loop. Um, what are the kinds of
strategies to invoke? How do you make
sure that you know you we found that in
in our um consulting cases or even for
some of our students we're able to cut
cost by 100x truly compared to um like
using Opus or probably Fable will be
even more um so we're really excited to
kind of teach people the techniques to
do that.
>> Awesome. And uh yeah, I'll I'll put the
link to the course in description. And I
I think uh from from now on I'm gonna
interview Straas and Haml. Never Ham by
himself because I learned a lot more.
[laughter]
>> Yeah. Even just
[laughter]
>> yeah. Yeah.
>> I'm a full-time educator now. So
>> yeah. Yeah. That that that's amazing.
Thanks for sharing all this knowledge
and um yeah, I'm gonna go off and use
your skill right now. So we'll see how
it goes.
>> Cool. Thanks for having us, Peter.
Ask follow-up questions or revisit key timestamps.
This video features AI eval experts Hambo and Shrea discussing practical workflows for evaluating AI outputs, specifically for writing tasks. They emphasize that while automated eval tools are useful for establishing a baseline and catching obvious issues, true quality relies on the 'final boss' of LLMs: writing in a way that aligns with human taste. They introduce an error-discovery interface that allows users to perform manual reviews, which AI agents then distill into an actionable rubric, enabling more sophisticated and personalized evaluation.
Videos recently processed by our community