HomeVideos

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Now Playing

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Transcript

1503 segments

0:00

You want to externalize your taste and

0:02

your judgment before you start writing

0:04

evals.

0:04

>> No matter how many stuff I add and how

0:06

many evals I add, it cannot get it

0:08

perfect in one shot.

0:09

>> Do you have the bottom up evals? I think

0:11

this is where my question mark comes up

0:12

for me. [music] Bottom up is like data

0:14

driven. Top down is not about the data.

0:17

Claude is very very bad at coming up

0:19

with bottom up evals. That's all you.

0:21

The interface that it comes up with is

0:23

going to be so much better than me

0:25

looking at my data in say Google

0:27

spreadsheets [music] or something. We

0:28

joke that writing is the final boss for

0:30

LLMs. To like have it right like you in

0:32

a way that you are satisfied with is

0:34

extremely difficult.

0:38

Hey everyone, my guests today are Hambo

0:40

and Shrea, the instructors behind the

0:42

most popular AI eval course online. I

0:44

think they've taught over 4,500 students

0:46

now. So, they're here to share a really

0:48

practical example about how to use AI to

0:50

automate evals the right way without

0:52

doing a whole bunch of slop. And uh

0:54

yeah, really excited to have you both.

0:56

>> Really excited. Likewise. [laughter]

0:58

>> Great. Okay. So, so Haml, last time you

1:01

came on the podcast, I feel like I

1:03

learned a lot from you. I think some of

1:04

the lessons were manly review customer

1:06

traces and conversations to figure out

1:08

where things are going wrong. Use like

1:10

yes, no, past fail evals instead of

1:12

making up random scores. But maybe

1:14

before we get to the demo, can you guys

1:15

talk about at a high level about how

1:17

eval like last year?

1:19

>> So, the fundamentals still apply. You

1:22

still want to start everything by

1:23

looking at data and you want to do error

1:25

analysis and you want to have a

1:26

structured approach to looking at data

1:28

so that you can externalize your taste

1:30

and your judgment before you start

1:32

writing eval. The main thing that has

1:35

changed is getting agents to help you

1:37

look at the data in a very thoughtful

1:39

way. So agents have become a lot more

1:41

powerful over time and we have figured

1:43

out a workflow that allows you to kind

1:45

of have agents running in the background

1:46

while you look at data supporting you so

1:49

that you can have a lot more leverage in

1:51

that first stage of looking at data. And

1:54

that's what Treya is going to show today

1:56

is an example of that.

1:57

>> The other thing I'll add to that is that

2:00

um we we teach LLM judges a lot in our

2:03

course. LLM judges are a way to have an

2:05

LLM look at a trace and make a judgment

2:08

about a very specific failure mode. So

2:11

for example, uh is is this output too

2:14

long or is this output too short? Does

2:16

it follow a particular structure? Um LLM

2:19

judges have gotten very good at

2:22

evaluating these very well- definfined

2:24

criteria more accurately. So we are much

2:27

bigger fans I would say of LLM judges.

2:29

>> Got it. Since I have two ex experts

2:31

here, I'd love for you guys to take a

2:33

look at my pretty pretty dumb evals for

2:35

a skill that I built and maybe give me

2:37

some feedback. Basically, I built a

2:39

skill for podcast post-production and

2:41

part of the skill is it writes a bunch

2:43

of it if tried to find the top 10

2:45

takeaways from the podcast and then it

2:47

writes an initial draft as a newsletter

2:49

post and I built a bunch of evals for

2:51

for this or I as in claude a bunch of

2:54

evals for this. So uh so for example if

2:57

you go down here there's some evals

2:58

around like are the takeaways the right

3:00

length between 240 and 330 characters

3:03

and then there's also some like judgment

3:05

things is each takeaway understandable

3:07

on its own without watch the episode

3:08

that does it give plain useful advice so

3:11

so I guess there's kind of two types of

3:12

take two types of eos right one is just

3:14

like checking if the character lens is

3:16

correct and the other one is more like

3:17

judgment based but yeah do you guys have

3:19

any thoughts about this or like how how

3:20

can I make this better? Yeah, this is

3:22

awesome. I love that you have a real

3:25

demo for us to take a look at. There's

3:28

some things that I want to back up a

3:30

little bit on, which is what's in the

3:33

skill, right? If we're going to start

3:34

doing evals, it's very hard to do evalu.

3:40

And uh what was your I love that you

3:43

tried to make evals. So then kind of

3:45

what was your thought process on making

3:46

the evals? Was it like I'm just going to

3:48

ask Claude to make some emails for me or

3:52

I don't know you saw some outputs or you

3:54

you saw some things that you noticed

3:55

were like off. Um and then you wanted to

3:58

create emails around that or like tell

3:59

us a little bit more about that.

4:01

>> That's a good question. Yeah. So um let

4:03

me pull up my little blog post. Um so

4:06

basically I do these interviews and uh

4:09

there's a video and then the people who

4:11

want to watch video I have a bunch of

4:12

takeaways down here. So that that's

4:14

basically what the skill does. it it

4:15

produces these draft takeaways and uh

4:19

you know just for people watching I I

4:20

don't just paste the stuff in and call

4:21

it a day. I also manually review and go

4:23

back and forth with it. That that's

4:25

basically what it does. So let let me

4:26

show you a real example. So so like it's

4:28

producing these takeaways for interview

4:30

I did with Tharic from Anthropic and

4:32

then I I I just give like a bunch of

4:34

feedback like hey go change all all this

4:36

stuff [laughter] and then did it again

4:38

and then here's some more feedback right

4:39

and and this is after running the skill.

4:41

This is after running a skill and the

4:42

email. So it no matter how many uh stuff

4:45

I add and how many emails I add, it it

4:46

cannot get it perfect in one shot

4:48

because like I guess each interview is

4:49

different and it kind of goes back and

4:51

forth. Yeah. But basically the output is

4:52

like this thing.

4:54

>> Yeah. So I think you hit the nail on the

4:57

head with one takeaway which is you know

4:59

eval are iterative. It's a work in

5:01

progress and by seeing more and more

5:03

examples you're going to see more

5:05

failure modes give more feedback that

5:07

you kind of want externalized into those

5:09

eval. So that makes sense. Um so let me

5:12

start with kind of there are two halves

5:14

of evals that you definitely should

5:16

have. One half is like the top down

5:19

evals and the idea here is if you think

5:21

very critically about the nature of the

5:23

task for example podcast takeaway

5:26

generation um at a very high level what

5:29

makes for good takeaways. So these are

5:30

absolutely things like word length.

5:32

These are absolutely things like make

5:34

sure you're using action verbs correctly

5:36

or make sure these are actionable

5:38

takeaways for humans. Um, so think of

5:40

top down as what if you're in a vacuum

5:43

just given the task description, what

5:45

would you come up with? And I think

5:46

Claude does a very good example with

5:48

help or does a very good job of helping

5:50

you with these top down um, evals. Now

5:53

bottom up evals are the other half.

5:56

Bottom up evals means when you look in

5:58

lots and lots and lots of sample

6:00

outputs,

6:01

>> what is your gut vibes and feedback that

6:03

you also want externalized into evals?

6:06

So these are the things that you're

6:07

talking about when you're iterating with

6:09

Claude on uh the takeaways itself. These

6:12

are things that you come up with that

6:14

you want externalized into the eval. Now

6:17

Claude is very very bad at coming up

6:19

with bottomup evals. That's all you um

6:22

and that's also why it it occurs over

6:25

time. So when I look at your skill here,

6:27

what I'm mentally thinking is do you

6:29

have the top down evals? It seems like

6:30

you really have it. Um and do you have

6:32

the bottom up evals? I think this is

6:34

where my question I the question mark

6:37

comes up for me because I would love to

6:38

know you know which which of these

6:40

bullets are things that are bottom up.

6:42

Uh maybe are how do you do you feel

6:44

confident that it is exhaustive of all

6:46

the feedbacks that you've had over

6:48

several different podcasts um and so

6:50

forth.

6:51

>> That's actually a really good point. So

6:52

basically my my my loop is like you know

6:55

I run this skill for like a raw

6:56

transcript and it produces a bunch of

6:58

output and then I go back and forth with

7:00

it and and I you know finally we get to

7:01

a result that I'm happy with right then

7:04

I tell you like hey okay now now reflect

7:06

on our entire conversation and and like

7:08

suggest what kind of updates you'll make

7:10

to the skill and the eval so that we

7:11

don't have to go through this whole

7:12

process again

7:14

>> right

7:14

>> exact that's perfect that's amazing I

7:16

think it's really great that you do that

7:18

>> but but I guess one thing I'm worried

7:19

about is it it'll just kind of overfit

7:21

things so so for example we we'll do one

7:23

podcast and then let's say the takeaways

7:24

are not mei like mei is like this

7:26

consulting thing right it's like they're

7:27

not mutually exclusive and they're not

7:29

comprehensive and then it adds this eval

7:32

but then the next time I come and do

7:34

this then actually the me thing is maybe

7:35

not that important for for this

7:37

interview and then it just like overfits

7:39

to

7:40

>> you know this episode is brought to you

7:42

by whisperflow whisperflow saves me at

7:44

least 3 hours a week and is one of my

7:46

favorite AI apps by far it's just so

7:48

much faster to dictate to AI using your

7:51

voice than to type. You just talk

7:52

naturally and it outputs clean, ready to

7:55

send text. Whisper Flow even removes

7:57

filler words and formats your sentences

7:59

for you. I use Whisper Flow for

8:01

everything, including drafting

8:02

newsletter posts, writing product specs,

8:04

replying on Slack, and more. It works on

8:06

Mac, Windows, iPhone, and Android across

8:09

all of your favorite apps. Try a free at

8:11

whisperflow.com and use my code

8:14

peterwisperflow to get six months free.

8:16

That's peterwisperflow. Now, back to our

8:19

episode.

8:19

>> Yeah. So one advice that I always give

8:22

is actually in the skill itself separate

8:24

your evals into top down and bottom up.

8:27

Then the other thing in the skill is if

8:30

you have lots of eval criteria use sub

8:33

agent to then go and evaluate each

8:37

criteria or each group of criteria. So

8:40

that way you can be a little bit more

8:41

confident that all of the criteria is

8:43

being looked at. The third thing I'll

8:45

say here is once the sub agent is

8:48

complete, you can detail this in the

8:49

skill, have the AI write a spreadsheet

8:53

or as Haml will probably say pivot table

8:56

of basically each of these criteria and

8:59

the comps or like an indicator yes or no

9:02

pass or fail. That way you could go look

9:05

and then you yourself right know the

9:06

hierarchy of what is important and for

9:08

example this MECI criteria might not be

9:10

as important in other things. So you can

9:12

kind of really quickly make a judgment

9:14

call yourself.

9:16

>> Oh, I see. Okay. So maybe like some

9:17

criteria is more important than other

9:18

ones and like I can I can try to like

9:20

add some weights or something.

9:21

>> Yeah, exactly. But you don't even need

9:23

to do that. It's as simple as just

9:25

having the AI as part of the skill

9:27

generate a spreadsheet for you.

9:28

>> Got it. And and and when you say bottom

9:30

up, Eva, just so we're on the same page,

9:31

you mean like stuff that I discover

9:33

through the process, right? Is that what

9:35

>> bottom? Yeah. Bottom up is like data

9:37

driven.

9:37

>> Top down is it's not about the data.

9:40

It's just like basics of a domain expert

9:43

understanding the task. What would they

9:45

look for?

9:45

>> And and I guess another thing that I

9:47

haven't done is just like run this skill

9:49

on let's say because I have all the raw

9:50

transcripts and the final output for

9:52

like the past 10 or 20 interviews that I

9:54

did, right? So maybe I can just like run

9:56

the scale on like 10 different raw

9:57

transcripts and get its output and then

10:00

maybe compare with my final, you know,

10:02

human edited stuff and be like, hey,

10:03

okay, so looking at all these 10 things

10:05

at once, what kind of patterns can you

10:06

find? Or that kind of stuff.

10:08

>> I love that. That's that's a great

10:09

example of bottom up directly derived

10:12

from the data. Um and because you have

10:14

the actual labels the the actual

10:16

takeaways, right? Then you can have AI

10:19

agents reflect on the difference between

10:20

the takeaways and then you know whatever

10:22

would come out of the skill by Elyle.

10:24

>> Awesome. Yeah. I I want to start with

10:25

example because like I I feel like this

10:27

kind of stuff anyone can build because

10:28

people are building skills all over the

10:29

place and anyone can build this eval

10:31

thing for their skills. Um I think this

10:33

is really interesting like you opened up

10:35

a a very interesting rabbit hole which

10:37

is writing and the thing that and I

10:41

discuss all the time we joke that

10:43

writing is the final boss for LLMs to

10:45

like have it right like you in a way

10:47

that you are satisfied with is extremely

10:49

difficult. So I just want to yeah I do

10:51

want to emphasize when you have this

10:53

many criteria and you want to

10:55

>> check this criteria post talk it is good

10:58

to fan out and like give you know sub

11:01

aents because if you give it the this

11:03

entire list of criteria and there's so

11:05

many bullet points here

11:06

>> it's kind of get washed out like a model

11:09

will tend to ignore something and get

11:11

lazy. we have a focus on one piece of

11:13

criteria, it's going to focus on that

11:15

piece of criteria.

11:16

>> You know, like we f if you say, hey,

11:18

like evaluate N2, okay, it's going to

11:21

like really focus on N2 and you can be

11:24

more is going to do that. That's one

11:27

thing. Second thing is, okay, it's worth

11:29

thinking about the workflow. So, it's

11:31

going to be really hard. There's a lot

11:34

of different criteria you have and

11:35

there's a lot of subjective tastes and

11:38

sometimes you might want to relax

11:39

certain criteria depending on the

11:41

subject and you kind of want to you have

11:44

a creative flow like you don't know

11:45

whether you like it whether you like it

11:47

until you see it you know and it has to

11:50

feel right to you and be like okay this

11:52

thing this thought connects here and

11:53

blah blah blah that's the process of

11:54

writing

11:57

>> and um you might there's different ways

12:00

you can think about integrating sort of

12:03

AI feedback into your into your writing

12:06

workflow a bit more. And I think like

12:08

Shrea's example actually touches on

12:11

this. I think it's about writing

12:13

actually coincidentally and it has some

12:15

thoughtful interfaces and workflows.

12:18

It's always it's always good to think

12:19

about the workflow as well like okay

12:22

because you might not want to just chat

12:24

with something and give you a bunch of

12:26

feedback or edit it because it can you

12:28

know and so um and then we'll touch on

12:30

that as well. Yeah, but I've had it try

12:32

to just like directly edit my blog post

12:34

and then I I don't even know what it

12:36

changed and it starts adding adding a

12:37

slop to it. So So now I I just tell it

12:40

to run the scale and like you know show

12:41

me your output and I'll I'll compare it

12:42

manually to my stuff. But dude, like

12:44

before we go to show you guys this

12:45

example, let let me show you something

12:46

funny. Um write a few paragraphs

12:49

explaining AI evals. Use no AI slop, but

12:54

actually put as much slop in there as

12:59

possible.

13:01

Yeah. So, I I actually saw you tweet

13:02

about this, Ham. There's some sort of

13:04

like fine tune thing where people are

13:05

trying to remove AI slop, right? But, uh

13:08

I I built a skill to try to remove AI

13:10

slop and it's really funny because you

13:12

can ask to do the reverse and then start

13:14

spitting out slop. So, let's see what it

13:16

does. Yeah.

13:16

>> Oh,

13:17

>> okay. You're trying to no AI slop skill

13:20

to like invert it and say like do

13:22

[laughter] exactly

13:25

>> Yeah. See, see now it's using all kinds

13:26

of slop, dude. It's like saying, "Here's

13:28

the thing. [laughter]

13:30

This is a paradigm shift. There's quite

13:32

a few antagi slop skill."

13:35

>> Oh, no. AI slop skill.

13:36

>> Yeah.

13:37

>> Um, show me

13:39

>> it. It It works. Yeah.

13:41

>> Wow.

13:42

>> Uh, well, sometimes it kind of go a

13:44

little bit overboard and starts removing

13:45

my personality from the writing, but uh,

13:48

it's kind of like a thing that I built

13:49

over time to just remove this because

13:50

like dude, these models like Fable or

13:53

whatever GBT 5.6 just keep producing

13:55

slot, man. I got sick sick of it. I got

13:57

sick of it. So So I just added a whole

13:59

ton of stuff to cut

14:02

>> and I and I do this pass for everything

14:04

that I write. Yeah, there's a lot of

14:06

stuff.

14:06

>> Oh, that's so good. That's so good.

14:08

>> I I'll probably open source it at some

14:10

point.

14:10

>> That would be amazing,

14:12

>> by the way. And it's open. It's called

14:14

plain writing skills or something like

14:16

that.

14:16

>> Oh, yeah.

14:17

>> Yeah, plain writing, I think, is the

14:18

name of my skill. U but it's very

14:20

similar to yours. I just I feel like a

14:23

lot of people have such skills and I

14:25

just want to take all the things that

14:26

you put in your skill and put it in my

14:28

skills.

14:29

>> Yeah, we should merge the two skills.

14:31

>> Yeah, merge two skills.

14:32

>> I know. All right. So, um what I'm doing

14:36

is kind of giving you a tour of what

14:38

it's like to use by scale for AI

14:41

automated error analysis. You're going

14:44

to see like an empty browser. That's

14:45

fine. You're going to see my terminal.

14:46

And I'll just show you what's in my

14:48

folder pretty briefly. So here is

14:51

basically I have um an AI assistant to

14:54

help me with writing and I've asked the

14:56

AI to generate if some AI to basically

14:59

write some articles, I want to show you

15:01

what it's like for me to externalize my

15:03

taste on writing into LLM judges. Um and

15:06

rather than look at all the data myself,

15:08

I'm going to use Claude to help me look

15:10

at the data. Okay, so I'm going to start

15:13

out by just invoking a Claude session

15:15

and I'm going to dangerously skip

15:17

permissions. Um, and I'm not going to

15:19

use Fable because I feel like that's a

15:22

bit much. Um,

15:23

>> is that Fable left?

15:25

>> I have a little bit of Fable left, but

15:26

I've definitely I'm sorry. Don't want to

15:28

use it on uh what? I don't want to use

15:32

it on definitely looking at the data.

15:33

It's going to just exhaust everything.

15:35

All right. So, I've I've set the model

15:37

to Opus 48. Um, and then what I am going

15:40

to say is use my error

15:44

discovery

15:46

skill to build me an interface to um

15:52

help me look at my AI outputs and give

15:57

feedback. Okay. So, what it's going to

15:59

do is going to go through my skill,

16:01

which I'll show you in a little bit, and

16:03

build this interface for me to review my

16:05

data. this I'm going to stress that the

16:07

interface that it comes up with is going

16:08

to be so much better than me looking at

16:11

my data in say Google spreadsheets or

16:13

something it's just very very hard on

16:14

the human eyes to read okay so this is

16:16

running uh it's telling me what it's

16:18

doing it's found the data let me un

16:21

examine them to understand the structure

16:22

of the data so these this data here is

16:24

just AI generated writing samples it

16:26

could also be endto-end traces like

16:28

whatever you really want to do error

16:30

analysis on is fine now let me show you

16:32

the skill while this is running um error

16:35

for discovery skill. Here we are.

16:40

Um, wow, people are really using it. So,

16:43

that's nice to see. So, what does the

16:45

skill do? There's a lot going on. So,

16:48

here are the five steps that I first

16:50

want to mention. First, it's going to

16:51

read the data set and figure out, you

16:52

know, what is the type of data, like the

16:55

semantic type really. So, not just like

16:57

a string type or whatever. That's kind

16:59

of silly, but is it an article? Is it

17:01

code? Is it a traces? Then what it's

17:04

going to do is figure out how to render

17:06

this data so that it's very easy for

17:08

humans to read. So that's what's called

17:10

visual encoding. So based on what varies

17:13

in the data such as you know if it's

17:15

agent traces maybe it's like the number

17:16

of turns in a trace. I don't know I'm

17:18

making that up. The AI agent is going to

17:20

reflect on that and when I say something

17:22

use just stall principles here. All that

17:25

means is carefully use things like

17:28

color, spacing, opacity uh to show me

17:32

the variance in the data. Um so again

17:35

this is something that like kind of

17:36

speaks to our visual cortex or visual

17:38

processing skills as a human being.

17:41

Third step is then to actually build an

17:43

interface build it in HTML. I call it a

17:46

review app here. Um it's a Python

17:48

backend. It's a JavaScript or HTML with

17:51

some JavaScript front end. Really

17:53

doesn't matter. Then the fourth thing is

17:56

it's going to figure out what samples

17:58

that I should look at. You might have

18:00

hundreds or even thousands of traces you

18:02

want to do error analysis on. What

18:05

sample do you look at? Um we leave it to

18:07

an AI agent to help us with this. And

18:10

then finally, last but not least, it

18:12

runs an interactive loop. So that is

18:15

cloud code or codeex whatever I'm going

18:17

to use basically looks at the activi the

18:20

interactions I make in the user

18:22

interface compiles those interactions

18:24

into open-ended feedback um and then

18:27

reads that feedback in real time and

18:29

proposes either new samples for me to

18:32

label or some um criteria for my rubric

18:35

you'll see what I'm doing as it's going

18:37

>> and and the data is uh you said is cloud

18:39

generated writing samples based on your

18:41

stuff or

18:42

>> yes well not bas bas on my stuff just I

18:46

um based on our evals course bas I

18:48

generated some synthetic data or

18:50

synthetic inputs that people might

18:52

presumably want AI to help them write an

18:54

article with I'm just showing you this

18:56

because I've a lot of people who build

18:58

AI products within their org, right?

19:01

Like have to do something more than just

19:03

like their own personal workflow.

19:05

>> Yeah. Yeah.

19:06

>> Yeah. Um so the unfortunate thing is it

19:09

takes a little bit of time like takes

19:10

like five minutes or so. I just wanted

19:12

to show it to you from scratch so you

19:13

know that I'm like I didn't you know

19:15

make something up for you like this is

19:17

real it's happening in real time. Um and

19:19

it's following the steps exactly as

19:22

that's in here right now. I think it's

19:23

on step four of clustering the data and

19:25

picking a diverse initial sample for us

19:28

to label

19:29

>> and and the goal of the eval is to

19:32

basically improve someone's writing

19:34

>> or

19:34

>> in this case what it's going to do is

19:36

going to come up with a rubric of

19:38

criteria that's important that defines

19:40

like good and bad outputs good and bad

19:42

writing in this case. So think of this

19:45

like as exactly what you would want in

19:47

your scale essenti like what are the

19:49

rubric criteria.

19:50

>> Okay.

19:51

>> It's going to help you discover more of

19:53

the bottoms up stuff.

19:54

>> Yes.

19:54

>> That we were talking.

19:56

>> Are you guys still mostly using cloud or

19:58

you switch over to GPT or both?

20:00

>> Um I can answer first. So I feel like

20:03

Fable was very exciting for me because I

20:05

feel like you can do like more meetings

20:07

meaty software engineering tasks with

20:09

it. Um, but then I've also been quite

20:11

impressed by Codeex's um, I don't know

20:14

if it's Codeex or GPT, but it's the new

20:16

Soul. I feel like Soul is quite good at

20:19

working on hard tasks and then having a

20:21

very good like human readable or

20:23

digestible output. Often when I work

20:26

with cloud models these days, like what

20:28

it's writing, like for example, let me

20:30

quickly smoke test the server before it

20:32

feel it just uses jargon that I would

20:35

never use in my day-to-day life. And

20:36

often I'm just so lost. So

20:39

>> yeah, how about how about you Hava?

20:40

>> Yeah, I've been using codeex a lot more

20:42

and it's mainly because of the codeex

20:45

app itself is so incredible in its

20:48

functionality. So there's two things

20:50

that are really important. One is mobile

20:53

support is at a completely different

20:55

level. You can see all the sessions

20:57

across all your devices on your mobile

20:59

phone.

21:00

>> Yeah.

21:00

>> You can sort of like have a command

21:01

center that like shows you everything

21:04

with no loss in functionality. like you

21:07

you can access all the same things. You

21:09

can cue messages, steer, switch models,

21:12

you can see artifacts, so on so forth.

21:14

>> The second thing that's really cool, and

21:16

not many people talk about this,

21:18

>> is you can let codecs steer codecs. So

21:22

you can have codecs create threads. You

21:24

can have it you can have threads talk to

21:26

each other. You can actually tag another

21:28

thread that's running and say, "Hey,

21:30

like help this thread out or front this

21:33

thread. This thread has like 20 tasks.

21:36

do some front run it, do some research.

21:39

You can and you can do all kinds of you

21:41

know interesting orchestration

21:43

um you know without trying to go full

21:46

code factory like you know sometimes

21:48

this is like a really helpful

21:50

>> um and you can have like codecs like

21:51

operate on codecs like hey rename all

21:53

these threads

21:54

>> do something and it just does that so

21:56

it's like really useful um and that

22:00

makes me gravitate towards that and then

22:01

like at least like the last part is the

22:03

economics they have a lot less limits in

22:06

codecs when it like the the Frontier

22:09

models is like fully within the

22:11

subscription. It's on like this 50%

22:13

thing. Then they keep resetting it

22:15

constantly.

22:16

>> Yeah. It's almost like a joke. They keep

22:17

resetting.

22:19

[laughter]

22:19

>> It's so nice. I mean, I'm not going to

22:21

complain. I saw this tweet recently that

22:24

made me laugh. Like we're in the Uber

22:26

Lift era of like AI models. Like

22:29

remember 10 years ago where Uber was

22:31

like $2?

22:32

>> We're in that right now.

22:34

>> Yeah. The subsidization is going to stop

22:36

at some point.

22:37

>> I know. Yeah.

22:38

>> Yeah. I don't know what's going on with

22:39

Anthropic. I I feel like Dario hates

22:41

Twitter or something. So like this is

22:44

completely botched all their

22:45

communication strate and and everything.

22:47

I feel like the people who actually want

22:48

to communicate like probably cannot

22:50

because of some of internal stuff. So

22:52

[laughter] I don't know what's going on.

22:53

>> What about you, Peter? What do you

22:55

>> uh Yeah, I I pretty much use Codex for

22:57

everything except for Fable that I used

22:59

to do like planning and try to find bugs

23:01

and stuff, but pretty much I use Codex

23:03

and GBT for everything.

23:04

>> Did you always use Open AI? No, I I I I

23:07

was like total cloud fanboy like even

23:09

like a few few months ago.

23:11

>> Wow.

23:11

>> But I but I discovered Codex. I I think

23:13

I think the stuff Haml mentioned, but

23:15

also like the browser use is just

23:16

incredible on CODEX. Like half the stuff

23:19

I use don't even have any APIs and I

23:21

just get to click around and [laughter]

23:23

figure things out.

23:25

>> Wow.

23:25

>> It's pretty incredible.

23:26

>> One of my most used skills is this

23:28

reverse engineering skill. There's just

23:30

like any website without a API. It just

23:34

it instructs the agent to like listen to

23:37

all the network traffic and document the

23:39

API and like like memorialize that

23:42

>> like scripts and whatnot so that I can

23:46

Yeah, I just make any website API now.

23:49

[laughter]

23:49

>> That's smart. Yeah, that's smart. Just

23:51

check out this thing called printing

23:52

press uh CLI or something. It does

23:54

something pretty similar.

23:55

>> Oh, cool. Okay,

23:57

>> so finally after Wow, it's didn't it

24:00

cooked for 15 minutes. Okay, so finally

24:02

after 15 minutes a built an interface.

24:04

Um, so let me show show you a little bit

24:06

what's going on in the interface. So

24:08

there's three different tabs. There's a

24:10

article by article pane. So this would

24:12

either be article by article or trace by

24:15

trace or whatever granularity you want

24:16

to look at. So this is looking at one

24:18

thing at a time. Um, there's a map view

24:21

which kind of shows you the clustering

24:24

that the agent did into different

24:26

semantically meaningful categories. Um

24:29

and then there's the progress view which

24:31

is kind of empty here but this view

24:34

tells us um based on our reviewing and

24:36

our annotations on the data what are the

24:39

common failure modes that are happening.

24:41

Okay. And so this is automatically

24:43

populated by an agent. So I want to talk

24:46

a little bit about the design philosophy

24:48

here which is have the human have

24:50

yourself or myself read this and give

24:53

open-ended feedback and then the job of

24:56

the agent is to really it's not to

24:58

invent new open-ended feedback but it's

25:00

to kind of group that distill that into

25:02

actionable rubric criteria so the human

25:05

doesn't have to do a lot of that work.

25:07

>> So let's go through this. Um so here I'm

25:10

looking at this AI generated article

25:12

about sleep hygiene. Um, and I'm just

25:14

going to go through a little bit about

25:17

um the how I kind of like feel. So, I'm

25:20

going to just assume that the top down

25:21

criteria Claude can take care of. Um,

25:24

and then I'll give some feedback on

25:26

this. Um, I don't really like AI

25:30

phrasing around, you know, it's not X,

25:32

it's Y, like the negative contrast. So,

25:35

I'm going to say something like I don't

25:37

like negative contrast

25:40

in my writing. And then notice that I

25:44

just gave that feedback what I'll call

25:46

in sichu. Okay. So I just I this

25:49

interface makes it really easy to just

25:52

externalize my thoughts as I have them.

25:54

And as I gave that feedback, my claude

25:57

agent is invoking the monitor tool to

25:59

basically look at whatever I give and

26:01

react to it. So here you can see that it

26:03

said my first note from you on the sleep

26:05

article. I don't like negative contrast

26:07

in my writing. And then it's going to do

26:09

something. Um let it do this in the

26:11

background, right? I just want to be in

26:12

flow of reading and giving feedback um

26:15

on the samples that it picked for me. So

26:17

remember it also picked a diversity of

26:18

samples and it will keep generating more

26:20

samples for me if it wants but that's

26:22

not important. Um anyways so I'm going

26:27

to keep reading uh

26:30

and then

26:32

>> like that kind of stuff like this one

26:33

humble me is like also really annoying.

26:37

>> Oh let's do that.

26:38

>> Yeah. H. Uh, this feels annoying.

26:42

Not exactly sure why. Can't give you a

26:47

pippy label yet, but reflect

26:53

on more feedback I gave you and reason

26:56

why.

27:01

>> So, let's do that. Um

27:07

>> so just so you know like the two things

27:08

that are really interesting going on

27:09

behind the scenes here is one you recall

27:12

the clustering that Treya showed the

27:14

diverse examples are being drawn from

27:16

the clusters so that diversity and then

27:19

the monitor tool the cloud monitor tool

27:21

that's built in the quad is being used

27:23

behind the scenes to to look at all the

27:25

stuff that he's

27:30

>> just I'm like giving a lot of feedback

27:32

back. Um, and it's totally fine that

27:34

like you know if you you might most

27:38

people like are going to have to go

27:39

through many samples to be able to give

27:41

feedback. Um,

27:42

>> yeah.

27:43

>> I've just I've done writing before so

27:45

I'm able to give it now but um that's

27:48

why it's very important to review lots

27:50

of samples especially in the beginning.

27:51

>> So just you're still doing the manual

27:53

review. You're not getting the AI to

27:54

review everything. Yes.

27:56

>> Well, here's the thing that's gonna

27:57

happen is I'm giving manual review and

28:00

in the background Claude is reflecting

28:02

on it. Um, so it's about voice rather

28:05

than facts, whatever. So once I give

28:07

around 10 reviews, then it's going to

28:09

start like thinking about a taxonomy and

28:11

then it's going to apply my feedback. So

28:14

for example, this was another negative

28:15

contrast thing. Um, after a few more

28:19

reviews, it would take all it would try

28:21

to look for all instances of negative

28:23

contrast. should automatically label it

28:25

for me. So I don't again

28:27

>> keep keep going down. I'll let it see

28:29

that. Yeah.

28:30

>> Back here. Okay.

28:34

Um All right. So that's fine. I've given

28:37

four feedbacks. Uh and then I'm just

28:40

going to also look at the right. See, so

28:42

if you look at the progress like it

28:44

doesn't like negative contra, it knows

28:46

that negative contrast is a thing. And

28:48

then this uh triadic parallel

28:50

enumeration thing. I don't know about

28:51

you, but I really don't like this list

28:53

of threes. I feel like AI does it all

28:54

the time. So,

28:55

>> yeah.

28:56

>> All right.

28:59

So, probably running something. I don't

29:02

know exactly, but let's just keep going.

29:14

>> It's like a treasure hunt for a sloth.

29:16

>> I know. Everywhere.

29:18

>> For some reason, this also

29:21

I feel like I see AI do this a lot. Like

29:24

a number, it makes like two or three

29:26

three more things to know or like keep

29:28

in mind.

29:29

>> Yeah.

29:33

>> Paragraphs

29:36

start topic

29:39

sentence

29:40

with small things.

29:48

And and uh this whole annotation thing

29:50

is uh part of the skill that you built.

29:53

>> Yeah.

29:53

>> Or

29:54

>> everything I'm showing you is like the

29:55

skill and it's this the skill was

29:58

invoked to do this.

30:01

>> Um

30:03

here's I want complete sentences.

30:09

I'm just giving giving feedback so that

30:11

way it can start doing the automatic

30:13

pass

30:14

>> because this suggestions number is right

30:16

now zero but this will become more um as

30:19

it's doing its automatic pass.

30:21

>> Okay. I I was going to ask you what's

30:23

the advantage of having it reflect live

30:25

versus like you know you give all the

30:26

feedback and it reflects at the end.

30:28

>> Oh, it's a good question. Um

30:33

I think that you can have it reflect at

30:35

the end. That works. Um, but I I also

30:39

feel like trying to interle like human

30:41

think time with AI think time is great.

30:44

Like otherwise it's going to take a lot

30:46

of time to reflect on the feedback at

30:48

the end and you're waiting on it. Like I

30:49

like the idea of like it reflecting in

30:51

the background and then being able to

30:53

suggest things for me once I hit to a

30:55

certain number. Um, and and in practice

30:57

I don't like always look at the AI side

30:59

by side like this. I'm just showing it

31:01

to you for the demo so that you can like

31:03

see that the AI is doing something

31:04

>> and giving these doing these annotations

31:06

is usually really it iterative meaning

31:08

you refine your taste over time after

31:11

you see more examples you get new ideas

31:14

on what you want and a lot of times it

31:16

could be helpful to see what the AI

31:18

thinks as well and just to help ground

31:20

you so it's like happening more

31:22

iteratively even if

31:24

>> that makes sense that makes sense. It's

31:26

like having a thinking partner.

31:28

>> Yeah. Okay. Well, I maybe I need to go

31:31

three more before it starts suggesting

31:33

cuz I can't remember the exact number.

31:36

[snorts] Start suggesting things

31:38

already. I'm just going to tell it.

31:48

You might be thinking, hey, like is

31:50

there a SAS that does this? But look at

31:52

like you don't need that. And look at

31:54

the degree of flexibility Shrey is able

31:56

to have by just integrating this process

31:59

directly with claw. She can do whatever

32:01

she wants.

32:01

>> Yeah, that that's why I this is a

32:03

tangent, but that's why I feel like SAS

32:04

is is actually in trouble, dude. Like

32:07

[laughter]

32:07

if I can just v code this kind of stuff,

32:09

why would I pay for SAS? Like just build

32:12

my own thing.

32:12

>> Now the other side of the coin is you

32:14

have to have a really good idea of the

32:15

workflow. Like a lot of people that may

32:17

not be watching this video would not

32:19

know where to begin, right? Really good

32:21

idea of the process, then you can

32:22

probably

32:36

And the other thing I that hopefully

32:38

comes to mind is like it's like really

32:40

tedious to give feedback. Like

32:42

>> yeah,

32:43

>> I've not reviewed that many. Um

32:47

I've only given eight notes. So just

32:50

like yeah, know that it's it's really

32:51

hard. Um, and also AI just can't read

32:54

your mind for you. I've never like all

32:58

of these notes that I'm having like some

33:00

some of them maybe AI can come up with,

33:02

but I don't know like not all of them.

33:07

>> Uh, yeah, you can probably just speed up

33:08

with like voice dictation or something,

33:10

>> but but yeah, you just have to read

33:12

through everything.

33:13

>> Yeah.

33:13

>> Yeah.

33:14

>> Or [snorts] read through some sample at

33:16

least. Um, okay. So, I'll start pushing

33:19

suggestions to your queue right now

33:20

instead of waiting. Yeah, probably I had

33:22

to get to 10 notes, but

33:24

>> okay.

33:25

>> I like having the 10 notes because it's

33:27

like a carrot. It like makes you read

33:29

some things like

33:31

>> makes you sweat just a little bit which

33:33

is good for you.

33:35

>> There we go.

33:35

>> Thinking about stuff. Oh, nice. Okay.

33:38

>> Yes.

33:41

>> Okay. So, it's going to uh annotate

33:44

annotate every single doc.

33:46

>> Mhm.

33:47

>> Yep. With the criteria I've already

33:50

found.

33:52

Okay.

33:52

>> Thing that's really great about this is

33:54

like okay this shre is showing it's it's

33:58

fit perfectly to this domain of writing.

34:00

Like if you notice the writing is being

34:03

shown like in situ as she said in the

34:05

way that you would read it as a user.

34:07

You're not looking at like a log file or

34:09

a trace viewer at what it actually looks

34:11

like

34:12

>> and you're just annotating it in place.

34:16

>> Right. So it's like it's um generated

34:19

361 suggestions.

34:22

Yeah. So it said the less than four

34:23

words staccato rule produced 249.

34:28

Um

34:29

>> um but but what happens so what happens

34:31

next after you accept all this? Like

34:33

does it actually suggest how how do you

34:35

fix this stuff or

34:37

>> so the idea here as I mentioned before

34:39

is to like then you can turn this

34:41

failure modes all the things that you

34:42

found into a rubric like this becomes

34:44

your evals. I see.

34:47

>> Um, and then so you can apply that as a

34:49

skill as you do. Um, you can have an LLM

34:53

judge for each. You can turn that into a

34:54

dashboard. You can have that monitor

34:56

agent traces or outputs in real time.

34:59

Like the whole the the hardest part of

35:01

eval is error analysis which is coming

35:03

up with this rubric.

35:05

>> Um, so this process just makes it really

35:07

easy

35:07

>> because of this annotation. You also

35:09

have a data set. So

35:11

>> yeah,

35:12

>> you can after your evals get created now

35:15

you can check those evals against the

35:17

data set

35:18

>> that I mean what I'm saying is part of

35:19

eval but

35:21

>> eval is not just a rubric you want to be

35:23

able to measure the rubric and say like

35:26

is it good and so because of these been

35:28

annotated you can you can you have this

35:31

thing to benchmark it against

35:33

>> right

35:35

>> okay so it found okay so the most

35:37

frequent example is um the staccato

35:40

fragments

35:41

Right.

35:42

>> And there was also colons and stuff,

35:44

>> but it makes sense, right? It's like, so

35:47

that's another thing like as somebody I

35:49

didn't know that that's the most

35:50

frequent example, right? Like

35:52

>> I would have thought it's the not X but

35:54

Y, but I I guess

35:57

I guess it is the stacato fragments.

36:00

I believe it. So I so again goes back to

36:04

the philosophy of like human is the one

36:06

driving the taste and the AI is the one

36:09

like scaling up or giving superpowers to

36:11

the human to apply that at scale.

36:14

>> Okay. Got it. Wait. So so so just to

36:16

understand like can you go back to the

36:17

progress tab?

36:19

>> So now we have these these themes and

36:21

these rubrics, right? So so what what do

36:23

I do with this now? Let's I say I'm

36:25

building a writing skill. Do I just like

36:26

uh point AI to this thing and be like

36:28

hey now create a scale for me? I mean, I

36:31

would be right here. Write a skill that

36:34

evaluates AI assisted

36:38

writing based on the failure modes

36:42

found. I like it would just ask for a

36:44

skill and there's the skill. Um, and I

36:47

want to emphasize that like you should

36:48

spend a couple more hours doing the

36:50

error discovery. Like don't don't just

36:52

look at five things. like actually go

36:53

through, you know, review all 18 like

36:57

feel confident that you've like kind of

36:58

saturated. Um, but yeah, you can just

37:02

ask it to write a skill and it will

37:04

write a rubric.

37:06

>> I I see. Okay. Then then basically every

37:09

time you have a writing sample, you just

37:10

run the scale, right?

37:12

>> Yeah. And the rubric will output like uh

37:14

this writing sample has a ton of

37:16

staccato fragments or something.

37:18

>> Um the rubric will have staccato

37:20

fragments as one of the criteria. Like

37:22

if you look at if you think about your

37:24

podcast evaluation skill, right? That's

37:26

a rubric of many criteria. So stat will

37:29

be one of the like five criteria that it

37:31

comes up with or that this came up with.

37:33

>> So one thing that's really interesting

37:35

is if you notice the way there is a okay

37:38

the way you eval something is often

37:41

should inform the way you might want to

37:43

design the interface for the user to

37:46

begin with. So, um, if you notice like

37:48

in the annotation workflow, she had the

37:50

AI, you know, suggest

37:54

different annotations. You can imagine

37:56

like if you were writing for real, you

37:58

would want an IDE that would make the

38:00

same suggestions to you and show you

38:02

why. Like, hey, like here's a staccato

38:04

thing, here's a negative contrast,

38:06

here's this, here's that.

38:07

>> Um, so that you you don't have to just

38:09

have AI just overwrite things. You have

38:11

to keep rereading things. You want to

38:12

accept and reject things.

38:15

And so yeah, it's actually forces some

38:18

amount of good design thinking to go

38:21

through this exercise because each

38:22

interface is going to be different like

38:24

Sha showed you one for writing. Okay, a

38:26

customer service support bot is going to

38:28

look totally different in the interface

38:31

or you know another

38:32

>> even in the error discovery interface,

38:34

right? My skill asks the agent to

38:37

reflect on the nature of the data and

38:39

design a visual encoding for it. It's

38:40

not going to look exactly like this.

38:43

>> Okay. And I guess the the main the main

38:46

benefit so uh in the previous process I

38:50

still have to review traces and come

38:51

over the errors but now the AI is

38:54

actively helping me find the themes

38:57

right so it kind of skips that

38:59

>> host step.

39:01

>> Yep.

39:01

>> Yeah. Got it. [snorts]

39:03

>> Well doesn't skip the step I would say.

39:06

You still have to look at the data.

39:07

There's there is no world in the future.

39:09

Even if you have AGI, if you're building

39:11

a product, you have to look at your

39:13

data. Otherwise, you're there's no way,

39:15

right? Like you cannot you have to be

39:17

able to inject your taste somehow into

39:19

the development of your product.

39:20

>> This is we would think is our is our

39:23

attempt to show you like how here's a

39:25

way you can do it effectively with with

39:27

all the benefit and might of AI to help

39:29

you accelerate.

39:30

>> Got it. What what what is the point of

39:32

the suggestions to re review? it just

39:34

gives more like gives more human input

39:36

to some of the stuff.

39:37

>> Um, so this is the agent scanning all

39:40

records for the known failure modes.

39:42

>> Um, and these are records that were not

39:45

labeled by me with the human. So, okay,

39:49

just to like I can accept all in

39:51

practice I always accept all of them

39:52

after like scanning briefly just to make

39:54

sure like I use this to make sure that

39:57

the AI understands what I meant by for

40:00

example negative contrast.

40:02

>> Okay. So, it's like, yeah, totally. It

40:05

understands what I meant by negative

40:06

contrast. Good job. Accept.

40:09

>> The thing I really love about this is if

40:11

there's something wrong with this

40:12

interface or she can't give the feedback

40:14

she wants. So, she's like, let's say she

40:16

has a theme that the AI is missing out

40:18

on, she can just go directly to Claude

40:20

on the left hand side and fix it. You

40:23

don't have you're not constrained to

40:24

this UI.

40:26

>> Yeah. Yeah.

40:26

>> Exactly.

40:27

>> Yeah. This is why like uh your own

40:29

product is like way more flexible than

40:30

some sort of SAS SAS product. It just

40:33

you have all these use cases that no

40:35

common SAS will think think about. Yeah.

40:38

Um so if I wanted to use you okay so now

40:40

this is getting me pretty excited. So

40:42

let's say I want to use my use your

40:44

error analysis your error discovery

40:45

skill for my previous example right for

40:48

the podcast stuff. So basically I would

40:50

uh give it so you got AI to generate a

40:53

bunch of writing samples. So, instead of

40:54

that, maybe I'll give it um like maybe

40:57

before and after examples or like maybe

40:59

I'll give it examples of like takeaways

41:01

that it made for previous episodes.

41:04

>> Yeah. What I would do is just create a

41:06

folder with your current skill for

41:09

generating takeaways as is and then

41:11

examples of podcast transcripts and the

41:14

takeaways that you said were good. And

41:17

then I would just invoke the skill and

41:20

say um help me do error analysis on my

41:24

podcast takeaway skill. So apply the

41:26

podcast takeaway skill to generate a

41:28

bunch of takeaways and then build me an

41:31

interface to review those takeaways

41:33

especially compared against the outputs

41:36

>> the the that's perfect. Yeah, that's

41:39

that's perfect. So yeah.

41:41

>> Yeah. and then just go through I mean

41:42

it'll take like 15 minutes for it to

41:44

give you the interface but you'll be

41:46

able to go through it fast and like the

41:48

whole point is confidence right like at

41:49

the end of the day I want confidence

41:51

that that my skill is do like I know

41:53

what behavior the LLM is exhibiting so

41:56

that's the whole point of error analysis

41:58

>> so so I should have like maybe I should

42:00

modify your interface to have my the

42:02

AI's output on one side and the ground

42:04

truth on the other so I can look

42:06

>> so um the interface is going to be

42:08

different when you invoke the skill on

42:10

your data. It's not going to come up

42:13

with this interface. The AI agent will

42:16

read your data. It will read those

42:18

takeaways and come up with a different

42:20

interface to show you and render that to

42:22

you as in step two here, designing the

42:24

visual encoding and building the HTML.

42:26

>> Okay, that that's amazing. Thanks so

42:28

much. Thanks so much for showing us.

42:30

Okay, and I just I just want to point

42:31

out to people who watching this that

42:32

this is free as far as I know. And uh

42:35

>> yep. Yep.

42:35

>> Just go to error discovery skill on

42:38

GitHub and and you'll find or you know

42:39

what I'll link in the description

42:41

>> of this episode. Cool. And I love how

42:43

you're honest about plot helping you

42:45

come up with this on the right.

42:47

>> Oh yeah, of course.

42:49

>> Yeah. I usually try to remove that to

42:51

pretend I built the whole myself.

42:54

>> No, no, no.

42:55

>> Yeah. Yeah.

42:56

>> Claude is very good at like writing

42:57

skills for Claude. Like I don't know how

42:59

to talk to AI. So that's that's great.

43:02

Um, but I mean like I wrote the readme,

43:04

right? Like I I I'm the one coming up

43:07

with this process, so I feel pretty

43:10

confident in it.

43:11

>> Cool. All right. Well, this is awesome

43:13

demo. I I think we have one more topic,

43:14

which is uh Haml, you you wrote a

43:17

article about do AI automations actually

43:19

work or like where do they work, right?

43:22

Do you want talk about that?

43:23

>> Yeah, I can quickly talk about that.

43:25

Okay. Um, so I have a blog post, do

43:27

automated evals work? And y'all, you we

43:30

can put the link in the description, but

43:32

basically what we did is,

43:36

let me just back up. Um, so I wrote a

43:39

blog post about that's titled, do

43:42

automated evals work? And the reason for

43:45

writing this blog post is there is a

43:48

bunch of different vendors that in the

43:51

last few months have created tools that

43:54

will that promise to just do evals for

43:57

you. And the way they work is you upload

44:00

your traces into those tools and you

44:03

chat with an AI. So here's an example.

44:06

Um here's one in Brain Trust where you

44:08

know is a trace viewer on the left. Um,

44:11

and then you can chat with your AI, you

44:14

can ask it some questions and basically

44:16

say do my evals for me. And this

44:18

particular product, this feature is

44:20

called loop. Um, and

44:23

uh, Arise has really the same thing, a

44:26

very similar thing. It's called Alex.

44:28

You can chat with it. You can say, "Hey,

44:29

look through my traces, do my evals."

44:31

Langmith has something similar. And um

44:37

you know we were curious like okay if we

44:39

take a data set that we have annotated

44:41

as humans how good is the AI at catching

44:44

those same errors and how good are these

44:47

automated eval systems and it turns out

44:49

that um and here's a fun meme. It's like

44:52

okay you're faced with this choice right

44:53

as a developer or as a product manager

44:56

like do you just press this auto eval

44:57

button sounds really convenient or do

44:59

you actually look at the traces? It's

45:01

actually a pretty hard decision because

45:03

like it's very tempting to press the

45:04

button and um and so I won't get too

45:08

much into the details but um we did kind

45:12

of a benchmark

45:13

and we found that okay like the tools

45:16

are able to recover a lot of the errors

45:18

that a human would.

45:21

But

45:23

u on the flip side it these automated

45:26

tools all sort of miss the same thing

45:28

which is

45:30

uh finding errors that are not obvious

45:33

things that require product judgment and

45:36

kind of taste. So like anywhere in the

45:38

trace where something obviously went

45:40

wrong like a tool call failed or

45:43

somebody's like expressing

45:45

dissatisfaction or something is just

45:47

like really obvious like you don't need

45:49

product expertise to know something went

45:50

wrong. AI is really good at catching

45:52

that. and it's not good at catching like

45:54

oh okay this uh sa this support bot

45:57

didn't handle sales objections correctly

46:01

you know so like what we uploaded is a

46:03

customer service uh it's it was a real

46:06

estate or rental sales agent that is you

46:11

know supposed to interface with people

46:13

looking to buy apartments or rent

46:14

apartments and like yeah for example one

46:17

of the items that were always missed

46:19

were sales objections like hey we don't

46:22

it would just say hey we don't have that

46:23

available have a nice day instead of

46:25

saying hey like there's alternatives for

46:27

you blah blah but another kind of very

46:31

interesting fact is we also benchmarked

46:33

this against coding agents so clawude

46:35

codeex so on and so forth and we found

46:38

that those pretty much work the same so

46:42

at the end of the day what you know

46:43

what's going on here is the harness is

46:47

somewhat thin it's just someone else's

46:49

prompt some you know so when you're

46:51

chatting with uh the AI these tools,

46:54

you're really just invoking one of those

46:56

models and they have a prompt behind the

46:58

scenes that's trying to do the same

46:59

thing. But what you really want to do is

47:03

you want to steer the agent more

47:04

actively. So the way to make this better

47:08

across the board for either you know

47:10

your coding agent or the eval tools is

47:12

to give it an idea of what's good and

47:14

bad. The only thing is you don't know

47:16

what's good and bad unless you do what

47:18

Shrea just did. It's really hard for

47:21

humans to upfront think of like from top

47:24

down like what are all the things that

47:25

are good and bad. It's like almost

47:26

impossible.

47:28

>> Um and so that's really what we found

47:31

here is like okay they're all missing

47:32

the same thing. So but it's useful. It's

47:35

not that it's not useful. It can catch a

47:37

lot of things. It just won't catch some

47:40

of the important things

47:43

especially those that are likely to make

47:45

your product work well. And so you

47:48

should think like okay all your

47:50

competitors are going to be using this.

47:51

This is like a baseline of like you can

47:53

of course you can point claude at your

47:55

product and say like find all the errors

47:58

and your competitor is going to do that

47:59

too. So the thing that's going to matter

48:02

at the end is

48:03

>> how much taste can you infuse into your

48:06

product beyond that uh to to like define

48:10

like what is good. And so that's why you

48:13

need to look at your your um you need to

48:16

look at your data and I spelled it out

48:18

here um you know I spell out like okay

48:21

these are the kinds of things missed for

48:23

example you know sales objections in

48:26

this example this example it was

48:28

multi-channel so customers of this

48:31

particular agent can talk on the web and

48:33

text make a phone call they the eval

48:38

didn't know that so the eval agent

48:40

didn't catch that like hey like we

48:42

shouldn't be putting markdown in t text

48:44

text messages it doesn't work.

48:45

>> Um yeah so things like that like of

48:47

course you can give this context to the

48:49

LLM but this is not something you might

48:51

think about giving as criteria up front.

48:53

You would have to see that failure.

48:55

>> Um so so yeah that's that's the gist of

48:58

it. Um we also uh have this recording.

49:02

Well you should actually cut that off.

49:04

So um that's the just blog post. Um

49:07

>> yeah.

49:08

>> Okay.

49:09

>> Yeah, just sharing that in case that's

49:11

useful.

49:11

>> Okay. So the takeaway is like this stuff

49:13

will get you to a good baseline, but if

49:15

you give a demo of your product, you

49:16

just do the manual reviews. That's the

49:19

take away, right? The manual

49:20

>> review. Use use the automated tools and

49:23

also look at data yourself. Like it's a

49:25

competition, right? You you got to be

49:27

better than your competitors.

49:29

>> Yeah. So one takeaway is okay the coding

49:32

agents are just as you know are at par

49:36

with these auto eval tools. The only

49:38

benefit of auto eval tools is they're

49:40

integrated into the the rest of the

49:43

stack. Like

49:44

>> yeah they're in the ecosystem. So if you

49:46

store your dasis traces in in Langmith

49:48

you should totally use lang like it's a

49:51

no-brainer.

49:52

>> I see. I think it was pleasantly

49:54

surprising for me because I wasn't

49:56

involved in writing this blog post,

49:57

right? It was Hamill. Um I I was so s I

50:01

was very happy to find that oh these

50:03

these off-the-shelf tools work so well.

50:05

So it's great like you should absolutely

50:07

use them and and find all the errors in

50:09

your data but also you know look at your

50:12

data for more errors and also look at

50:14

the errors that for example linksmith

50:16

found just to make sure you agree with

50:18

them right because the precision on

50:19

these is is like 80% to 90% in the best

50:22

case. Um so that means 10 to 20% of the

50:26

errors found by the automated eval tool

50:28

are actually not errors. they're like

50:31

red herrings and might distort your

50:33

product if you just listen to that

50:34

automatically. Right? So, it's very

50:36

important to look at both recall and

50:38

precision.

50:39

>> Okay. So, basically actually looking at

50:41

stuff is is the edge. [laughter]

50:43

Actually reading and looking at stuff

50:45

>> always.

50:47

>> Yeah. Cuz because like you know with all

50:49

this AI generated stuff you can just

50:50

like let let it go. But like actually

50:52

reading stuff this is what I learned

50:53

from another guest that I had to just

50:55

actually read what what it produces.

50:58

>> Oh yeah. Yeah. And even like sometimes

51:01

with code, that's an edge.

51:03

>> Yeah.

51:03

>> You need selectively code, it's

51:05

definitely an edge.

51:06

>> Cool. All right. So, why don't we kind

51:08

of summarize what we covered in this

51:10

episode? I think we covered a lot. Um,

51:12

so I I think my number one takeaway is

51:14

to actually use SH's scale because uh I

51:18

think it's pretty awesome and um give a

51:20

bunch of examples and then have it help

51:22

you identify like the common themes of

51:24

errors, right? That's kind of number

51:26

one. And number two from from your

51:27

example is like use all the coding

51:30

harnesses and like the eval stuff, but

51:32

also just read the stuff. That's

51:34

[laughter]

51:34

that's kind of my two takeaways. So uh

51:36

thank you for sharing all this

51:37

information for free. Can you guys talk

51:39

about what is covered in your eval

51:41

course that's that we didn't cover here

51:43

or like you know what was the more

51:45

advanced stuff?

51:46

>> Yeah, so here hopefully you got a taste

51:47

of our philosophy of like how you're

51:49

really going to differentiate your AI

51:50

generated product. Um, so there's some

51:53

our our course has five modules. It's

51:55

totally revamped for this era of like

51:58

agentic AI. So beyond just teaching you

52:00

the error analysis philosophy, we also

52:02

teach things like okay, how do you

52:04

really productionize this work and scale

52:06

it up to work within your company? For

52:08

example, like if you have multiple

52:09

people working on the same product, I

52:11

need to align on criteria. How do you

52:13

integrate this stuff into, you know,

52:15

CI/CD for example? um how do you monitor

52:19

over time like drift over many months of

52:22

your deployment. We also have a module

52:24

on safety and kind of adversarial

52:26

evaluation. So it's really important you

52:28

know how do you make sure you're not

52:29

leaking tenant data? How do you make

52:31

sure that you can evaluate that over

52:33

time especially when you want to deploy

52:35

to your actual customers. Um and then

52:38

finally the last module is all about

52:40

improvement. So it's cost improvement

52:42

and accuracy improvement. really those

52:45

both of those go hand in hand, right? If

52:46

you switch to a cheaper model, how do

52:48

you improve the accuracy of your cheaper

52:49

model? So, it's both cost and accuracy

52:52

improvement. Um, and really no one else

52:54

talks about this, right? If the the

52:56

benefit of having evals is that you can

52:57

then put put it in an auto research kind

52:59

of loop. Um, what are the kinds of

53:01

strategies to invoke? How do you make

53:03

sure that you know you we found that in

53:05

in our um consulting cases or even for

53:08

some of our students we're able to cut

53:09

cost by 100x truly compared to um like

53:14

using Opus or probably Fable will be

53:17

even more um so we're really excited to

53:19

kind of teach people the techniques to

53:21

do that.

53:22

>> Awesome. And uh yeah, I'll I'll put the

53:23

link to the course in description. And I

53:25

I think uh from from now on I'm gonna

53:28

interview Straas and Haml. Never Ham by

53:30

himself because I learned a lot more.

53:32

[laughter]

53:34

>> Yeah. Even just

53:37

[laughter]

53:39

>> yeah. Yeah.

53:39

>> I'm a full-time educator now. So

53:42

>> yeah. Yeah. That that that's amazing.

53:44

Thanks for sharing all this knowledge

53:45

and um yeah, I'm gonna go off and use

53:47

your skill right now. So we'll see how

53:49

it goes.

53:50

>> Cool. Thanks for having us, Peter.

Interactive Summary

This video features AI eval experts Hambo and Shrea discussing practical workflows for evaluating AI outputs, specifically for writing tasks. They emphasize that while automated eval tools are useful for establishing a baseline and catching obvious issues, true quality relies on the 'final boss' of LLMs: writing in a way that aligns with human taste. They introduce an error-discovery interface that allows users to perform manual reviews, which AI agents then distill into an actionable rubric, enabling more sophisticated and personalized evaluation.

Suggested questions

3 ready-made prompts