HomeVideos

/handoff is my new favourite skill

Now Playing

/handoff is my new favourite skill

Transcript

361 segments

0:00

A few weeks ago, I noticed myself doing

0:02

something with agents that I thought was

0:05

very clever, but I thought it was just

0:07

too simple to require a skill. For those

0:10

who don't know, I'm constantly thinking

0:12

about skills. I'm constantly thinking

0:13

about how to package my instincts and

0:16

coding practices into reusable skills,

0:18

and this has meant my skills repo has

0:20

almost 100,000 stars at the time of

0:22

recording. The skill that I started to

0:24

think about was a handoff [snorts]

0:27

skill. And the theory was that this

0:28

skill would take the context window of

0:31

the current session and compress it down

0:33

into a markdown file that could be

0:35

handed off to another session. And so, a

0:37

couple of weeks ago, I shipped this.

0:38

It's inside skills, inside productivity,

0:41

and it's inside handoff here. And it's a

0:43

very, very simple skill. It says to

0:46

write a handoff document summarizing the

0:48

current conversation so a fresh agent

0:50

can continue the work. Save it to the

0:52

temporary directory of the user's

0:53

operating system, not the current

0:55

workspace. I put this into my skills

0:57

folder as an experiment to see how much

0:59

I would use it. And it turns out I used

1:01

it a lot. In this video, I'm going to

1:03

show you a deep dive of the skill, kind

1:05

of why I designed it, what is the point

1:07

of it, how it compares to built-in tools

1:09

in some of these harnesses like compact,

1:11

and also how you can get the most out of

1:14

it to make the most of your grilling

1:16

sessions. And if you dig the kind of

1:17

stuff I've been showing you, then you

1:19

will love the course that I've put

1:21

together, which is AI coding for real

1:23

engineers. A two-week cohort for folks

1:25

who want to use AI coding tools for

1:28

shipping quality code, not slop. It

1:30

starts on June the 1st. We're doing a

1:32

discount right now. Get into the link

1:34

below so you can check it out. Let's

1:36

start first of all by explaining why I

1:38

made this skill and how it differs from

1:40

compaction, which you may have heard of

1:42

before. When we're inside a session like

1:44

this, a coding session, we essentially,

1:46

as we, you know, converse with the

1:49

agent, as it does tool calls, as it

1:50

makes file edits, then this context

1:52

window is going to be filled up and

1:54

filled up with more and more stuff in

1:57

it. More and more tokens will fill up

1:58

the context window. Now, in the harness

1:59

I use, Claude code, it's the context

2:02

window is huge, right? You get 1 million

2:04

tokens worth of context window, but

2:07

there is actually a smart zone and a

2:10

dumb zone in these context windows.

2:12

Early on in the context window, you are

2:14

going to get much better performance

2:15

from the agent because the attention

2:18

relationships are not so strained there.

2:21

Because there's much fewer tokens to

2:23

calculate, fewer attention relationships

2:25

between those tokens, then the agent's

2:28

attention isn't so diffuse. In other

2:30

words, it's better able to focus when

2:32

there's less content in there. This

2:34

means that as your conversation

2:36

develops, you're going to get dumber and

2:38

dumber and dumber responses from the

2:40

agent all the way up to going up to, you

2:42

know, 800,000 tokens, which personally

2:45

I've never been in because around by the

2:47

120k token mark, I start to feel like

2:51

I'm in the dumb zone. So, this means

2:52

yes, that even though Anthropic

2:54

advertises a ton of context window on

2:56

these models,

2:57

really for, you know, proper smart

3:00

tasks, you've only got about 120k to

3:02

work with, which means you need to

3:04

budget really efficiently and you need

3:06

to be aware of your context window at

3:07

all times. So, the question then

3:09

becomes, what do you do when you're

3:10

starting to hit up against this dumb

3:13

zone? How do you recover your

3:15

conversation? How do you continue the

3:17

conversation beyond the dumb zone while

3:19

staying smart? And the answer to that is

3:21

compact. What compact does is it will

3:23

take a large conversation like this and

3:26

summarize it, so you go essentially from

3:28

near to the dumb zone to

3:31

all the way into the smart zone here.

3:34

And there's even sometimes an auto

3:36

compact buffer depending on what harness

3:38

you're using and whether you've got it

3:39

turned on, which means that when you're

3:41

near to the end of the context window,

3:42

let's say deep in the dumb zone, the

3:44

auto compact buffer will kick in and

3:46

automatically summarize your

3:48

conversation inside a new session. This

3:51

summary usually looks like the files

3:54

reference, so it's just a list of files

3:56

that have been referenced, the things

3:57

that you said in the conversation are

3:59

usually included, and the general tone

4:01

of the conversation as well. This is

4:03

then included as a little nugget at the

4:05

start of the new session, and as you

4:07

build up context in the new session,

4:09

then you're continually referencing the

4:11

old session. This means as you continue

4:13

to compact and compact, you're going to

4:15

end up with this kind of sediment of

4:16

different layers here from previous

4:18

conversations. And this can be a little

4:21

bit inefficient, but it's also a decent

4:23

way if you want to do certain types of

4:25

sessions where you just need to barrel

4:28

on on the same problem again and again

4:30

and again. It can be really useful for

4:31

debugging, actually, because you can

4:34

compact all of the other options that

4:36

you've tried, and then continue to try

4:38

different things, hit the barrier, and

4:40

then compact again to just save your

4:43

state, essentially. So, it's a way of

4:44

doing a long-running session, but it's

4:47

only really one session. So, I continue

4:50

to find compact a really, really useful

4:52

tool for creating these long single

4:54

sessions. But what I started to notice

4:56

was I wanted to do other things with

4:58

compact. I wanted to compact into

5:01

another session. For instance, let's say

5:03

I was in one session here, and while I

5:06

was in this session, I noticed a little

5:08

refactoring opportunity. Something that

5:10

was totally out of bounds, out of scope

5:12

for my current session, but I knew I

5:14

would need to get there eventually. So,

5:15

what were my choices? I could extend my

5:18

current session, but then I would end up

5:20

with this sort of like diluted context,

5:22

where I was half working on one thing,

5:24

half working on the other, and I would

5:26

definitely hit the dumb zone, right? So,

5:28

I probably wouldn't be able to finish my

5:30

initial goal. I could compact, but then

5:33

I would clobber all of the progress that

5:34

I'd made in my current session, right?

5:37

What I really wanted to do was just say,

5:39

"Okay, I want to complete this other

5:41

thing in a separate session, and keep my

5:43

current session pure." In other words,

5:45

this was what I wanted. I wanted to

5:47

essentially take the context or take

5:50

just the slice that pertains to this

5:52

extra bug fix, hand it off to another

5:54

session, and then these two could just

5:56

run independently. And so, for a while,

5:58

what I was doing was saying, "Okay, take

6:00

the stuff in my current session. I want

6:02

to fix this particular bug. Write me a

6:04

handoff.md document so that I can then

6:07

just pass that into another agent." And

6:09

it turned out I was doing this so

6:11

freaking often that I just decided,

6:13

"Okay, I need a skill for this." I most

6:15

often use handoff while I'm grilling

6:17

here. Here, I'm inside a grilling

6:19

session that I did for planning some

6:20

future features for Sandcastle, which is

6:22

my sort of software factory. And what

6:25

you can see here is that I'm kind of

6:26

answering some questions. I'm only in Q2

6:29

of this grilling session, so not a long

6:30

one. And I say here, "I think in future

6:33

we may want to move the iterations and

6:34

the completion signal onto a separate

6:36

API. In fact, let's hand off that task

6:39

to a separate agent." You can see here

6:41

that when I'm defining handoff, when I'm

6:44

saying, I'm saying the reason why I'm

6:47

handing off and exactly what should be

6:49

in that document. This does two things.

6:51

First of all, it actually sharpens the

6:53

current grilling session I'm on. So, it

6:55

says that given that constraint, Q2

6:57

collapses. So, it doesn't actually like

6:59

it helps my current grilling session

7:01

because I'm saying that's out of scope,

7:02

we'll pick that up somewhere else. It

7:04

then goes and creates a markdown file

7:07

just here with the focus for the next

7:09

session, file a GitHub issue, and

7:11

eventually design for splitting

7:12

iterations and the completion signal

7:13

into a separate API. And then later, I

7:16

just pass this into a another agent in

7:19

order to create the issue. Simple.

7:20

Another pattern that I really strongly

7:23

recommend is handing off during a

7:25

grilling session to prototype. When

7:27

you're grilling, when the agent is

7:29

asking you questions from a grill me or

7:31

grill with docs, which are more of my

7:33

skills, you will often find there's two

7:35

categories of questions you need to

7:36

answer. There are the kind of known

7:38

unknowns, the ones that the agent can

7:41

ask you about, and then there's stuff

7:42

that you really need to see in code or

7:45

need to see prototyped. This can be

7:47

really true with like UI prototypes or

7:50

complicated bits of logic that you're

7:51

not quite sure how to deal with yet. So,

7:53

in this grilling session, we're down to

7:55

question 13, actually, and we've got a

7:57

sort of final uh resolution from the

8:00

agent. And then we can see I say, "Hand

8:03

off to prototype the difficult bits

8:04

here, the window communication, the teal

8:06

draw SDK integration," which was

8:08

something I was building at the time. It

8:10

creates the hand off, and then I go and

8:12

implement the prototype on that branch.

8:14

So, in the prototype session, this ended

8:15

up being a huge session, so 169 K

8:18

tokens, so way bigger than would have

8:21

fit inside the grilling. And what I did

8:23

was I created this prototype of the UI

8:26

and the kind of interaction that I

8:27

wanted to see. And then I said, "Okay,

8:30

let's hand this off back to the grilling

8:32

session that spawned this. Take all of

8:34

the learnings from the prototype,

8:35

anything that's not directly captured in

8:37

the prototype itself or that's

8:38

non-obvious, give me a handoff document

8:40

that I can pass back to the planner."

8:43

This is actually a really common pattern

8:45

that I'm using here, where you have the

8:47

initial session where you do some work,

8:49

you hand off to another session, that

8:51

session then creates another handoff

8:53

document, and then passes it back to the

8:55

original session. It's almost like

8:57

you've done a kind of DIY sub-agent,

8:59

where you're able to use a context

9:01

window for one specific task, compress

9:04

your learnings from that task, and pass

9:06

it back to the parent. Then I was able

9:08

to finish the grilling session and

9:10

create some proper PRDs and issues with

9:13

the prototype in there. So, it's

9:15

incredibly rich pattern for actually

9:18

getting what you need out of AFK agents

9:21

and using prototypes. It's very, very

9:23

cool. It's worth saying, too, that the

9:25

thing that's cool about just using like

9:26

a markdown documents here and not

9:28

relying on kind of native agent stuff is

9:31

that you can have this first session be

9:33

Claude code, but you can just pass this

9:35

to another agent, right? You can pass it

9:37

to Codex or pass it to, you know,

9:39

Copilot CLI, whatever you're using. So,

9:42

if you want to do any kind of

9:43

adversarial review or any kind of, um,

9:46

you know, interaction between different

9:48

coding agents, this is a very, very

9:50

simple way to do it. We should also just

9:51

read through the final bits of the skill

9:53

here just so you understand the

9:54

reasoning behind everything.

9:56

The theory here is include a suggested

9:58

skill section in the document which

10:00

suggests skills that the agent should

10:01

invoke. I added this because sometimes

10:05

it would

10:06

I use skills to kind of define the

10:08

flavor of that session and so having a

10:11

suggested skill section means that you

10:13

can kind of just paste the handoff

10:15

document into the new session. It will

10:17

invoke the skills needed like grill with

10:19

docs or diagnose or prototype or

10:21

something and then you're kind of good

10:23

to go. So, you don't need to think about

10:25

the skills that you need to use in the

10:26

next session. It's pretty handy. Another

10:27

one is do not duplicate content already

10:29

captured in other artifacts. I would

10:32

often find these handoff documents just

10:33

got really big and they were just

10:36

duplicating stuff that was already

10:37

present either in other markdown files

10:40

or in resources like GitHub issues or

10:42

things like that. So, it's basically

10:44

saying just use pointers instead of, um,

10:46

you know, repeating everything that's in

10:48

the documents. I also really strongly

10:49

believe that you should save these

10:51

handoff files to the temporary directory

10:53

of the user's OS. In other words, these

10:55

handoff files are disposable. They are

10:57

not something to be kept around for a

10:59

long time to rot in your code base's

11:01

documentation. Another one is redact any

11:04

sensitive information, API keys,

11:05

passwords, or PII. This is, you know,

11:08

pretty essential. You don't want these

11:10

floating around in markdown files in

11:11

just random places. And finally, if the

11:14

user passed arguments, in other words,

11:15

what the next session will be used for,

11:17

treat those as a description as to what

11:18

the next session will focus on and

11:20

tailor the doc accordingly. I think of

11:21

this is essential for handoff because in

11:24

order to write a decent document, the

11:27

agent needs to know what the next agent

11:29

session is going to focus on. Every time

11:31

I used handoff, I always describe the

11:33

purpose, the reason that we're handing

11:36

off because I just can't see how you

11:38

would write a good handoff document

11:40

otherwise. And of course, dictation

11:42

makes this really easy cuz I just blast

11:43

it out and then we're good to go. So,

11:45

there we go. That's handoff. This is an

11:47

essential skill in my toolkit that, you

11:49

know, just like a lot of my other skills

11:50

didn't exist but a few weeks ago. If

11:52

you've been enjoying my skills, then you

11:53

should check out the Cohort course. It

11:55

is an absolute banger. We had about

11:57

2,500 people take it last time and I'm

12:00

expecting, you know, a decent whack this

12:01

time, too. Other than that, thank you so

12:03

much for watching. My bookshelf behind

12:05

me is filling up with new coding books

12:07

that I'm going to be reading over the

12:09

next couple of weeks. I'm thinking about

12:11

maybe making a sort of what's on my

12:13

bookshelf video of recommended books.

12:15

And I don't know. If you like that, then

12:17

maybe give us a like and a comment or

12:18

let me know what you want to see next.

12:20

Either way, thanks for watching and I'll

12:22

see you very soon.

Interactive Summary

The video introduces a new 'handoff' skill for AI coding agents, designed to bridge separate sessions by compressing the context window into a portable markdown file. The speaker explains how this tool helps maintain focus by avoiding the 'dumb zone' of large context windows, allows for parallel work on different tasks, and enables a pattern of delegating sub-tasks to separate sessions before returning results to the original planner.

Suggested questions

4 ready-made prompts