HomeVideos

Burn through the backlog from hell with /triage

Now Playing

Burn through the backlog from hell with /triage

Transcript

302 segments

0:00

The problem with most skill based setups

0:02

is that they are great for solo

0:03

developers, but they're not so great

0:05

when you get into teams. With teams, or

0:08

when you're working on other people's

0:10

stuff, or building ideas for other

0:11

people, you will often need to triage

0:14

other people's ideas to figure out if

0:16

they're good or not, to even figure out

0:19

if they're worth building, to figure out

0:21

if it's a bug report, to see if you need

0:22

to reproduce it, all that stuff. And

0:24

this is very familiar to anyone who's

0:26

been building anything of reasonable

0:27

size for a while. And so, I have built a

0:31

skill for that. This skill is called

0:33

triage, and I run triage on every single

0:35

open source repo that I run. It's great

0:38

for working through GitHub issues, you

0:40

can plug it into Jira, you can plug it

0:42

into any backlog to essentially flesh

0:45

out that backlog and turn it into

0:47

actionable tasks that can be picked up

0:48

by an agent, or to reject it, or to do

0:52

all sorts of various things with it. The

0:54

thing that drives this is that it has a

0:56

essentially a state machine encoded into

0:59

labels. So, you have two category roles

1:02

currently, this may change in future. We

1:04

have a bug role and an enhancement role,

1:07

which is very, you know, that's well

1:09

known to anyone who's done any work with

1:11

triage. And then, you have five state

1:13

roles, and these are kind of up for

1:14

grabs as well, like this is um something

1:16

that I probably will expand on in the

1:18

future, but for now this is working for

1:19

me.

1:20

Each of these five state roles

1:21

corresponds to a state that the ticket

1:24

can be in. So, the ticket might be needs

1:26

triage, in other words, it needs a

1:28

maintainer to look at it. You know, it's

1:30

a paused the maintainer has actually

1:31

done that work. It might be that it

1:33

needs info, waiting on the reporter for

1:36

more information. It might be ready for

1:38

an agent to pick up, in other words,

1:39

it's fully specified, ready for an AFK

1:42

agent to go and slam through the task.

1:44

Or it might need a human to look at it

1:47

now and to actually implement it.

1:50

Or we do a won't fix, where it just will

1:52

not be actioned. This is a state machine

1:53

because every single triaged issue

1:55

should carry exactly one category role

1:57

and one state role. In other words, you

2:00

can't have it be ready for human and

2:02

needs triage at the same time. This is a

2:04

good old state machine, and that's

2:06

familiar to anyone who's followed my

2:07

career, you'll know I would love state

2:09

machines. So, the way you use triage is

2:11

you can use it in a bunch of different

2:12

ways. You can either use it to triage a

2:15

specific issue, you can triage the

2:17

entire backlog and say show me anything

2:19

that needs my attention now. You can

2:21

say, "Okay, let's move this one to ready

2:22

for agent." And by the way, ready for

2:24

agent is a nice one because in order to

2:26

move something to ready for agent, you

2:28

need to write a brief for the agent

2:30

that's going to pick it up. And we have

2:32

inside here a agent brief template,

2:36

which allows you to kind of write this

2:38

ticket really nicely. So, you're

2:39

probably thinking, "Okay, I get the

2:40

basics of this. Show me this in action."

2:43

Turns out we have a bunch of issues that

2:45

need triage in Sandcastle. So, that's

2:47

what I'm going to do. I'm going to show

2:48

you how to use this skill by triaging my

2:50

own repo. You can see there's a bunch of

2:52

stuff that's already in here. Some of

2:54

the stuff that's ready for agent that's

2:55

been put into a PR already, some that

2:59

kind of needs triage that I need to look

3:00

at again, some that's totally unlabeled,

3:02

some marked as won't fix. So, I'm going

3:04

to open up a new Claude session inside

3:07

my repo, and I'm going to say triage,

3:09

and then I'll say, "Just give me all of

3:11

the open issues." Or rather, just the

3:13

ones that I haven't triaged yet. So,

3:15

let's ping this off and see what

3:17

happens. It should go and explore the

3:19

repo. It should pick up that I'm using

3:21

GitHub issues, and yep, look at it go.

3:23

Okay, it can see there are nine

3:24

untriaged open issues. I would like it

3:27

to Could you just walk through each of

3:29

these and add the basic labels to them?

3:32

Just doing initial triage for me so that

3:34

I don't need to make that many

3:35

decisions. One thing you're probably

3:36

thinking is how does the agent know what

3:40

can be actioned and what's even like a

3:43

good candidate to be actioned? Well,

3:45

inside the repo I have a dot out of

3:48

scope directory right at the top here,

3:50

which is a few things that I've already

3:51

triaged with the skill that marks it.

3:54

Basically, anything that I say, "Okay,

3:55

we're not going to do this." It

3:57

basically says, "Okay, you're going to

3:59

mark this feature as out of scope for

4:02

the future." So, for instance,

4:03

Sandcastle does not provide an

4:05

abstraction layer for composing

4:06

Dockerfiles or managing base images

4:07

programmatically. This is essentially

4:09

like an architectural decision record,

4:10

but specifically for features that we're

4:12

not going to implement. So, this means

4:14

that the agent, when it's looking at

4:17

issues, especially enhancement issues,

4:20

is essentially looking at the ADRs that

4:22

I've created at these out of scope

4:24

references and going, "Okay, I can just

4:26

close this straight away because this is

4:27

an enhancement that we're not going to

4:29

commit to." Okay, it's gone ahead now

4:31

and it's labeled a bunch of these. So,

4:33

it's labeled all of these as bugs and

4:36

all of these as enhancement and mark

4:38

them all as needs triage. Let's just

4:40

work down from the bugs first since

4:42

those are going to be, I think, a little

4:43

easier to grok.

4:45

Could you start with 477 for me? Let's

4:48

fire this off and see what happens. All

4:50

right, so it's come back with saying,

4:52

"Okay, a recommendation, it is a bug."

4:54

And it's recommending that I move it to

4:55

ready for agent. It's already had

4:57

substantial triage notes pinpointing the

5:00

exact root cause, reproduced with a

5:02

stack trace, etc., etc. However, I think

5:05

I want it to actually reproduce it

5:07

itself. It's being a little bit too

5:09

credulous about the

5:12

you know, about the stuff that it's

5:13

received from the person who's tried

5:15

this. So, I would like it to use another

5:17

skill of mine to say, "Diagnose this

5:20

yourself." going to make it walk through

5:22

essentially

5:24

reproducing the bug and trying to fix

5:26

the bug. So, we're actually going to do

5:27

this from within this same session. Now,

5:29

of course, whenever I'm doing anything,

5:31

anything with the LLM, I'm thinking

5:33

about my context usage, which is in the

5:34

bottom left here. I've sort of have a

5:36

budget of 100K for this session. We're

5:39

at 46.5K, that seems fine to me. We've

5:41

certainly got enough context to diagnose

5:43

this and fix this. I I a feeling this

5:45

error is just a small error, so that's

5:48

what's at the back of my mind all the

5:49

time. What's nice about the triage too

5:50

is that, you know, we've already marked

5:52

everything as it needs to, essentially,

5:55

so it's already labeled correctly. In a

5:58

future triage session, I would just go

5:59

back to the issues and I would just say,

6:01

"Okay, let's do this issue now." It's

6:02

funny how much of AI and how much of

6:04

managing AFK agents is essentially queue

6:07

management. This is what we're doing

6:08

here. We're just sort of pruning a queue

6:11

and acting as a translation layer

6:13

between the humans creating these

6:15

tickets and the AI that's going to

6:16

implement them. The way that my plan

6:18

prompt works for Sandcastle, which is my

6:20

AFK agent kind of software factory

6:23

thing, is it looks for the label ready

6:25

for agent and it only touches ones that

6:27

have been explicitly marked ready for

6:29

agent. Okay, it looks like it has found

6:31

the issue, so diagnosis confirmed. So,

6:34

the issue is is that a couple of

6:36

variables contain a literal task ID and

6:39

this particular syntax here, if it's not

6:42

passed in by the user, actually results

6:44

in an error. So, what it's saying is we

6:46

need to replace these task IDs with an

6:49

ID, so this kind of pattern here, and

6:51

that non kind of double curly brace

6:53

placeholder will actually not cause an

6:56

error, so that'll be fine. We can see

6:57

it's also added an idea of a feedback

6:59

loop here. So, a unit test scaffolding

7:01

simple loop and asserting that the

7:02

resulting prompt dot md

7:04

contains no unresolved task ID. That

7:07

looks great and it seems to be just

7:10

busting on with it without asking, which

7:11

is is fine in this case. We can see that

7:14

the diagnose skill actually uses a

7:15

similar setup to my TDD skill, where it

7:18

gets it to create the regression test

7:20

first, create the feedback loop first,

7:22

and then fix it within that feedback

7:23

loop. So, it's looking good. It's

7:25

created the full suite and the type

7:27

check and it looks like it's just ready

7:29

to go. So, this is another nice feature

7:31

of triage, which is once you've pulled

7:33

in the issue and understand it and

7:34

diagnosed it, you can just fix it there

7:37

and then if you want to. One thing I

7:38

often like to do is when I'm doing these

7:40

is actually have a two sessions running

7:41

at once. Is have one session over here

7:43

where I'm fixing a or triaging one issue

7:46

and then one session over on the left

7:48

here where I might actually be using

7:49

main to fix the issue. Okay, it's giving

7:51

me some kind of command request here.

7:53

I'm just going to say yes and then I'm

7:55

going to bump it into auto mode. All

7:57

right, so looking pretty good. It's

7:59

added a change set. It's added a fix

8:02

here. I'm going to get it to push this

8:05

to main and close the original issue.

8:07

One thing I could do is I could get it

8:08

to create a PR and then the PR would

8:11

reference the issue originally so that

8:12

when I closed or

8:14

when I merge the PR it would close the

8:16

issue. That's how I usually like to

8:17

work. But let's just push this to main.

8:19

I actually kind of did a bit of vetting

8:21

on this beforehand. I actually know what

8:23

the fix is and this is the fix. So a bit

8:25

of movie magic for you. So that is the

8:27

triage skill. It's essentially a way of

8:30

you and the AI working together to turn

8:33

messy human ideas into actual real tasks

8:38

that the agent can pick up and work on.

8:40

And this is specifically designed for

8:41

AFK agents. So you, in order to get your

8:44

agent working properly, you need a

8:46

backlog of tasks that it can pick up

8:48

over time. And having specific labels

8:50

for each of those means that the agent

8:52

is not going to stumble around on crap

8:54

tasks that aren't ready for it yet. And

8:57

the ready for agent kind of signal is a

8:59

really cool one because you just get to

9:01

see your backlog fill up with these

9:03

little marks that you know will be

9:05

implemented. Very cool. This also

9:06

matches well with my existing skills.

9:08

For instance, we can create PRDs here.

9:11

For instance, this is one about work

9:12

tree locking to prevent concurrent agent

9:14

access. This is a PRD and I've marked it

9:16

as ready for agent. So the agent will

9:18

see this, will pick it up. You can also

9:20

do this for tickets as well. So tickets

9:22

based on PRDs. This has the parent of

9:24

the PRD that we just saw and it's also

9:27

ready for agent. So that's going to be

9:28

seen by the AI picking this up. If you

9:30

dig this then you should check out the

9:31

skills page on AI Hero below here. I'm

9:34

adding all of the change logs onto this

9:36

page which are video change logs and

9:38

article change logs, so you can keep up

9:40

to date with what's happening. And you

9:41

can sign up to this newsletter, so you

9:43

get all my latest and greatest skill

9:45

updates, as well as tips on getting the

9:47

most out of agents. You may have also

9:48

noticed that this video is in a slightly

9:50

different style from my usual videos.

9:53

It's a little less cutty. I'm

9:55

experimenting with like a different

9:57

silence detection algorithm, and I'm

9:59

interested to know what you think. I

10:01

hope these sort of slightly longer takes

10:03

have a slightly more relaxed feel, a

10:05

little bit less relentless, and you get

10:07

a sense for my kind of natural speaking

10:09

cadence, sort of separate from YouTube.

10:12

So, if you dig this, let me know.

10:13

Anyway, thanks for watching, and I'll

10:15

see you very soon.

Interactive Summary

The video introduces a custom 'triage' skill designed to streamline project management for solo developers and teams. By using a structured state machine approach encoded in GitHub labels, the tool helps maintainers organize backlogs, filter out-of-scope issues, and prepare actionable tasks for autonomous 'AFK agents'. The presenter demonstrates how the agent can autonomously diagnose, test, and fix bugs using this framework, emphasizing the importance of clear queue management when working with AI developers.

Suggested questions

3 ready-made prompts