HomeVideos

LIVE: Poteto (creator of pstack) on shipping 1,000's of PR's a month at SpaceX

Now Playing

LIVE: Poteto (creator of pstack) on shipping 1,000's of PR's a month at SpaceX

Transcript

1707 segments

0:00

So, hello folks. I've got another treat

0:02

for you today. Last time on this kind of

0:05

podcasty thing, I suppose, we had Uncle

0:07

Bob and we talked about software

0:09

quality. We talked about agents. We

0:10

talked about lots of cool stuff. Now, we

0:13

have uh an incredible guest, someone who

0:15

I'm delighted to welcome on, who's been

0:18

exploding on Twitter recently about

0:21

software factories, um about increasing

0:24

the quality of your work, about

0:26

increasing your velocity and climbing

0:28

the trust ladder with agents so that you

0:30

can ship more and more and more. And it

0:32

is potato. Welcome. Thank you so much

0:35

for joining.

0:36

>> Thanks for having me. Yeah, very excited

0:38

to be here. Yeah, big fan of yours

0:41

>> and a huge fan of yours. I [laughter]

0:42

think people have been talking about

0:44

this like it's like the meeting of the

0:45

skill minds, the skill Mount Olympus or

0:49

something because both of us have very

0:50

popular skill libraries. Um I've not, as

0:52

I was saying before we started, I've not

0:54

used a ton of yours and like I want to

0:57

get all of the juice out of your brain

0:59

so that I can go and use it properly and

1:01

use it better. And I think where I want

1:04

to start with this is you gave a talk um

1:06

pretty recently like um about 10 days

1:08

ago and posted on X which went

1:10

absolutely nuts as about how I shipped

1:13

2,500 PRs last month to production got

1:16

about 3 million views or something on X

1:19

and I watched it and I loved it and I

1:21

recommended it and I kind of want to run

1:24

this as almost like a Q&A of that talk

1:26

basically of giving you because it just

1:29

I just had tons of questions about it

1:31

and I wanted to dive into it. And I

1:33

think where I want to start is you talk

1:36

about a trust ladder with agents where

1:39

you as you trust agents more, you can

1:42

get them to do better and better things

1:44

and or scale them to up to use more and

1:47

more agents. So what is your story of

1:50

how you climbed the trust ladder and how

1:53

did that work when like you got SpaceX

1:55

and

1:56

>> started climbing more and more?

1:58

So I think this the the journey sort of

2:00

began even before I joined cursor uh

2:03

which is now SpaceX AI. Uh [snorts] so

2:06

the story is um

2:09

after Meta so I I used to work at Meta

2:12

on the React team. [snorts] Uh I took a

2:14

month off uh because I was feeling kind

2:16

of burnt out and of course when what

2:19

what do you do when you're burnt out?

2:20

You go and start a new side project. Um

2:22

and so I started a side project. you

2:24

know, I was uh of course using AI to to

2:27

write code. Uh [snorts] but then I

2:29

started to realize uh you know, I was

2:31

spending like so many hours just

2:33

micromanaging one agent, right? And you

2:36

know, at the time, this was back in

2:38

February, maybe February, early February

2:41

or January, you know, people were really

2:43

obsessed with this idea of like

2:44

orchestration. This was like, you know,

2:46

before, you know, things like cursor,

2:48

you know, like the agents window was had

2:50

become popular. So people were still in

2:53

like like 2 land you know in their

2:55

terminal and they were all talking about

2:56

okay here you know I built a custom

2:58

orchestrator right and so of course I

3:01

had I was a bit nerd sniped by that

3:03

[snorts] and you know as I was building

3:04

my toy project uh I got nerd sniped by

3:08

oh how do I make my AI coding setup more

3:10

efficient and so you know I I kind of

3:13

started the journey there where I just

3:16

you know took a step back and realized

3:18

you know I was spending all this time

3:20

micromanaging a single agent you know I

3:23

was [snorts] creating skills and I was

3:24

like finding it quite difficult to

3:26

measure the output or the the result the

3:29

impact of the skill as well so I was

3:32

kind of flying blind but I was you know

3:34

iterating really fast um and um so that

3:39

project eventually sort of became the

3:42

basis of PAC even though I didn't know

3:44

it at the time um and a lot of some

3:48

tricks I had learned like building that

3:50

early set of skills. Actually, it's

3:52

still open source if you want to if

3:54

anybody wants to take a look. It's on my

3:56

GitHub like potato

3:58

noodle n o d l e. Um, and in there you

4:02

will see some skills and a brain

4:04

directory. [snorts] And so I was really

4:06

interested in this idea of how do I, you

4:09

know, extract my own ability, if that

4:12

makes sense, and give it to the agent,

4:14

right? cuz I was I I realized that you

4:16

know all I was trying to do was trying

4:18

to teach the agent to write code more

4:19

like me you know do do you do do

4:22

workflows more like me. So you know the

4:24

skills were like an entry point to doing

4:27

that.

4:28

Um and then you know after I joined

4:31

cursor uh I was starting to work on the

4:33

agents window and uh it had a lot of

4:36

performance issues. Uh it was it was it

4:39

was pretty laggy. Uh and so since I had

4:41

experience working in React, I was asked

4:43

like, "Hey, do you want to come and help

4:45

out uh with the agents window?" Um and

4:48

so the the this beginning of the cursor

4:52

journey was very manual. [snorts] Uh I

4:55

was deep in like looking at like flame

4:58

graphs and heap snapshots and trying to

5:00

see like why exactly is the app so slow.

5:03

Uh but then coming back to the same

5:05

realization like you know I was sort of

5:07

the bottleneck. I was doing everything

5:09

manually. I was sort of the meat proxy

5:11

in a way, right? I was the meat proxy

5:13

between my agent and Chrome DevTools. Uh

5:16

and I was like really annoyed by that.

5:18

>> And what month of the year is that?

5:20

Let's say where are we in the timeline?

5:22

>> Uh so I joined Cursor in March. So this

5:25

was like early early April probably

5:28

early April is when [snorts] you know uh

5:30

I joined and I didn't have any skills,

5:32

right? I had I I sort of abandoned my

5:35

personal skills because I didn't think

5:36

they'd be relevant anymore. Uh but then

5:39

working on the agents window uh and

5:41

[snorts] now working on grockbot uh I

5:44

sort of realized that a lot of the

5:46

lessons I had learned from those skill

5:48

time building the the initial set of

5:50

skills were very relevant especially

5:52

around things like verification

5:55

uh you know being very rigorous in your

5:57

work um [clears throat] because I think

6:01

from my experience even the the frontier

6:04

ones tend to

6:07

tend to take shortcuts. Uh they tend to

6:10

do the easy thing. Uh so uh a lot of the

6:14

skills that I've built have been around

6:16

how do I make the easy thing the right

6:19

thing? You know, how do I make that the

6:20

best thing?

6:22

>> The idea of sort of distilling your

6:25

expertise and turning what you do every

6:28

day into processes, that's something

6:30

that feels super familiar to me. That's

6:32

exactly what I've been doing with the

6:34

skills. And I suppose there's something

6:37

in that which is a lot of people think

6:39

domain expertise is getting less useful

6:41

now as people uh start to rely more on

6:44

AI where what do you think about that

6:47

just as a sort of vibe check before we

6:49

start talking [laughter]

6:51

>> I actually feel like domain expertise is

6:52

is more important than ever you know uh

6:56

I think I wrote this on my ex at some

6:58

point but you know at times I sometimes

7:01

think of you know AI as is like

7:05

especially as the models get smarter and

7:06

smarter and more capable and the

7:08

frontier models are just getting so good

7:10

like I [snorts] love Opus 5.5 by the way

7:13

um uh you know as the models get really

7:17

really really good it almost becomes

7:19

like the bottleneck is no longer the

7:21

agent right it becomes your ability to

7:24

express your intent and your goals in a

7:27

clear way that [snorts] the agent can

7:29

understand and actually carry out and

7:33

That's why I think you know like people

7:34

with a lot of domain expertise are

7:37

extremely have a have a huge advantage

7:39

in my opinion especially if you're a

7:42

little bit like you know tech technoc

7:44

curious you know so I I think of people

7:46

like you know like uh like a doctor or a

7:49

lawyer or you know someone who who has a

7:52

deep expertise in a particular

7:54

non-engineering domain and [snorts] if

7:56

they're actually just a little bit

7:57

techsavvy and they can figure out how to

7:59

use agents they can actually build

8:01

really really great products, right? If

8:04

they if they have a clear enough vision

8:06

in their head and they can articulate it

8:09

in a way that the agent can build it,

8:11

you know, I think that that is really

8:14

the the bottleneck these days is is like

8:17

the transfer of your intent, right, and

8:20

your vision to the agent.

8:23

>> Yeah. I've been obsessed with language

8:25

basically since agents um dropped. are

8:27

just obsessed 100% and thinking

8:30

constantly about the the composition of

8:32

words, how I can make things sharper,

8:35

what um what might be hidden in the

8:37

phrases that I'm using. And it's and

8:40

finding what I love is when you find a

8:42

word that the agent then hooks on to and

8:45

then goes, "Okay, I'm going to reinforce

8:47

that word. I'm going to reuse that in my

8:49

thinking traces." You know, I found that

8:51

with um TDD was an early example of

8:53

that. a lot of chat about TDD recently

8:55

of like, you know, people say, should

8:57

you use TDD with agents? Doesn't matter.

8:59

What you're doing is you're getting the

9:00

agent to think about TDD, getting it to

9:03

write tests, getting it to prioritize

9:04

things in a different way than it did

9:06

before. And that's why sort of grilling,

9:08

I think, works effectively. Grilling is

9:10

a

9:11

>> Yeah. Yeah. It it draws those words out

9:13

of you, right? or or at least it helps

9:16

the agent understand your thinking so

9:19

that they can propose those words to you

9:20

and you can pick up and say yes exactly

9:22

that.

9:23

>> Uh I've actually copied some of the the

9:25

tips that you've shared as well where

9:27

you know one of my favorite ones that

9:28

you've shared recently or or not or like

9:30

maybe in the past couple weeks is

9:32

[snorts] about uh reducing or

9:34

eliminating tautological tests. Like one

9:37

of my pet peeves of agents is like all

9:39

of the useless tests that they write.

9:41

And so, you know, that was one thing

9:44

where, you know, the word tutology,

9:45

right, is is is I guess, you know, not

9:47

many people necessarily know that if if

9:49

especially if English isn't your first

9:51

language, but there's a lot of meaning

9:53

to that word. And it's like it's almost

9:56

like compressed, right? Like you

9:58

compress a lot of intent and meaning

10:00

into words. And so I I I totally agree

10:04

with you. I think language I've always

10:06

been interested in language actually uh

10:08

like programming languages natural human

10:11

languages and how they came to be and

10:13

it's so interesting that now with agents

10:15

it's sort of like this meeting of

10:18

natural language with programming

10:20

language but it's all it's all language

10:21

out of the hood it's all communication

10:23

>> totally I did a drama degree right so

10:25

you know I've been thinking about

10:26

language and Shakespeare and stuff for a

10:28

long time [laughter] and so this all

10:29

feels very familiar

10:31

>> um so okay there's sort

10:34

before we get into like because I think

10:36

the thing I want from you is like

10:38

software factory stuff, right? Software

10:40

factory is the big buzzword. Software

10:42

factory is the thing that I'm thinking

10:43

about too. I'm sort of releasing a

10:45

course in that direction too.

10:47

>> And it's this sort of scaling yourself

10:50

up to un unrealistic numbers of PRs

10:53

basically or PR numbers that sound

10:56

ridiculous to people who don't

10:57

understand how this works. So where I

10:59

want to get to is sort of from people

11:01

who are doing kind of like one to five

11:03

agents today up to, you know, hundreds

11:06

of agents running at once and how that

11:08

sort of functions. And so I'd love to

11:10

hear about your metaphor of the Michelin

11:12

Kitchen instead of the software factory

11:15

because I think that says a bit about

11:16

the way you think about this stuff.

11:18

>> Yes.

11:20

Yeah. I I I've I've never really liked

11:22

the term software factory. Not because

11:25

you know it's not accurate but I think I

11:28

think the a lot of people when they

11:29

think factory right they don't

11:31

necessarily equate that with quality or

11:34

craft right things which are very

11:36

important to me and a lot of people and

11:39

technologists who work you know building

11:41

products we care about the user

11:44

experience we care about the things

11:45

we're building. So while so while I do

11:48

think software factory is an apt term,

11:51

it also I guess maybe conjures up

11:53

negative, you know, maybe sometimes

11:55

negative connotations. So Michelin

11:58

Kitchen is the thing that I've sort of

12:00

landed on where it's much more I feel

12:02

like it's much more aspirational and uh

12:05

I like the metaphor a lot cuz you know I

12:06

like food. I'm called potato of course

12:09

[snorts] and I like cooking [laughter]

12:10

and I see a lot of parallels right like

12:13

with food right when you're cooking a

12:15

meal for yourself for example

12:17

[clears throat] it's both utilitarian

12:19

like you're trying to just feed yourself

12:20

right and and survive uh but it can

12:23

actually be transformed into art right

12:25

and that's what what a Michelin starred

12:28

chef or even just a chef or a cook can

12:31

do with food is take something very

12:33

ordinary and turn it into a delicious

12:36

meal that you know takes you back to

12:38

your childhood days or something like

12:39

that. Um and so it almost like mirrors

12:43

that trust letter that I talk about

12:45

where uh you can sort of imagine your

12:48

own journey as a home cook, right? Uh as

12:50

a home cook, you are doing all of the

12:52

food, the cooking yourself. You cut all

12:55

the vegetables, you do all the prep

12:57

work, you do all the cleanup, you know,

12:59

you are the one man or one woman show

13:03

really. Um, and it's an interesting

13:07

thought experiment like, okay, if you

13:09

were to cook a meal and then you add

13:11

people, right, your your your partner

13:14

trying to your brother, your sister, and

13:16

now suddenly you have your whole family

13:18

in the kitchen. I think most people

13:19

would get very stressed by that, right?

13:21

The thought of, oh, so many people are

13:23

just mocking around in my kitchen. They

13:24

have no no idea where all the utensils

13:26

are.

13:27

>> I have a max capacity of one person in

13:28

the kitchen. Yeah, absolutely. So, I

13:31

feel like that that's really apt because

13:33

when you ask yourself that question of

13:35

how do I go from being a solo cook,

13:37

right, to having an army or even not not

13:41

even an army but a few sue chefs, right,

13:44

that that are helping me in the kitchen.

13:45

How do I think about dividing the work

13:48

in a way that makes sense? You know, I'm

13:50

not dividing work just for the sake of

13:52

it, but in a way that actually makes the

13:55

sum the to the better than, you know,

13:57

the total of its parts. And so the

14:00

Michelin kitchen metaphor to me like

14:02

works really well in that regard because

14:05

you know as a chef you're you know if

14:08

you become a chef you're in a position

14:10

where you're not necessarily cooking all

14:12

the food yourself anymore but you are

14:14

thinking you're almost like the tech

14:17

lead right for the kitchen where uh you

14:20

know chefs have to think about you know

14:21

not just cooking but they have to

14:23

basically organize the whole kitchen and

14:25

they're like the CEO of the kitchen they

14:27

have to think about when do you order

14:29

ingredients, how do you store them, how

14:30

do you prepare them, when do they have

14:32

to be prepared, you know, it's a whole

14:34

it's a whole job, right? That's not just

14:36

cooking. Um, and I think that again it

14:39

mirrors so much of how engineers write

14:42

code today where you are not writing the

14:45

code yourself anymore. You have agents,

14:47

right? But you as the human are still

14:49

responsible for the final outcome,

14:51

right? your name still is associated

14:53

with the work that you do, your

14:55

reputation and you know so how you set

14:58

up your kitchen right and how you set up

15:01

your skills your environment your

15:03

codebase I think are ultimately the new

15:07

ingredients that go into um building

15:10

product

15:12

>> yeah I think what I love about your

15:14

approach is the amount of focus that you

15:16

put into the environment that the agent

15:18

operates in right because I think a lot

15:20

of people they think, right, the agent

15:22

is good. I'm probably not going to be

15:25

able to make it better. Let's just trust

15:27

what these magic model people have put

15:29

into the harness and the model

15:31

combination. Uh, there's nothing I can

15:33

really do, like I can't mess about with

15:35

claw codes internals or something or

15:36

whatever you're using. Um, but what I

15:40

love about your approach, and it's

15:42

something I advocate for too, is that

15:43

you can change the environment the agent

15:45

operates in, right? you can make changes

15:48

in the codebase and also give it tools

15:51

for verification as well and allow it to

15:54

verify its own work. So the thing I I

15:56

loved about watching that talk is the

15:58

amount of focus you put in verification

16:01

and like that is the lever that you can

16:03

start to generate trust. Can you talk

16:05

about that and what that concretely

16:07

looks like? Let's start like looking at

16:09

practical ways that people can improve

16:10

their own processes, their own kitchens.

16:14

Yeah, I've I've I've said this a lot

16:16

actually that you know even if you don't

16:18

use PAC or you know your skills I think

16:21

that the single most important skill

16:24

that should be in your toolkit is

16:27

verification because without

16:29

verification and for for by the way for

16:32

those watching who don't know what that

16:33

means it's this idea that you can give

16:35

you can sort of give your agent uh hands

16:38

and eyes in a way that's the the analogy

16:42

I

16:43

where the agent is able to run the code,

16:46

right? And actually

16:48

>> [snorts]

16:48

>> uh interact with it like a normal human

16:50

user would and also do things like you

16:53

know debug it, you know, take traces and

16:56

snapshots. Um and uh funnily enough like

17:00

that was actually the first skill I

17:01

built when I joined Cursor. uh that gave

17:03

me a lot of that was that was the thing

17:05

that actually started to let me ascend

17:07

the trust ladder a little bit in a way

17:09

that some of the other skills I had

17:11

looked at or built had not really let me

17:14

do because no matter how good you know

17:17

some of the other skills were like the

17:19

how skill, the why skill, the unsop

17:21

skill were, I was still relying on me

17:24

right as the proxy between my agent and

17:27

the output. So that you know if the

17:30

agent can't actually see the result of

17:31

its work there's no way it can actually

17:34

iterate right and so this is where

17:36

people start to talk about loops this

17:38

idea of a loop and really I think the

17:40

term loop you know seems kind of uh

17:43

almost abstract like people like what

17:46

what what is a loop what is an agent

17:47

loop but really to me like the most

17:50

important part of a loop that allows it

17:52

to be a loop is the verification part

17:55

because the agent is able to to verify

17:58

by its own work and uh you know that

18:01

takes you out of the equation where now

18:04

I can actually do something like so the

18:06

very one of the very first use cases I

18:07

had for verification was you know like

18:09

the performance work that I was doing on

18:11

cursors agent window and I want I wanted

18:14

to get to a point where I could do

18:15

something called hill climbing uh which

18:17

is a term that I I think the labs uh

18:20

talk about a lot which is this idea that

18:22

you know you have some kind of rubric or

18:25

a way to judge or score something And

18:28

now because you have a loop, you can

18:30

have an agent continually try to make

18:33

improvements to that. Uh I think

18:35

Carpathy, Andre Carpathy also famously

18:38

released uh something called auto

18:40

research that has a lot of these ideas.

18:43

Um but yeah, verification I would say is

18:46

probably the most important skill in PAC

18:49

uh and many [clears throat] other you

18:51

know tool sets. Uh and I think it's the

18:55

most important thing to focus on. So a

18:57

lot of the a lot of I spent a lot of

19:00

time actually you know tuning the

19:02

verification the creative verification

19:05

skill um and also internally the the we

19:09

have so many verification skills now

19:10

like every app that cursor has or

19:12

spaceexai has has a uh verification

19:16

skill that is automaintained as well

19:19

[clears throat]

19:19

>> uh and it's become critical

19:21

infrastructure for our team because

19:22

everybody uses it

19:24

>> and you went pretty far with that too

19:26

right like you had a um in your talk I

19:28

saw that you actually built a custom CLI

19:30

for that too. So what does that CLI do?

19:32

Like how does it [clears throat] execute

19:33

things and why did you I mean that's

19:35

proof of how deep you're going right of

19:37

how much you're pushing that.

19:39

>> Yeah. So this is actually a tip I

19:42

learned early on where

19:45

um I guess you know back in January or

19:49

or late last year the thing that people

19:51

were concerned about was context window,

19:54

right? That was the big the big topic at

19:56

the time was how do I you know manage

19:58

the context window because you know

20:00

compaction summarization wasn't really

20:02

that good yet [snorts] and people were

20:05

always people had this there was almost

20:07

this meme in the community that you know

20:09

once your agent summarized or compacted

20:12

once it would become sort of stupid

20:15

right for the rest of your session. So

20:17

there was a lot of thinking around like

20:20

you know being very efficient with your

20:22

context usage and so that was actually

20:25

the inspiration for some of the uh the

20:28

CLI work inside of the verification

20:31

skills. I guess now it's less so about

20:34

context because uh you know agents are

20:38

much better or harnesses have gotten a

20:40

lot better with summarization.

20:42

Um, I still think there's some benefits

20:45

to, you know, uh, having a clean context

20:49

window. Uh, so the CLI is really just

20:51

more of a way for me to take the

20:53

deterministic parts of what the skill

20:56

does and encode that into a script or

20:58

CLI to reduce to kind of take away the

21:03

judgment that would otherwise

21:05

unnecessarily be used because with

21:08

judgment so I also think of you know

21:10

agents and skills in sort of like it's

21:13

like a gradient

21:15

you have some parts of the work that are

21:18

entirely judge measurement based right

21:20

you know something that requires thought

21:22

you know putting together multiple

21:24

pieces of context thinking um and then

21:27

you have the more deterministic parts

21:28

like I don't know if you wanted to uh

21:31

refactor some code right from one

21:34

pattern to another that's very

21:35

mechanical right you don't you don't

21:37

need an agent to think about it and come

21:40

up with it in a novel way each time

21:43

right and so that that was really the

21:45

inspiration for the CLI and you'll see

21:47

this in a lot of the other skills that I

21:48

built is like I try to extract out the

21:52

deterministic parts and turn that into

21:54

code and just leave only the parts that

21:56

actually require judgment to the agent.

22:00

So in a way I think of the seal as kind

22:01

of like your wrapper, right? It's a

22:03

wrapper with some light instructions

22:04

around how to use these custom tools

22:06

that are inside of the skill. Um but

22:10

yeah, I don't think the CL is really

22:11

that interesting in its own really. It's

22:13

not like a novel piece of software. It's

22:16

just something that interacts with like

22:17

Playright and the Chrome DevTools

22:20

protocol and calls a bunch of APIs. It's

22:22

like it's just a bunch of glue.

22:24

>> No, it's fascinating because it's a way

22:25

of hiding information from the skill,

22:27

right? It's a way of conserving the

22:28

skill, keeping the skill quite small, I

22:31

imagine, and then you're able to

22:32

delegate more of the complicated

22:33

deterministic stuff into a script within

22:36

the skill. So it's almost you're

22:38

compressing information and making the

22:41

agent do more consistent things more

22:44

consistently.

22:45

>> Yeah,

22:45

>> that's fascinating.

22:47

And it it [clears throat] also helps I

22:49

guess if you care about context window

22:50

it it does help because now the agent

22:53

doesn't need to uh you know re reinvent

22:56

uh things cuz uh one thing I had noticed

22:59

early on when we didn't have a CLI

23:02

was that uh well the the agent would try

23:05

to verify it work but it would basically

23:07

rebuild the world each time and then

23:09

every agent did it differently and I was

23:11

starting to notice like that's very

23:12

inefficient right I was wasting it it

23:14

was actually not just about context

23:16

usage but also speed, right? Like

23:18

because now an agent had to actually go

23:19

off and write the scripts or the CLI and

23:22

test it and you know and it doesn't work

23:24

and the last agent did it and it worked

23:26

but it discarded it. So it was just very

23:28

obvious at that point like I should just

23:30

turn this into a CLI and put that inside

23:32

of the skill uh so that every agent that

23:35

uses it now benefits from that same

23:38

piece. Um but I I also think like you

23:42

know it's a good push for people to

23:43

think about is how much of your skills

23:46

and rules could actually be

23:48

deterministic.

23:50

Um that's like another core thing or or

23:53

one of my core principles that I like to

23:55

think about is yeah how do I uh make

23:59

very efficient use of determinism and

24:02

non-determinism

24:04

and you know let Asians shine at the

24:07

non-deterministic parts right because

24:09

that's what they're trained to do. Um

24:11

and the other parts which are much more

24:14

mechanical or you know straightforward

24:16

can be just pure determinism. Um, and

24:19

you'll see this as well for things like

24:21

doing migrations. Um, which is another

24:24

big thing that I've I've talked about

24:26

is, you know, going from one technology

24:29

to another, especially one that is

24:32

better for agents, right? And a lot of

24:34

how you can do that migration is, I

24:36

think, through things like scripts and

24:38

CLIs, like the deterministic parts like

24:40

code mods, you know, like crawling the

24:44

abstract syntax tree and transforming

24:46

code literally mechanically, right? like

24:48

a script does it for you instead of the

24:50

agent.

24:52

>> Totally makes sense. I I mean I think

24:55

what there's another thing there which

24:57

is you're taking stuff away from the

25:00

agent and you're kind of putting it in

25:03

the environment too a little bit which

25:05

is let's say you have a a thing that you

25:08

notice the agent always gets wrong. You

25:10

want to make that um just impossible

25:13

within the environment. And that sort of

25:15

comes down to code quality as well. I

25:18

mean, I talk about a lot like having a

25:22

what a good codebase means, right? What

25:24

is a good codebase? And there's a

25:26

definition I like which is a a good

25:28

codebase is a codebase that's easy to

25:30

make changes in, right? Easy to um

25:34

change stuff without things screwing up.

25:36

And that means that you have a lot of

25:37

guard rails that you have a lot of um

25:40

the agent or the human is constrained to

25:42

very narrow paths. And that's again

25:44

something you talk about in your talk.

25:46

>> And you talk about this not only on the

25:48

kind of sort of automated checks side of

25:52

things. So linting and type checking

25:53

blah blah blah but also in the way you

25:55

design abstractions. And you guys even I

25:58

think built a framework uh for your

25:59

agent to work in too.

26:01

>> I think what I'd love to hear is you

26:04

obviously think of that as very

26:06

important, right? And that's how

26:08

important is that compared to other

26:10

things you could be doing like building

26:11

features or shipping work.

26:14

Yeah, I think that's a um I almost feel

26:17

like the new job of the engineer is

26:19

really to to spend time on the

26:21

environment. Um I almost actually wrote

26:23

a tweet about this yesterday, but I but

26:25

I didn't. But I think that

26:29

I think that you you know if you if you

26:31

haven't really spent time, you know,

26:33

building trust in your agents and

26:34

building skills and tools, you can get

26:37

stuck in this mode where you're very low

26:39

on that trust ladder, right? you don't

26:41

have a lot of trust in your agents work.

26:43

And so the only way to cope in that when

26:47

you're in that situation is just to kind

26:48

of lock in and micromanage your agents.

26:51

And that's very time consuming. And when

26:53

you're stuck in that mode, you don't

26:55

really have the luxury to think about,

26:58

you know,

27:00

uh higher level things like like making

27:03

yourself more productive. In the same

27:05

way that uh I guess analogy would be

27:09

like if you've never taken the time to

27:11

learn like your tools right as a

27:13

developer when you were writing code

27:14

yourself and you know you've never heard

27:17

of VS Code, you've never heard of Vim,

27:19

you only knew about Notepad [laughter]

27:22

uh and you had hadn't even heard about

27:23

Git. That's sort of the analogy. It's

27:25

like you you haven't spent the time

27:27

sharpening your own knives, right? And

27:29

so, of course, if you have a dull knife,

27:32

then everything's going to take a long

27:33

time. Um, and you're going to be you're

27:36

just going to be and and especially if

27:38

you know deadlines are looming, then you

27:41

don't have the now you're stuck in this

27:42

rut, right? Where where you you you

27:45

don't have sharp knives, you don't have

27:47

good tools, but you're under all this

27:49

pressure to ship, right? And so, all you

27:51

can do is just focus on that. But I do

27:54

think that, you know, if you can find

27:55

yourself the time to actually spend time

27:58

thinking about your setup, it's again

28:00

going back to the cooking, you know,

28:02

analogy,

28:04

it's like uh, you know, if you, for

28:06

example, if if cutting cutting cutting

28:09

the garlic is like super slow, right?

28:11

There are garlic mashers, right? You can

28:13

you can buy and you put it in the thing

28:14

and you like squeeze it out, right? It's

28:17

super fast. Uh, machines and tools were

28:20

invented for a reason, right? And so if

28:22

you're operating a Michelin kitchen and

28:25

your your your cooks had no tools, then

28:27

of course everything's going to be

28:29

extremely inefficient, very very, you

28:31

know, every every every cook is going to

28:33

make something up of their own. So I

28:35

think the tools and the determinism to

28:37

me are you know taking that part away

28:41

and and just like you said about

28:42

constraints as well. It's the

28:44

constraints are are to me as well like

28:48

uh actually a slight tangent on that is

28:50

uh I think we should talk about

28:53

TypeScript cuz like we we actually both

28:55

share like a background in Typescript

28:57

where you know you obviously have done a

28:58

lot of work with TypeScript and total

29:00

TypeScript and you know you're a leader

29:02

in that space and I uh had adopted

29:05

TypeScript pretty early and I had given

29:07

like a talk or two at Typescript conf uh

29:10

many years ago and so one of the the

29:13

talk that I did actually was about type

29:15

systems and constraining the

29:18

constraining types. Like one of my most

29:20

favorite things about Typescript is

29:22

actually type narrowing, right? This

29:23

idea that you go from a very broad type,

29:26

right? That could be anything and then

29:28

you through type guards and you know

29:31

type narrowing and you know runtime

29:33

checks you can actually narrow the space

29:35

and say like oh this isn't just a string

29:38

this is a very special type of string.

29:40

It's a constant, right? like I but I I

29:42

determine that through the type system

29:45

and in a way it's like uh there's a lot

29:47

of parallels I think to that with

29:49

constraints in your codebase where is it

29:52

you're you're constraining the space

29:55

right if you if you think about category

29:57

theory as well you know you're

29:58

constraining the the number of possible

30:01

types right that can can exist

30:05

and you're saying there's only one type

30:07

right and for for us like that framework

30:10

that I'm called Dune. Uh it's not an

30:13

open source framework. It's the the way

30:15

I describe it to people. It's it's kind

30:17

of like a internal Nex.js for our

30:20

Electron apps. Uh but it comes with a

30:23

lot of really really restrictive lit

30:25

rules and the codebase is designed in a

30:28

way that there's really only one way to

30:30

do something. So we make use a we make

30:32

use of a lot of conventional

30:35

patterns. So like features all go into a

30:38

specific directory. Well, every feature

30:41

has its own directory. As an example,

30:43

you know, there's like a a thing that

30:44

discovers features like through a

30:46

registry and like crawling the codebase

30:48

and stuff like that. But this

30:51

conventional pattern and the lint rules

30:55

make for an environment where it's

30:57

actually very hard to write bad code.

31:01

And that sort of frees up the it both

31:04

frees up your own mental uh you know

31:09

capacity as well as the agent sort of

31:11

doesn't have to think about that anymore

31:14

where it's just like oh there's only

31:15

there's I should just if I want to add a

31:17

new feature it just goes in the feature

31:18

the new feature directory and all the

31:20

code goes in there and I'm not going to

31:22

append to a god file right that was

31:25

really actually the inspiration for

31:26

those feature directories is the very

31:29

first couple of versions of Grockbot

31:32

were composed of like eight god files

31:36

which were like at least 10,000 lines

31:38

long if not longer and so I kind of had

31:40

to break it up into smaller pieces.

31:43

Uh but it was just observing you know

31:46

actually that's another important part

31:47

is observing how agents fail and then

31:50

every time you see a mistake every time

31:52

you see something that could be done

31:54

better you think you step back and think

31:57

how do I turn this into a lint rule? How

31:58

do I make it so that the code base makes

32:01

this impossible? Right? And it comes

32:03

back to me for my you know my background

32:05

learning Typescript and types uh type

32:08

systems is how do I constrain the space

32:12

so that you know I know precisely what

32:15

I'm working with and I think yeah

32:16

there's a lot of parallels there.

32:18

>> Totally makes sense. And don't I mean

32:21

it's funny that you mentioned TypeScript

32:22

and Goth files in the same sentence

32:23

because Typescript famously has a

32:25

25,000line type uh file.

32:29

>> [laughter]

32:30

>> Although I don't know if they've

32:30

rewritten that and go as they probably

32:32

have, haven't they? Um,

32:34

okay. So, environment is important. You

32:38

should watch your agent like a hawk to

32:39

make sure that any mistakes it makes.

32:42

You turn them into things in the

32:44

environment. And the benefit of the

32:45

environment is you're not overloading

32:47

your agent, right, in terms of rules, in

32:49

terms of things it has to remember. It's

32:51

just in the environment. And so it

32:53

stumbles into the rules and exactly um

32:56

you know bounces off them and hits them

32:57

at the right moment.

32:59

>> So okay, we still haven't talked about

33:01

the 2,500 PRs. Where do those come from?

33:03

Like how do you you've built your trust

33:06

ladder, you've worked on your

33:07

environment, and you understand, okay,

33:09

um I now want to scale up. So what are

33:13

the mechanics of that scaling? Are you

33:16

um initiating 2500

33:19

like chats per month? That can't be

33:21

right. So there must be are there any

33:23

kind of automated triggers that trigger

33:26

stuff in your repo? Like how do you get

33:28

the software factory kind of triggering

33:30

work by itself?

33:32

>> Right.

33:34

Um I'll definitely say that the

33:36

prerequisite to you know something like

33:40

a very high volume of of pull requests

33:43

um is the environment. you know, the the

33:46

stuff we just talked about where I

33:48

definitely would not have been able to

33:49

do this if I had not spent the time, you

33:52

know, thinking about the kitchen, right,

33:53

and the knives and the tools for my

33:55

Asians. And so, in a way, I I think of

33:58

this as

34:00

I've spent the time building one kitchen

34:02

and one restaurant. And now I I'm in a

34:05

position where I don't actually have to

34:07

be there anymore because the

34:10

environment, you know, that the same

34:12

analogy, right? It works really well.

34:14

[laughter] Yeah. You open chain of

34:15

restaurants, right? That's

34:16

>> Yeah. Exactly. Yeah. Exactly. It's like

34:19

you're Gordon Ramsay and you know, you

34:21

you've taught your executive chef like

34:23

all the tricks of the of coming up with

34:25

great menu. Uh and like the kitchen is

34:28

set up really well. Everything's just

34:30

perfect and you're now in a position

34:32

where you can open your second your

34:33

third restaurant. And [snorts]

34:36

I guess I I sort of see each project

34:38

that I work on, like each big chat is

34:41

sort of like a restaurant, right? and

34:43

and I'm I'm I have multiple of them

34:45

operating at the same time and I'm sort

34:48

of like helicoptering between them

34:50

sometimes some more than others

34:52

depending on how in the loop I am but

34:55

yeah definitely I think there's there

34:56

there are external triggers and context

35:00

that those projects don't have that

35:05

for a long time I was the proxy for that

35:08

so uh the best example I have is like

35:10

you know you have a project that's

35:11

working on a feature

35:13

uh or you're trying to fix a bug and

35:15

you're getting bug reports, but the bug

35:17

reports are going to things like Slack

35:19

or linear or X, right? And these are

35:24

external systems that aren't connected

35:26

to your inner loop. So, I like to talk

35:28

about this outer loop and the inner

35:30

loop.

35:32

Uh I don't know if I'm using the

35:33

definition correctly but to me my inner

35:35

loop is like basically my engineers my

35:38

agent engineers working on the code to

35:41

an building towards an intent or

35:45

snapshot of my intent right and the

35:48

thing about that is that the snapshot

35:50

can go stale right new information comes

35:52

to light that I then have to be the

35:54

proxy of and you know transfer that

35:57

context to my agent so you know if you

35:59

if you don't have these triggers pulling

36:02

information back into the interloop,

36:04

then you sort of have to play that role

36:07

where you're you're off, you know, in

36:09

Slack or X or or whatever and you're

36:12

gathering context, right? You're getting

36:14

context about bug reports, about feature

36:16

requests, about, you know, something

36:19

someone said about, you know, our

36:20

backend infrastructure has some

36:22

limitation, you know, all that

36:24

information, you have to f that across

36:27

to your agent. So that's where I think

36:29

like tools like Grogbot are really good

36:32

because they help you automate the outer

36:35

loop as well. And when you connect those

36:37

two loops, it's very very powerful

36:39

because now all of a sudden your agents

36:41

have the ability to get context for this

36:44

for themselves, right? If for example

36:48

uh you know either through just as a

36:51

simple example like maybe you have the

36:52

Slack MCP, right? Or you have uh your

36:56

own harness, right, that you've built a

36:58

Slack subscription into for a particular

37:00

Slack channel. Now all of a sudden you

37:03

can tell your agents, okay, subscribe to

37:05

the Slack channel. Every time there's a

37:08

uh, you know, bug report about

37:11

something, go off and triage that thing,

37:14

right? Go reproduce the issue, right?

37:16

Using the verification skills that we've

37:18

already spent time building and all of

37:20

those other skills that we've set up so

37:22

that I have a lot of trust, right? I

37:24

have a lot of trust that these agents

37:25

can actually go off and understand the

37:28

bug, you know, uh verify that the bug

37:31

actually still exists on main and it

37:34

wasn't

37:35

something about you maybe the users

37:38

setup or their data or maybe I don't

37:40

know they didn't install a dependency or

37:42

something like that like uh basically I

37:45

think uh creating that yeah creating

37:49

those two loops and connecting them is

37:52

really a very important part of the job

37:53

these days. Um, especially if you are

37:56

thinking about how to scale yourself.

37:58

So, a big theme here is really just like

38:01

always thinking about like what where am

38:04

I the bottleneck in this process? Why do

38:07

my agents need me, you know, to answer

38:09

this question? I I always like to think

38:11

about that. And so I try to think about

38:13

how do I actually get the agent to

38:15

answer its own question, right? But not

38:17

by hallucinating, not by guessing, but

38:19

actually real data. And you know, a lot

38:22

of people talk about this idea of a

38:23

company brain, right, or a context

38:25

graph. I feel like those terms are

38:27

unnecessarily

38:29

complex uh or even abstract. To me, it's

38:33

just about um how do I take information

38:37

that my agent needs that I would

38:39

otherwise have to go and pass it myself

38:42

and just teach it how to do it, right?

38:44

And that removes me from the equation.

38:47

And so how I arrive at 2,000 or however

38:50

many PRs is the fact that I have all

38:53

these loops set up, right? And so uh it

38:57

allows me to open chain restaurants,

39:00

right? I can I can really parallels

39:02

myself. So yeah, I'm not sitting there

39:04

creating 2,500 chats, right? Of course,

39:06

it's really like these projects are um

39:10

actually cursor has a new feature called

39:12

projects which are these like

39:13

coordinator agents. Um and so the

39:17

coordination co coordinator agents are

39:19

really good at sort of delegating and

39:21

not doing work of their own but they

39:23

manage and supervise like almost a list

39:25

of tasks and they spawn sub agents to go

39:28

and do them. And so I'm just constantly

39:30

feeding context or teaching the agents

39:32

how to get their own context and then

39:33

they're going off and doing the work for

39:35

me. Uh and really the big the last thing

39:38

I'll say to this is like the big unlock

39:40

for me for getting to 2,000 PRs is

39:43

starting from the question and working

39:45

backwards of how do I get to the point

39:48

where my agent can merge its own code?

39:51

Because

39:53

the obvious thing people ask me when

39:55

when they when I tell them, "Oh, I

39:56

shipped 2,000 and 2,500 pull requests

39:59

last month." They'll be like, "How did

40:00

you review that?" Right? That that's a

40:02

lot of PRs to review. Like your team

40:04

must hate you.

40:05

>> Do do you mind if we go there in a

40:07

second? Because a good question about

40:09

that.

40:09

>> Yeah. Yeah. Yeah.

40:10

>> I want to like this analogy is great. I

40:12

want to like deepen it a bit which is

40:15

before if you're like manually

40:17

initiating all those chats it's like

40:18

you're bringing the orders to your chefs

40:20

manually right whereas if you've got an

40:22

agent sort of like doing the expo then

40:25

you're able to sort of run it yourself

40:26

itself what is what does that concretely

40:30

look like then you've got these sort of

40:31

grock bots that are um subscribing to

40:33

channels pulling in Slack messages and

40:36

you it sounds like have a couple of

40:37

coordinator agents or like chief of

40:39

staff agents that like monitor that or

40:43

something like when you look at your

40:45

computer to manage your agents, what

40:47

does it look like?

40:49

>> Yeah. So, [clears throat] so uh this is

40:52

I guess somewhat confusing but we're

40:53

working on you know simplifying and

40:55

unifying but so uh there's graphbot uh

40:59

which or you know you can use other

41:01

tools of course as well but I I largely

41:03

think of these tools as like your outer

41:05

loop. These are tools like you know

41:07

Grabbot that have connectors right these

41:09

are connectors I guess they a lot of

41:11

people call them personal agents um but

41:14

they're connectors to things like your

41:16

email your calendar slack uh plaid I

41:20

don't know like all these different

41:21

services and they are a great source of

41:24

pulling context in to your work so the

41:29

same way that a human like you know if I

41:31

were if I was a manager and I was

41:34

leading a team of engineers years. Um,

41:37

you know, like when I used to work in

41:38

Netflix, one of the biggest things that

41:40

managers would talk about was this idea

41:42

of context not control, which funnily

41:45

enough, you know, has so much uh has so

41:48

much uh carry over to the agents world.

41:51

Uh, of you know, you you know, you you

41:54

of course can drive to an outcome you

41:56

want by control, right? Like by

41:58

micromanaging, but what you want is to

42:00

provide context instead, right? like

42:02

teach the agent, teach your engineers

42:04

how to be self-sufficient and then you

42:06

don't have to micromanage them.

42:09

>> Um, and so I see a lot of parallels

42:11

there. Uh, but yeah, graphbot. So,

42:13

concretely, I have some graph bots that

42:16

look at my Slack channels, look at my X,

42:19

uh, or my emails, uh, or linear, and

42:23

they're just constantly they have

42:24

routines that subscribe. So they're

42:26

constantly watching and I have I I'll

42:29

tell them things like you know uh I'll

42:31

watch for issues with uh bugs in the

42:34

graphbot desktop app as an example. Uh

42:37

and whenever you find that send it to my

42:40

cursor project. So one of the really

42:42

cool things about grabbot is it connects

42:43

to cursor. So cursor has uh like I I

42:47

just mentioned this new feature called

42:49

projects. And a project is really a uh

42:52

again like a you get a coordinator agent

42:55

that's in the cloud. It has its own

42:57

computer and all it really does is like

42:59

it's a manager of agents. It's like your

43:01

executive chef, right? Your your chief

43:03

of staff. It doesn't do the work itself.

43:06

It delegates and orchestrates and

43:08

manages the work of other sub agents to

43:12

you know that report to your chief your

43:16

chief uh of staff. And it basically is

43:20

responsible for driving the work forward

43:22

and managing things and uh passing

43:25

context to them.

43:27

>> So if you get a sudden burst of issues,

43:28

let's say you get 30 issues at once in

43:30

one payload or something or very quickly

43:32

the coordinator agent can figure it out

43:34

and delegate.

43:34

>> Yeah, exactly. It gets like uh you know

43:36

30 the 30 or so payloads and spawns a

43:39

sub agent or a single coordinator agent.

43:41

It can actually do a bunch of different

43:42

topologies of agents and it will sort of

43:46

figure out the best way to uh you know

43:48

efficiently distribute the tasks to your

43:52

team of agents. Um so I use uh cursor

43:57

projects a lot um and I also use grapot

44:01

a lot and cursor projects are my inner

44:03

loop and grabbot is my outer loop.

44:05

Grabbot takes all the context, external

44:08

context, gives it to the projects

44:11

because it can actually just send

44:12

messages to those projects, right? You

44:14

don't even have to open cursor. You can

44:16

just tell your grabbot, okay, create a

44:18

project, right, for these series of

44:20

tasks. They're all related, right? Maybe

44:23

as an example, you know, you've had a uh

44:26

[clears throat] a big burst of issues

44:28

that are all about performance, right?

44:29

Your app is slow uh and they're all

44:33

connected, right? Maybe some of them

44:34

even have a similar fix, right? But and

44:38

you can certainly go off and just spawn

44:40

one agent per task, but then you've lost

44:42

that sort of thread between them, right?

44:45

And and you may duplicate work or you

44:47

may not really think about the higher

44:48

level problem. You know, sometimes when

44:50

you you you solve bugs, you know, it

44:52

helps to have multiple bug reports that

44:54

are are slightly different because it

44:56

helps you really, you know, zoom out and

44:58

see actually, you know, the problem when

44:59

I looked at this one report, I thought

45:01

the bug was here, but actually when when

45:03

I see the other multitude of bugs is

45:05

actually up here,

45:07

>> right?

45:07

>> Yeah. Got you. So that that's why you

45:09

have so many agents in that loop then,

45:10

right? Because it's not just you have um

45:13

like you have a bug report comes in, you

45:15

spawn a single agent to look at that bug

45:17

report. that a that single agent will be

45:19

duplicating work with other um other

45:21

agents, right? Because if there are

45:23

multiple bug reports coming in through

45:24

the same thing, that can be duplicated

45:25

work.

45:26

>> Yeah.

45:26

>> Fascinating. [snorts]

45:28

>> That's really fascinating. Okay. And so

45:30

this just this endless series of

45:32

triggers um coming from real users

45:34

reporting real reports um builds up this

45:38

sort of and accelerates the factory sort

45:40

of adds more orders in. Other than bug

45:43

reports, are there any other sources

45:44

that you use for like um accelerating

45:48

for pushing these PRs?

45:51

>> Uh well, funnily enough, it's some of it

45:54

comes from uh reading the code, too. So,

45:57

I guess I have sort of uh well, so to

46:01

clarify that, you know, the 2,500 PRs,

46:03

they're not obviously like 2,500

46:06

features, right? they are a lot of the

46:08

work actually is spent on gardening like

46:12

another term that I really love. Uh

46:15

where

46:17

so I guess this is more important when

46:19

you have a big team of engineers human

46:21

engineers that you work with where and

46:25

also this goes back a little bit to what

46:26

I was talking about with the

46:27

environment. You know, setting up a

46:28

really good environment that doesn't

46:29

just help you and your agents, but

46:31

everybody on your team, right? Think of

46:33

a new hire who doesn't have a lot of

46:34

context on all of your engineering

46:36

practices joining your your team. And if

46:39

you have a really good environment, they

46:41

can be productive from day one, right?

46:43

They can they don't have to like, you

46:44

know, make open a bunch of lowquality

46:46

PRs. They can start, you know, they can

46:49

start just turning out really good code.

46:52

Um,

46:54

and uh, yeah, I think I I sort of lost

46:57

my train of thought.

46:58

>> I've got a I've got a followup, which is

47:00

what's what are the mechanics of like

47:02

triggering

47:03

>> like how when when do you trigger a a

47:06

agent

47:08

to go and look at the code, right?

47:10

Because some people might say, "Oh,

47:11

let's just do that every hour or

47:13

something or like on a chron job or

47:15

what's

47:16

>> Oh, yeah. Yeah. Yeah. Yeah. Uh I saw

47:19

some of your recent tweets as well about

47:20

like you know the some of the tweets

47:22

you've been doing which are great for

47:24

setting up your routines. Uh I have some

47:27

routines like that as well. Um so uh one

47:32

of them is uh

47:35

like looking through just another simple

47:38

example is you know [clears throat]

47:40

React has a lot of foot guns. Um, so,

47:43

uh, as as I'm sure you're aware. And so

47:45

I have an agent that's just constantly

47:47

looking for band patterns. And the

47:50

interesting thing about that one is that

47:52

I don't actually tell it to fix the

47:53

issue first. I tell it to append it to a

47:55

document. And then every couple of days

47:58

I look at it and I see actually these

48:00

are all the same thing, you know, and so

48:02

that gives me, you know, you almost want

48:04

like a buffer, a queue. Sometimes that's

48:07

actually more effective than just

48:09

spawning off a couple of like a lot of

48:10

sub agents to fix every single thing

48:13

because when you are [clears throat] in

48:15

kind of pure execution mode and just

48:17

trying to like you know f uh you know

48:20

execute on the orders that are coming in

48:21

very fast you sometimes miss the big

48:24

picture. So sometimes having a buffer

48:26

forces you to think about the big

48:28

picture because you you you have these

48:31

artifacts and things that you can look

48:33

at as a human um and sort of use your

48:36

own human judgment to or I guess you can

48:39

use an agent to do that as well. But you

48:41

give the [clears throat] agent and

48:42

yourself a way to identify patterns,

48:46

right, that you might otherwise miss if

48:48

you're just only solving each bug at a

48:51

time. And that's also really the benefit

48:53

of having something like a chief of

48:54

staff agent is uh it can see the forest

48:59

right uh in addition to actually doing

49:02

the execution.

49:04

>> Fascinating. That's I mean my brain is

49:07

exploding a bit there with the sort of

49:08

chief of staff at the software factory.

49:10

I might have to change some of the

49:12

course that I'm filming next week.

49:13

That's [laughter]

49:15

>> uh all right. Let's talk about let's

49:16

talk about review, right? because this

49:18

is the reply that you get, you know, is

49:20

>> did you read did you taste all 2500 of

49:24

those dishes as they swept past you?

49:27

>> And I assume the answer is a variety is

49:30

a version of no.

49:34

>> Yeah, I think you you you don't want to

49:36

be in a position where you're not

49:37

tasting your food ever again. Uh but you

49:40

also, you know, for scale, you cannot be

49:42

tasting every single dish that comes out

49:44

of your kitchen, especially if you have

49:46

multiple restaurants. So it becomes more

49:48

about sampling right and thinking about

49:50

the processes in the same way that you

49:53

know if I guess maybe this is where the

49:55

the the factory analogy is a bit more

49:58

apt is you know as a quality supervisor

50:01

on a factory you you can't look at every

50:04

single item you sample right you take

50:06

you you you go in there every day and

50:09

you look at the quality of the pull

50:11

requests you look at the code that the

50:12

agents are writing and you scrutinize it

50:16

very rig rigorously and you think about

50:19

all the inefficiencies, the bad patterns

50:21

that the agents are doing and then you

50:24

think about how to course correct the

50:26

environment, right? Not not that single

50:28

agent. Uh because if maybe if it if it

50:33

was a one-off incident, it's fine. You

50:35

know that maybe there's nothing to fix

50:36

there. But if you actually notice that

50:39

multiple agents are are having the same

50:41

issue, right? They're taking the same

50:43

shortcut. they're they're propagating

50:45

the same workaround everywhere. Uh

50:48

that's a sign that you should go off and

50:50

think about how to uh amend your kitchen

50:53

or your factory, right? Like thinking

50:55

about your skills, your constraints,

50:57

your lints, your type systems um and

51:02

setting or adjusting it so that that

51:04

problem doesn't happen again. And when

51:06

you do that enough times, then you get

51:08

to a place where the codebase is again

51:10

like the environment is so constrained

51:13

and so it guides you so well that you

51:16

can just you can just step away, right?

51:19

That's the dream. And I'll I'll

51:20

definitely say it um it's very hard to

51:23

get to this point. I don't want to sell

51:26

this as like, you know, something that

51:27

you can just do easily by using PAC.

51:30

Like it takes a lot of time and effort

51:32

to think about your code and where you

51:35

see your agents failing and thinking

51:38

very thoughtfully, intentionally and

51:40

setting up guard rails and constraints

51:43

so that they do the right thing by

51:45

default.

51:47

>> And you're not [clears throat] like if

51:49

to go back to the software factory

51:51

analogy, this isn't a dark factory,

51:53

right? This the lights are on, right?

51:55

>> It kind of is. Yeah, actually.

51:56

>> Is it?

51:57

>> Yeah. Well, it's dark in the sense that

51:59

so um it's dark in the sense that well I

52:02

think if my agents are merging their own

52:04

pull requests it's sort of become dark

52:07

where I go to sleep my agents now work

52:09

24/7

52:11

uh I have I have the equivalent of like

52:14

more than 10 chiefs of staff right each

52:16

working on a different area like for

52:18

example I have one that's working on

52:20

performance of the Grockbot desktop app

52:23

I have one that's working on uh fixing

52:25

bugs that users report I have one that's

52:28

exploring rewriting it in a different

52:31

language just for fun, you know, like

52:32

what if what if, you know, just

52:34

reimagining what what it would be if it

52:35

was like a native app. It's just a toy.

52:38

Um, but the idea is like yeah, I uh I

52:41

when you spend the time setting up your

52:43

environment, I've gotten to a point

52:45

where I review the pull request after

52:47

it's landed, right? I I tell my agents

52:50

full autopilot is is something that you

52:52

can do in in PAC and that will trigger

52:56

off this very intense rigorous

52:58

verification loop where it will spawn a

53:00

bunch of verifier agents for every pull

53:03

request and it will fuzz right fuzzing

53:06

meaning that it will actually run the

53:07

application. It's going to click around

53:09

and try to use it like a real human.

53:11

look for regressions, look for bugs in

53:13

your implementation and um it will try

53:17

to find issues with the thing and then

53:19

it will fix it itself. It'll do that

53:21

again and eventually get the PR to a

53:24

state where it can land.

53:26

Uh so it does it does it is quite token

53:29

intensive. You can tune this of course.

53:32

Uh so you know instead of like 10

53:34

verifier agents you might do like one,

53:37

right? Or you just tell the agent to

53:39

verify it's done work. But yeah, the key

53:41

thing is the verification part is really

53:44

the key piece that gives me a lot of

53:46

confidence

53:48

that I guess verification plus the

53:49

environment, right? It's these the

53:51

combination of these two things that

53:52

allow me to step away and say agents go

53:54

off and merge your thing. I'll review it

53:56

in the morning by looking at my commit

53:58

history

53:59

>> and if I see problems, I go and course

54:02

correct,

54:02

>> right? And I'll go and revert or modify,

54:06

add new link rules and whatever. Um, so

54:10

it does it does take time to get to that

54:12

point, but once you get it, oh, it's so

54:14

it feels so magical. Uh, I I I tell

54:17

people like I'm sleeping so much better

54:19

now because, you know, it took it the

54:22

very first day I turned on the sort of

54:24

dark factory was very scary because I

54:26

was like, "Ooh, what if I call the SE,

54:28

right? What if I break something

54:29

overnight?"

54:30

>> Uh, and it took a lot of it took a lot

54:34

of uh bravery, I think, to do that, but

54:37

>> somehow I did it. And yeah, now I'm in a

54:39

place where my Asians are are merging

54:42

their own code while I sleep.

54:43

>> It sounds like

54:44

>> I think it's dark in that sense.

54:46

>> Yes, it's dark sometimes, right? You do

54:48

sometimes.

54:48

>> That's true. That's true.

54:49

>> Because I think of a dark factory is

54:51

like almost like if you take the

54:53

original definition of Kapathy's vibe

54:55

coding, right, which is the code almost

54:57

doesn't exist. You forget that code

54:59

might be a thing. I think your approach

55:01

is totally different from that, which is

55:03

that code and the environment is

55:06

essential. And if the code in the

55:07

environment are bad, then you will get

55:09

bad outputs. Garbage in, garbage out. So

55:12

I I I think this is a this is a

55:15

different thing. It's like, you know,

55:18

the I don't know, maybe there's a dimmer

55:20

switch or something, right? Like, you

55:21

know, some parts of dark, some parts

55:23

were light. This is why the maybe the

55:24

the kitchen is a better analogy

55:27

>> the restaurant because you know even as

55:29

a as a as a

55:32

as a [clears throat] restaurant

55:33

restaurant you still might go to your

55:36

restaurants every now and then to take

55:37

take a peek in taste the food right

55:40

>> uh I think that's

55:41

>> the idea of sampling instead of blocking

55:42

I think is really important

55:44

>> I think what would you say to people who

55:46

are in I guess you're obviously in a

55:49

pretty security conscious environment

55:50

where you're working very security

55:52

conscious

55:53

>> mh [clears throat] Um maybe there are

55:56

folks working in like um medical

55:58

applications or law or finance or

56:01

something. I think of the like some PRs

56:04

are kind of like two-way doors which is

56:06

you can merge it and then revert it,

56:08

right? It's cheap back through. But

56:10

there are some PRs that are one-way

56:11

doors, right? That will cause data loss

56:13

of some kind that will

56:16

>> do something that can't be easily walked

56:18

back. How do you deal with situations

56:21

where most of your PRs, let's say, are

56:23

one-way doors? Like, is this something

56:25

you just wouldn't recommend or like what

56:27

do you think?

56:28

>> Yeah, I think that's a really good

56:30

question. I think that

56:32

it all comes back to me to the quality

56:34

of the verification that you're able to

56:37

um get out of your agent. And I think

56:41

for domains where the work is

56:44

verifiable,

56:46

this is easier, right? and the the

56:48

oneway doors become two-way doors in a

56:51

sense.

56:53

But I guess I don't know if you're

56:54

working on something that is like

56:56

extremely

56:57

is very hard to verify programmatically

57:00

then I think yeah you're definitely in a

57:02

position where it's very hard to get to

57:04

that point. Um so I do think like yeah

57:07

verifiability of the domain is an

57:10

important aspect to be able to do this.

57:13

Um and software engineering is just one

57:15

of those things where it's quite

57:16

verifiable in in a lot of cases maybe

57:19

not totally um you know like other

57:22

domains like mathematics I think are

57:24

another example of not all of it of

57:26

course but some aspects of mathematics

57:30

can be verifiable

57:32

if you write a proof for example um and

57:35

so yeah I think it's a great question

57:38

that I don't really have the answer to

57:40

and I think that this is something the

57:41

industry and us as engineers will have

57:43

to figure out is you know my sort of uh

57:48

hope and prediction for the future is

57:49

that we'll see more and more interesting

57:52

new agentoriented programming languages

57:56

and one of the most fascinating ones

57:57

that I've seen so far is this one called

57:59

bend bend d bend um and that language is

58:04

one where it kind of marries programming

58:08

with proofs right there used to be a

58:12

time, you know, where you actually had

58:14

to write your proofs in a different

58:16

language. And proofs, by the way, for

58:18

those uh who who aren't familiar is this

58:20

idea of uh that you can sort of formally

58:24

verify that some code is correct

58:26

mathematically, right? Especially if

58:29

you've written your code in a very

58:31

functional programming way.

58:34

uh but for the longest time you had to

58:36

do that in a separate language like lean

58:38

or tla+ or uh I'm blanking on some of

58:42

the other other examples but uh like

58:45

languages like that where you would

58:46

construct the mathematical proof and

58:48

then use a solver essentially to det

58:52

that that you've covered all the cases

58:54

you don't have like a race condition or

58:56

or whatever.

58:58

So yeah, I think

59:02

trying to sum up the question, I think

59:03

yeah, if you are in a position where you

59:05

can figure out how your agents can truly

59:07

verify the work in a way that gives you

59:10

confidence, you can actually, you know,

59:13

uh have the PRs merge cuz if it

59:16

compiles, right, if it if the proofs

59:18

show you that it's correct, then why

59:21

wouldn't you just merge it? Um but of

59:24

course, yeah, not all the means are

59:26

verifiable. Yeah, it's a tough one. Um,

59:31

okay. I think we've got to think about

59:34

wrapping up because we are nearly on the

59:35

hour. Have you Have you got something

59:36

after this? I mean, I've got something

59:38

before I give my son dinner, but

59:40

>> I I can go a bit longer after you.

59:42

>> Okay, let's let's go five minutes longer

59:43

then. Um I think I just want to have one

59:46

more question which is

59:49

I think I want to ask how you see Pstack

59:55

and how you see skills in general like

59:57

in terms of we talked about this before

60:00

we went on air which is like people

60:02

think of as like my skills versus your

60:05

skills and how do you combine frameworks

60:08

together? How do you use Pstack with my

60:11

stuff? like what should you take from

60:13

each one and because I think I see

60:16

skills as sort of just derived from

60:19

process basically like they're just

60:21

processes turned into words and I would

60:25

love to know

60:27

how you recommend people take Pstack and

60:29

take my stuff as well and turn it into

60:31

their own processes.

60:35

I think you shared a tip actually today

60:36

that I thought was actually very

60:38

relevant, which is this idea that you go

60:40

off and look at your previous

60:42

transcripts, right? And you sort of mine

60:44

for information of, you know, your own

60:48

look through your own your own prompts,

60:50

right, to the agents where you correct

60:51

them where you have to constantly

60:53

intervene and uh you know take that

60:57

higher level learning and turn that into

60:59

a a reusable skill, right? So that

61:00

agents stop repeating that mistake. I

61:04

think that uh PAC and your skills are

61:08

very complimementaryary. I I totally

61:10

agree with you that they're like a skill

61:12

is really much just process. I mean it's

61:14

just at the end of the day a skill is

61:15

just English or or language. It's just

61:18

markdown.

61:19

>> Um and I think you can you can

61:21

definitely weave them, combine them in a

61:24

way that makes sense to you. But I do

61:27

think that uh everyone should have their

61:30

own set of knives, right?

61:33

I I keep going back to the the the

61:35

cooking analogy, but it's so apt because

61:37

like, you know, every chef when they go

61:39

to a different job, right, when they go

61:41

to a different restaurant, they carry

61:42

they bring their knives with them. The

61:44

tools go with them, right? And so trust

61:46

to me is really about trust in your own

61:48

tools. And when you spend the time

61:51

sharpening them and understanding them

61:52

really, really well, you can do great

61:54

things. And everybody's skills and tool

61:56

set is going to look different. you

61:58

know, someone might find a lot of

62:00

success combining, you know, like your

62:03

grill me with docs, uh, or wayfinder

62:05

skill with some of the execution skills

62:07

in Pstack as an example. Some people

62:10

might use more of your skills, some

62:11

people might use more of my skills. I

62:13

think at the end end of the day, it

62:15

really just comes back to how much do

62:17

you trust, you know, me and Matt, right?

62:19

Like if you if you trust us both, of

62:21

course, use our skills, but I also

62:23

encourage you to, you know, look at your

62:25

own transcripts. Um um tell the agent to

62:29

look through, you know, some of the all

62:31

of the patterns that you've used, the

62:33

the times you've had to intervene, you

62:36

know, suggest turning them into lint

62:38

rules or new skills, right? The the past

62:41

chats I I often say is like a a treasure

62:44

trove of context because that, you know,

62:46

it's it's like the process materialized,

62:51

right? Like it's the real process. It's

62:53

not an abstract idea in your head. it's

62:55

the actual thing right and you can

62:57

actually see how it happened in practice

62:59

and extract so much information from

63:01

that and there's so much so that I

63:02

actually turn I have a skill in pac

63:04

called recall which is exactly that um

63:07

where this was a pattern where

63:10

you know I was working in I was working

63:12

on a similar problem so specifically I

63:14

was working on virtualization for the

63:17

cursor application and there were a lot

63:18

of bugs and so you know every time I

63:21

started a new chat I was like ah this

63:22

there's so much good context from the

63:24

last one So, you know, I want to bring

63:25

it over to the new chat. How do I do

63:27

that? And that's where the transcript

63:29

came, uh, you know, looking at the past

63:31

transcript came about. And then recall

63:33

was just a way for me to collapse

63:35

collapse and compress that workflow into

63:37

a skill. So that I didn't have to just

63:39

say I didn't have to write a long essay

63:42

every time. Go look at all these chats,

63:43

right? And, you know, blah blah blah.

63:45

So, I I largely think of skills,

63:47

especially as agents get more capable as

63:50

really encoding workflows. you know

63:52

skills from last year were really more

63:54

about like almost like implementation

63:56

details like here are the exact script

63:59

commands you know you should use right I

64:02

think with the latest models you can

64:04

just delete those parts and just really

64:06

focus on the workflow right until it

64:09

it's more the skill becomes more like a

64:11

series of steps

64:13

a series of your process uh and I think

64:16

over time we'll see that skills get

64:18

smaller and smaller you know more

64:20

compact Um, and yeah, they're very

64:24

compatible. Or you can, you know, if you

64:26

want, why not read our skills, right,

64:29

and com and combine them in your of your

64:31

own, right? Combine Wfinder with potato

64:34

mode and make your own custom mode,

64:36

right? Like like skills are the the

64:38

thing I love about skills that is are

64:40

that they're so malleable. You can do

64:42

anything you want. It's just language.

64:44

>> Absolutely. There's nothing magical in

64:46

them, right? They're just words. And

64:48

>> Exactly. If if there is any magic in

64:49

them, it's just the words chosen and the

64:52

phrases used and the thinking that's

64:55

been done to turn those like take

64:59

abstract process and turn them into

65:01

language. And once that thinking has

65:04

been done, then it's just there. It's

65:05

available. It's on the surface and you

65:06

just nick it. Um Lauren, thank you so

65:10

much. This has been glorious.

65:11

>> Yeah, this has been super fun. I really

65:14

enjoyed talking to you. Hope we can do

65:16

it again. I'd love to do it again. I'd

65:19

love to do it again. Absolutely. Um

65:20

yeah, we'll check in in uh

65:22

>> I don't know. Yeah, Monday. Let's do it.

65:25

>> Yeah, [laughter] let's do it. Part two.

65:28

>> Well, thank you so much. I'm going to

65:29

close the stream here. Laura and I will

65:31

uh chat a little bit and stay here. But

65:33

thank you guys so much for watching. The

65:34

glorious.

Interactive Summary

This episode features a discussion on how to scale software development using AI agents. The guest, Potato, explains the 'trust ladder'—a framework for increasing the autonomy of AI agents by building robust environments, implementing rigorous verification, and using deterministic tools. The conversation touches on the 'Michelin Kitchen' metaphor for managing agents, the importance of domain expertise, and how to automate an 'outer loop' that feeds context into the 'inner loop' of code execution, allowing for massive scaling in PR volume.

Suggested questions

4 ready-made prompts