HomeVideos

How I Shipped 2000 PRs Last Month — Trusting AI Agents|Grok Bot|Lauren Tan

Now Playing

How I Shipped 2000 PRs Last Month — Trusting AI Agents|Grok Bot|Lauren Tan

Transcript

989 segments

0:01

Hi, my name's Lauren. You might know me

0:03

as potato on X and I work on Grok Bot at

0:07

SpaceX AI.

0:09

So, last month I did something pretty

0:12

crazy.

0:13

I shipped 2,000 pull requests to

0:16

production.

0:19

A lot of how I'm able to do this is

0:21

through trust.

0:22

Um and I think a lot about trust in

0:25

terms of

0:26

how I can trust my agents to produce

0:30

high-quality work even when I'm not

0:31

there.

0:33

And my argument and thesis for this talk

0:35

today is that if you set up your

0:39

environment for your agents really,

0:41

really well,

0:42

you can end up with something that looks

0:44

more like a personal or even team

0:47

software factory

0:49

where you're producing very high-quality

0:51

code at much greater rates than before.

0:56

Uh but I'm not a fan of the term

0:58

software factory.

1:00

I like the analogy of a Michelin kitchen

1:02

better

1:03

where, you know, um as technologists,

1:06

we're not really producing we're not

1:09

mass producing a product on an assembly

1:11

line.

1:12

But the work that we do looks very

1:14

creative. It's very

1:16

um you know, it's the act of building a

1:17

product. And it's art in some sense.

1:21

So, you know, even though with agents

1:24

we're not

1:26

cooking the individual components that

1:28

go into the product anymore,

1:31

um

1:32

we're still responsible for the final

1:34

outcome

1:35

and thinking about our kitchen setup.

1:39

Right? Because depending on how you set

1:41

up your line cooks, your sous chefs, you

1:43

know, the kind of equipment they have,

1:45

the kind of training they have,

1:47

you know, dishwashers and the ratio I

1:50

guess of line cooks to dishwashers,

1:53

um all of these ingredients go into

1:56

making the final product. Um and I think

2:00

this analogy is really apt.

2:02

So,

2:04

uh I wanted to talk a little bit about a

2:07

story before we begin.

2:09

Where 6 months ago, when I first joined

2:12

Cursor, before we were a part of SpaceX

2:15

AI,

2:16

I obviously had, you know, no agent

2:18

skills to to use, right? I just joined

2:22

the company. It was

2:23

um a fresh code base and a fresh product

2:26

that I was working on.

2:28

So, at the time,

2:30

um Cursor was building the replacement

2:33

for the Cursor IDE,

2:35

which is the new agents window.

2:38

And uh before I joined, the Cursor agent

2:41

window had quite a lot of performance

2:43

issues.

2:44

Um and my manager at the time asked if I

2:47

was able to help them.

2:49

Um and since I had spent some time on

2:51

the React team before I joined Cursor,

2:54

uh it seemed like a good fit.

2:57

Uh but when I first started, I quickly

3:00

realized how manual this process was.

3:03

Uh obviously I've done performance work

3:05

before,

3:07

um but at the rate at which pull

3:09

requests were being landed, it was just

3:12

it felt almost an insurmountable wall of

3:15

pull requests that just kept coming in

3:17

and I had no idea whether or not the

3:20

performance of the app would be

3:21

regressing,

3:22

right?

3:24

So, a lot of my early time on the Cursor

3:27

team was spent looking at uh Chrome

3:31

DevTools and the performance and doing

3:33

performance traces and taking heap

3:36

snapshots.

3:37

Uh and it was extremely extremely

3:40

manual.

3:41

It was it It was so manual it it got to

3:44

the point where um I just got really

3:47

frustrated and started to think about,

3:49

you know, wait, we have agents, what am

3:51

I doing?

3:52

And so, I started thinking about

3:54

verification skills. Where you know,

3:57

what if my agent could actually run the

3:59

application itself and take the traces

4:02

for me, understand the traces, and find

4:04

the hot spots, and basically hill climb,

4:07

you know, to uh better

4:10

uh performance in our application

4:12

automatically.

4:14

And throughout the last 6 months uh of

4:17

being at Cursor and SpaceX AI,

4:20

you can really see that my productivity

4:21

has skyrocketed.

4:24

And, you know, I never set out to ship

4:27

2,000 pull requests a month. That was

4:29

not a goal of mine at all.

4:32

But, I realized that all of the skills,

4:35

all of the tools, and and code base

4:37

changes I were making

4:39

laddered up to this idea of trust. You

4:42

know, I didn't know it at the time, but

4:44

um I had this

4:46

thought in the back of my head head,

4:47

which was, you know, I am the

4:49

bottleneck,

4:51

and I need to be able to take all of the

4:53

knowledge that I have as an engineer and

4:56

impart them into into my team of agents

4:59

so that I didn't need to

5:01

uh be the blocker for everything. And um

5:06

you can clearly see that it's it's paid

5:09

off.

5:11

So,

5:13

I think it really comes down to trust.

5:17

Um but, how exactly do you build that

5:19

trust? And where do you start if you

5:22

are, you know, wherever you are in your

5:25

journey of using agents?

5:27

So, uh when I started, obviously, I was

5:30

in this category, you know, the one to

5:32

one to five range,

5:34

where

5:36

um you still feel like you have to baby

5:38

sit every chat uh and and conversation,

5:42

and you're just constantly

5:44

course correcting, you know, you're

5:46

intervening, you're correcting your

5:48

agent so that it does the right thing.

5:51

Um and if you're not there, basically

5:53

nothing get nothing happens and the

5:56

agents do the wrong thing.

5:59

And I would actually argue that this

6:00

part this phase of, you know, using

6:03

agents is actually the hardest to get

6:04

out of

6:06

because it's not always very clear how

6:08

exactly you get out of it.

6:11

And um again, it just comes down to

6:13

trust.

6:14

Um you know, the reason you're unable to

6:17

go from one one or one to five agents to

6:21

something more like 100

6:23

um is because you don't have trust in

6:25

your agents work yet.

6:26

So, if you don't have trust and you try

6:28

to spawn 100 sub agents or cloud agents

6:31

you're going to quickly find that you're

6:33

just going to get a ton of slot pull

6:35

requests and you know, a bunch of

6:37

regressions and a bunch of bugs shipped

6:39

and no one's going to be very happy with

6:41

that.

6:44

So, that begs the question

6:47

how do you trust your agents more?

6:50

So, for me it came uh it it really

6:55

started when I uh like I said, I meant I

6:58

joined the cursor team and I was

7:00

starting to work on performance where I

7:02

realized the need for verification.

7:05

So, when I say verification uh there are

7:09

sort of levels to that.

7:11

Uh because um you know, on the I guess

7:14

the lower end of the scale for

7:16

verification, you have things like um

7:19

the verification skills that I talk

7:20

about where you teach your agent how to

7:23

run your application

7:25

um you know, and use things like the

7:27

Chrome DevTools protocol or whatever

7:30

other protocol that you have for

7:31

debugging and teach them how to um

7:36

you know, debug the application, take

7:38

performance traces um

7:41

take heap snapshots and so on.

7:43

Um and then on the opposite end of

7:46

spectrum of the spectrum, which is much

7:47

much harder and still very much an open

7:50

question,

7:51

is like more

7:53

formal verification, where, you know,

7:55

maybe you rely on uh formal methods or,

7:59

you know, languages like Lean or TLA+

8:03

uh to do, you know, verification so that

8:05

you can check that, you know, your

8:08

business level or business logic uh

8:11

invariants are, you know, are always um

8:15

true.

8:16

Um and that um

8:19

uh you can formally verify that your

8:21

application is always in a correct

8:23

state.

8:24

Uh but I would say that, you know, even

8:27

if you don't have um the ability to run

8:29

formal methods, and very few people

8:32

really do,

8:33

um that with verification skills, you

8:35

can get very far.

8:37

So, when I joined Cursor, uh and started

8:40

to work on the Cursor agent window, the

8:43

first skill that I built was the skill

8:44

called control glass,

8:46

which is a verification skill that

8:49

teaches the agent how to

8:51

um run the application and take traces,

8:54

like I mentioned,

8:55

and it does that through

8:57

uh the Chrome DevTools protocol.

9:01

Um something interesting about that

9:03

skill

9:04

is uh and I kind of iterated my way to

9:07

this, uh it didn't start out this way,

9:09

um but the control

9:12

uh the control or verification skill

9:15

really has two components to it.

9:18

Um

9:19

the first part is obviously a CLI.

9:22

So, you want your agent to be able to

9:25

reproducibly

9:27

um be able to run the application and

9:30

collect traces and collect evidence

9:32

that, you know, empirical evidence that,

9:35

you know, code is working,

9:37

uh that your performance bar is being

9:39

met. Um and rather than have your agents

9:42

create scripts every time,

9:45

um you know, that can differ between

9:47

agent sessions, uh you can actually in

9:50

you can actually create a CLI that's

9:52

within the skill directory, and then

9:55

your agents will just use that every

9:57

time.

9:58

And then,

9:59

of course, you need to actually invest

10:01

in it and make it good so that it can

10:03

handle all sorts of different use cases,

10:06

um

10:07

and um

10:09

be able to

10:11

uh you know, run the application

10:13

correctly.

10:15

Uh another thing that's also really

10:17

important is uh this idea of a feature

10:20

map.

10:21

So, a feature map really is um something

10:25

that I kind of coined, I guess, where uh

10:29

you know, we started using these control

10:31

skills with Incursor.

10:33

Then we quickly realized that, you know,

10:35

a Slack report would come in and a user

10:37

would post a very vague screenshot,

10:40

right? Like a very small slice of the

10:42

UI, and just like three question marks.

10:45

And the agents using our control skills

10:47

had just no idea. Like it could run the

10:49

application, but it would just be

10:50

guessing at what exactly the the user

10:53

meant.

10:54

And so, I had this idea to create

10:55

something called a feature map,

10:57

um which is

11:00

I guess kind of inspired by like a site

11:01

map.

11:02

Um it essentially is a

11:06

uh form of materialized memory, you

11:08

know, like how exactly does your

11:10

application work? What features does it

11:12

have? How do it does a user reach it? Um

11:15

you know, in terms of like keyboard

11:17

shortcuts or what DOM elements to click

11:19

on, uh you know, things like that, and

11:22

what all the different features do.

11:25

Um and uh they are these this feature

11:29

map is stored in the skill itself,

11:32

um in the code base as part of the skill

11:34

directory.

11:35

And we have an automation that uh

11:37

maintains this feature map as well.

11:41

But when we combine the CLI and the

11:42

feature map, we quickly realized that um

11:45

this combination was very, very powerful

11:47

because now agents could not only,

11:50

you know, reproducibly control the

11:52

application and take traces,

11:54

but it could also understand

11:56

uh requests that came in from internal

11:59

users as well as external users.

12:02

And so, we very quickly realized that

12:05

the control verification skills

12:08

were so useful that they've become more

12:11

or less

12:12

uh critical infrastructure for our team,

12:15

and we constantly uh maintain it.

12:18

But the ability for an agent to verify

12:20

its own work is extremely powerful,

12:23

and uh we have um

12:27

spent a lot of the time on the skill,

12:29

and it's very, very powerful for

12:30

building that trust.

12:33

Um in addition to verification, of

12:35

course,

12:36

uh you want uh because ver- uh

12:39

verification is really about

12:41

correctness.

12:42

Correctness to me is really about does

12:45

the thing does the feature or the code

12:48

do the thing that you want it to do?

12:50

Right? Like does the you know, the the

12:52

checkout button, does it actually check

12:54

out the cart?

12:56

Um verification's really important for

12:58

that because you can get empirical

13:00

evidence that the feature actually

13:02

works.

13:03

But it doesn't tell you much about, you

13:06

know, the performance or

13:08

uh the code quality of that feature.

13:11

Um and that's where you start thinking

13:13

about skills that teach agents to work

13:15

like real software engineers.

13:18

So,

13:20

um I've built a plugin called Pystack.

13:24

Um I'm not going to talk about the

13:25

plugin too much today.

13:27

But

13:28

um a lot of the inspiration for that

13:30

plug-in,

13:31

uh which is a collection of skills that

13:33

I've created,

13:34

are inspired by the kind of workflows

13:38

that I personally have used as my in my

13:41

time doing software engineering

13:44

for all sorts of different types of

13:45

tasks. So, debugging, you know,

13:49

uh feature development, prototyping.

13:52

There's a whole bunch of different

13:53

playbooks and skills that ship in Piece

13:56

That that teach your agents how to write

13:59

code the way that you want them to.

14:01

And you know, this is where

14:04

uh you know, the more experienced

14:05

engineers on your team can really

14:07

contribute

14:08

to set up a team repository of skills

14:12

that um just make your agents a lot

14:15

smarter.

14:16

Um

14:18

and when you combine those skills with

14:20

uh verification skills,

14:23

then you you're able to get to a point

14:25

where your agents are able to not just

14:27

verify that the work that they're doing

14:29

is correct, but that it's also high

14:31

quality.

14:32

And again, because of verification, you

14:35

can collect real performance metrics.

14:39

You can get real numbers and statistics

14:41

and telemetry

14:42

on uh the like performance of the

14:46

application.

14:49

Um so, I think that's a really important

14:51

part to invest in.

14:54

Um

14:56

Another thing I think is really

14:57

important is refactoring and rewriting

14:59

your architecture to be more

15:02

agent-friendly.

15:03

Um and I almost want to say that

15:07

uh this is one of the most important

15:08

things you can do

15:10

um as a as a software engineering team

15:13

because um if you really truly believe

15:16

that agents are going to be writing all

15:17

the code in the future,

15:19

then we need to design our code bases so

15:22

that they do the right thing by default.

15:24

>> Mhm.

15:25

And you'll find that there's this

15:27

uh I I almost want to say like a

15:29

a scale or continuum between

15:34

uh you know, these different pieces of

15:36

building trust in your agents

15:38

where

15:39

um on you know, I have like five of

15:41

these points here where

15:44

uh

15:45

the code base is really like the best

15:48

form of memory

15:49

because agents love to extend existing

15:52

patterns that they see.

15:54

Um and you know, uh I think this is just

15:57

the nature of how LLMs work where uh

16:00

they're more likely to use whatever it

16:03

is in their context window um to make

16:07

changes. And of course, the files that

16:10

the agents read and opened are part of

16:13

its context window.

16:15

Um and therefore, the code base is a

16:17

really important part of that.

16:19

Um because you know, agents aren't going

16:21

to just refactor your code every single

16:23

in every single PR. They're going to

16:25

just look at what's already there and

16:27

just extend. The next level, I think, is

16:30

about static analysis

16:32

where you have linters, you have

16:34

compiler diagnostics, you have

16:37

uh continuous integration

16:39

um

16:40

and these are

16:42

guidelines and constraints that you can

16:45

enforce in your code base so that

16:48

whenever you correct your agent, you

16:49

find that, you know, they just keep

16:52

making the same mistake.

16:54

Um you can add those as lint rules or

16:57

even better, you can refactor your code

16:59

base so that the mistake that the

17:01

agent's making becomes categorically

17:03

impossible.

17:05

Um and then a step above that, right?

17:07

Where uh where and this is where we're

17:09

starting to get more into

17:11

uh less of like a hard constraint and

17:14

enforcement and more into the realm of

17:16

guidance.

17:18

You have things like rules, you have bug

17:19

bot, you have skills,

17:22

um which, you know, your agents will

17:24

obviously sometimes and mostly use uh

17:27

when they're doing their work, but

17:29

there's also a chance that it might, for

17:31

various reasons, you know, forget to

17:33

read a rule or maybe the user that is

17:37

piloting the agent ignores them.

17:40

Uh so, these aren't quite as

17:41

enforceable, but also an important part

17:44

of

17:45

um

17:47

setting up your environment so that you

17:48

can really trust what your agents are

17:50

doing.

17:52

Uh and then finally, you have a style

17:54

guide, which is really only enforceable

17:57

by humans in a code review. I guess you

17:59

could put the put these in your rules

18:01

and and bug bot and skills as well.

18:04

But if you don't, then you have this big

18:06

glaring hole in your

18:09

uh review process, where now humans have

18:12

to, you know, look at every single line

18:14

that's being changed and remember the

18:16

comment and

18:18

with the rate of pull request that are

18:19

coming in, it just it just becomes

18:21

impossible.

18:23

So, I definitely wouldn't recommend, you

18:24

know, relying only on the style guide. I

18:27

think the style guide or, you know, like

18:30

looking at human reviews is a good place

18:31

to start

18:33

um in terms of what's missing,

18:35

but you should really invest the time to

18:38

think about the other four parts,

18:41

uh you know, your code base, making

18:43

things categorically impossible to

18:45

better

18:46

data structures or algorithms,

18:49

um static analysis, and then of course,

18:51

you layer that with rules and bug bot

18:53

and skills.

18:56

Um on the code base front,

18:59

um

19:01

and in the Grokbot code base,

19:04

we actually have invested into

19:07

setting something up that we call

19:09

Dune, which is our agent-friendly

19:11

framework.

19:12

So,

19:14

um the inspiration for Dune really came

19:16

about from a lot of the performance

19:18

issues we were seeing we were seeing in

19:20

um the cursor agent window.

19:23

And so, a lot of lessons came out of

19:25

that exploration,

19:27

but um

19:29

the key principle that we landed on is

19:31

really that agents love taking

19:33

shortcuts.

19:35

So, what if we designed a framework such

19:39

that the shortcut, you know, the easy

19:41

path is the right path for agents. And

19:45

also one that would be a code base that

19:47

is, you know, maybe pretty annoying for

19:50

humans to work in uh because it's so

19:52

locked down in terms of what you can do

19:54

and what you can't do,

19:56

but uh it actually creates the perfect

19:58

environment for agents, especially ones

20:01

that have very minimal context because,

20:03

you know, not every contributor to your

20:05

code base is going to be an engineer

20:07

anymore. You can have designers, you can

20:09

have product product managers, you can

20:11

have CEOs, you know, going into the code

20:14

base and and shipping features. So, we

20:16

want to really think a lot about how we

20:19

invest and set up our code bases so

20:22

that, you know, even agents that are

20:24

piloted by busy people with not a lot of

20:27

context can do a good job by default.

20:31

And like I mentioned before,

20:34

um

20:35

your code base is really a form of

20:37

memory for agents because they love to

20:39

extend the existing patterns that they

20:42

see.

20:43

And the reverse is actually also true,

20:45

right? You can invest the time to set up

20:47

your code base um in a way that things

20:51

are, you know, bad patterns are

20:52

categorically impossible, or you have

20:55

like lint rules that prevent them,

20:57

but the reverse is also true in the

20:59

sense that um if you have existing

21:02

anti-patterns, you'll actually find that

21:05

these will spread kind of like a virus.

21:08

Um

21:10

where you have like one small workaround

21:13

or a comment that explains a workaround,

21:16

and you'll quickly find that agents just

21:18

love to copy that. And then in a matter

21:21

of a few days or a few weeks, you'll

21:23

find that the workaround has spread

21:25

everywhere and it's now becomes it has

21:28

become like a de facto pattern for all

21:31

agents. And that's a really, really bad

21:33

place to be to be in.

21:35

And that goes back to what I was saying

21:36

about why it's really important, you

21:38

know, whenever you're correcting your

21:40

your agents,

21:42

that you invest the time into thinking

21:43

about codebase changes and static

21:46

analysis to and and of course the

21:49

layering them with rules, good rules,

21:52

and bug bot and skills

21:54

uh because of of that reason.

21:58

Um and

22:00

another analogy that I like in addition

22:02

to the Michelin kitchen is this idea

22:04

that, you know, your codebase is kind of

22:06

like a garden

22:07

where you have

22:09

uh

22:09

workarounds uh you know, that seem

22:12

seemingly that seem kind of innocent at

22:14

first, but then because of the nature of

22:17

agents, you just copy that pattern over

22:19

and over again,

22:20

and you quickly end up with a very, you

22:23

know, vibe coded codebase that is a you

22:27

know, a pain in the butt to maintain and

22:30

has a lot of performance issues.

22:32

So, in my opinion, the perfect agent

22:35

codebase is one that's so locked down

22:38

that, you know, it's again it's like

22:39

really annoying for humans to to write

22:41

code in,

22:43

but it's so conventional, it's so

22:45

standardized that

22:48

uh you know, even innocent-looking

22:49

patterns are just forbidden.

22:52

Um and the best example of this I have

22:55

is actually something that seems very,

22:57

very innocent when you look at it,

22:59

but when you think about it, it's

23:01

actually really bad.

23:02

And that pattern is

23:04

uh agents leaving comments in the code.

23:08

Now, you know, when I first saw agents

23:10

starting to do this, I

23:13

initially wasn't Well, I I did think

23:15

that a lot of them were slop, but I also

23:18

thought that, you know, it actually

23:19

doesn't It's It's not a bad thing,

23:20

right? I guess if agents are leaving

23:22

comments in the code because

23:24

as humans, we left

23:26

comments in the code whenever we saw,

23:28

you know, edge cases or we needed to

23:30

actually make a workaround or, you know,

23:34

leave a note to ourselves or a colleague

23:37

on a particularly tricky part of the

23:39

code base.

23:40

But,

23:42

what I quickly realized when we saw this

23:44

happening in the Cursor code base

23:47

was that agents were just using the

23:49

comments

23:50

around the code as justification for why

23:54

it wasn't going to solve the actual

23:56

problem.

23:57

And instead paper over it with a

23:59

band-aid or

24:01

a short-term solution.

24:03

So, in Dune, which is again the

24:06

framework that powers Grok Bot,

24:09

we made the choice to

24:12

um

24:13

we made the choice to actually ban

24:15

comments

24:16

before that reason where so that agents

24:19

would not just copy

24:20

that pattern and uh

24:23

you know, propagate it everywhere in the

24:25

code base.

24:28

So, uh

24:30

my pitch here is that every team really

24:32

needs something that, you know, like a

24:34

role that I'm calling a gardener.

24:37

Um in the same way that uh you know,

24:40

with a real garden, you need someone who

24:42

is thinking a lot about

24:44

the

24:45

uh you know, things that can kind of

24:47

creep in and grow

24:49

in ways that you don't want. Like, you

24:51

know, you have weeds, you have

24:54

uh you know, just other types of organic

24:56

growth.

24:58

I don't really know much about gardening

24:59

that well. Uh I

25:00

But you have things that you know

25:02

unwanted pests and and and stuff like

25:04

that that kind of creep in into your

25:06

code base.

25:07

And so you want to nip them in the bud

25:09

as soon as possible before they start

25:11

propagating everywhere.

25:14

A lot of the principles behind Dune are

25:17

really centered around these three

25:19

things.

25:20

First of all, we want to delete tech

25:22

debt that we already have for you know

25:24

for reasons I just mentioned.

25:26

We want to

25:28

keep or enforce a single paved path for

25:32

most blessed patterns.

25:34

You know there should be one

25:36

conventional way to do some things.

25:38

So that agents don't really need to

25:40

guess and there should be enough

25:42

guidance in the code base in CI

25:46

in lint rules

25:48

so that the agents are guided to do to

25:50

to follow that path.

25:53

And then finally whenever you see tech

25:55

debt or bad patterns, your instinct

25:58

should be I need to write a lint rule

26:00

against it.

26:01

You don't always have to clean up

26:03

immediately

26:04

because if you write a lint rule, you

26:06

can at least stop the bleeding.

26:08

And which you know doesn't solve the

26:10

problem entirely, but it at least

26:11

prevents it from growing.

26:13

So I definitely recommend you know

26:15

really thinking a lot about

26:17

how you can guard against anti-patterns

26:21

so that they don't spread like a virus

26:24

and then also spend time to actually you

26:27

know get your agents to clean them up so

26:30

that your code base is just constantly

26:31

kept in a state where

26:35

you would be happy if an agent would

26:36

have copied it. That's the kind of

26:38

mindset that I would recommend having.

26:42

And then I won't actually go through all

26:44

the details of Dune itself. But I'll

26:46

just kind of gloss through some

26:47

interesting parts.

26:49

So again as a reminder, Dune this

26:51

architecture the client framework that

26:53

we built to power GraphBot.

26:57

We've invested a lot into, you know, all

26:58

the things I was saying, where we have

27:01

conventions. We have a lot of

27:02

conventions about where code should

27:04

live.

27:05

And

27:06

um where and how

27:08

uh code should be imported between them.

27:11

So, in Dune applications, you know,

27:13

there's different concepts where like,

27:16

for example, features are all co-located

27:18

in a single folder.

27:19

Um you have an entry point um that's,

27:22

you know, in the React part of the code

27:24

that determines uh you can kind of think

27:26

of it like a route.

27:28

Uh you have transcript cards that show

27:30

up in the GraphBot application.

27:32

You have a host that runs on the

27:35

you know, the GraphBot virtual machine,

27:37

and then of course, you have your client

27:39

uh which uh powers the uh overall Dune

27:43

application.

27:44

And we have a lot of um strict

27:48

boundaries between these things, where

27:51

um

27:52

just an as an example, things that run

27:55

on the main process

27:57

uh or the main thread in Electron aren't

27:59

allowed to be run on the renderer

28:01

thread. And we keep that separation very

28:03

intentionally because uh of lessons we

28:06

learned from Cursor's agent window,

28:09

where we would sometimes see

28:11

code accidentally uh get imported into

28:15

the renderer thread, and you know, slow

28:17

code. And uh since

28:20

uh on the renderer thread, you want your

28:22

UI to be very smooth and

28:25

uh performant, uh you need to make sure

28:27

that you don't have any long tasks or,

28:30

you know, things that take longer than

28:32

16 milliseconds for if you want like 60

28:35

frames per second, or 8 milliseconds if

28:37

you want 120 frames per second.

28:40

And so, your renderer has to be

28:42

constantly in a state where

28:46

uh it is uh you can really kind of chunk

28:48

up the work um, and not do them all at

28:50

once.

28:52

Um, and so, we have code within Dune

28:55

that enforces this

28:58

these boundaries through the import uh,

29:00

independency graph.

29:02

Uh, but, yeah, this is just an example

29:04

of a pattern that we saw lead to really

29:06

bad performance that we um,

29:10

categorically eliminated through uh,

29:13

the architecture of the of Dune.

29:17

Um,

29:18

and then, all of these other pieces

29:20

aren't that interesting.

29:22

Uh,

29:23

but, again, the

29:25

the core theme here you know, it's not

29:27

about Dune but the idea that

29:30

um,

29:32

an agent-friendly framework of your own

29:34

is actually very, very powerful.

29:36

And you can encode all of the learnings

29:38

that you and your, you know, your best

29:41

engineers on your team have

29:43

tribal knowledge of. Um, and um, I think

29:47

the lesson here is that how do you take

29:49

that away

29:50

from, you know, what used to be in the

29:52

style guide process of reviewing code

29:55

and, you know, in engineering is

29:58

reviewing other engineers' work and

30:00

leaving comments

30:02

um, to

30:03

extract, almost like extracting that

30:05

knowledge and encoding that into the

30:07

framework

30:08

into the code base itself so that the

30:10

code base acts as the memory. Right?

30:13

It's the the thing, like you're coming

30:15

back to this idea that, you know, the

30:17

code base is just the thing that it's

30:19

the

30:20

the materialized snapshot of the state

30:23

in which you want your agents to extend.

30:26

And you want that code base to be so

30:28

pristine, so great that the next agent

30:31

that comes along is just very likely to

30:32

continue that pattern and keep it

30:35

really, really good.

30:37

Um, and if you spend enough time on this

30:40

process

30:41

uh, like I mentioned, you can really set

30:43

up um, a Michelin kitchen or a software

30:46

factory

30:48

where

30:49

um, because you spent so much time on

30:54

you know, all of these pieces that allow

30:55

you to trust your agent,

30:57

whether it's in the code base, whether

30:59

it's lint rules,

31:01

um, whether it is

31:04

uh, you know, diagnostics or rules or

31:06

bug bar or skills,

31:09

these layers come together

31:12

um, and provide you a lot of trust.

31:14

Because now, you know, just imagine for

31:16

a moment, you're working in the GrokBot

31:18

code base.

31:19

Um,

31:21

it's super locked down, you know, it's

31:22

like almost impossible to write bad

31:25

code. So, you know, you can even an

31:27

agent with very little context,

31:30

uh, you know, even a an agent with not a

31:32

lot of reasoning can come in and and

31:34

actually write code that's good.

31:37

And going back to my example about the

31:39

Michelin factory, I think there's

31:41

a lot here, right? Where, you know,

31:43

we're setting up our agent, our bot with

31:46

skills and tools, you know, we're

31:48

training them, uh, we're setting up our

31:50

kitchen in a way that makes sense,

31:53

right? For

31:55

the agents and bots to do the right

31:57

thing by default.

31:59

Uh, you know, whenever we see, for

32:00

example, in the in the kitchen example,

32:02

if we notice that uh, one of our cooks

32:05

or dishwashers is constantly tripping

32:07

over something, of course we need to fix

32:09

that. Right? We need to problem solve

32:11

and ensure that, you know, others don't

32:13

trip as well because, you know, in a

32:15

kitchen is a very dangerous place and

32:16

you don't want to hurt yourself.

32:18

It's the same

32:20

mindset, I think, that we should have

32:22

with our code bases. How do we

32:25

set it up so that even agents, uh,

32:27

without a lot of

32:29

uh, knowledge can and do a good job.

32:33

Um, and I think with

32:36

uh, you know, GrokBot,

32:37

GrokBot and Cursor play an interesting

32:39

role together where GrokBot is really

32:43

great at uh providing what I call the

32:46

outer loop because you can connect

32:48

GrokBot to lots of different uh you

32:51

know, different connectors like Slack to

32:54

DataDog, Sentry,

32:56

uh PlanetScale, whatever services that

32:58

you use. And you can aggregate all of

33:01

that information together and use that

33:02

to make really good decisions for

33:04

itself.

33:06

Uh

33:07

Some people call this like a company

33:09

brain. Uh I don't really think I

33:11

personally don't think you need anything

33:13

that sophisticated here because agents

33:15

are really good at using tools. And so,

33:18

if you connect these tools to GrokBot

33:21

um and you start having your GrokBots

33:24

auto kick off things like cloud agents,

33:28

you can actually find that it's really

33:29

not uh

33:32

you don't really have to invest in a lot

33:34

of infrastructure to build a software

33:36

factory.

33:37

Um in fact, I'm going to you know, cross

33:40

cross out this

33:41

this term cuz I don't like this term.

33:44

Um uh I think you can set up this

33:47

personal mission control kitchen for

33:49

yourself through GrokBot, things like

33:51

GrokBot routines which let you subscribe

33:54

to you know, Slack threads to Sentry

33:57

alerts that let you kick off things

33:59

automatically.

34:01

And when you combine all of these things

34:03

that I've been mentioning, you know,

34:04

your code base, your rules, your skills,

34:08

uh they all compound and GrokBot uh will

34:13

be able to you know, automatically

34:14

respond to events that come from the

34:18

outer loop and then kick off cloud

34:20

agents. Um and you can also set up

34:23

Cursor automations and use our SDK uh to

34:27

set up um bot additional bots as well

34:30

that reuse a lot of these

34:33

pieces of agent infra that you've set up

34:36

um and allow them to do much more

34:38

complicated tasks.

34:41

So, if you if you if you've done all

34:44

this, then I think you can get to a

34:45

point where

34:46

uh you know, uh I have some screenshots

34:49

here of some of our automations and our

34:51

agents in

34:53

uh that that work on cursor

34:56

where we are automatically re-

34:59

reproducing bug reports, we're

35:01

automatically opening pull requests, we

35:04

are um essentially adding a lot of value

35:08

to the entire team

35:10

uh because all of these things compound.

35:15

So,

35:16

if we kind of zoom out again,

35:18

um and go back to this graph,

35:24

I think that to kind of close off the

35:26

talk,

35:27

um if you spend a lot of time thinking

35:30

about all of the pieces that you need uh

35:32

to be able to ascend the trust graph,

35:35

uh you start getting to a place where

35:37

you can

35:39

uh really trust your agents more and

35:41

parallelize your work, and also empower

35:44

your entire team to build on top of

35:46

these uh pieces of infrastructure for

35:50

your agents, um and empower everyone,

35:54

you know, every engineer on your team,

35:55

every builder,

35:56

to be extremely productive and be able

35:59

to write high-quality code.

36:01

So, the last thing I want to leave you

36:03

with is actually this piece.

36:06

Uh sorry, not that piece, but this

36:08

piece.

36:09

Um I think if there's only one thing you

36:10

take away from my talk, it should be

36:12

this this slide here,

36:14

uh which is, you know, that these are

36:16

the activities that will help you build

36:18

up towards a high-trust environment. You

36:22

know,

36:23

whenever you find yourself correcting

36:25

and intervening your agent,

36:27

um you really want to think about it

36:29

from these five pieces and where is the

36:33

most effective step in this

36:37

sequence

36:39

in order to

36:41

get make your agent you know much more

36:43

trustworthy

36:45

and of course I definitely recommend

36:47

thinking about thinking about it in this

36:49

order where you know you

36:52

either invest the time to make that

36:54

pattern categorically impossible through

36:57

your code base and architecture and data

37:00

structures or you start looking at

37:02

things like static analysis and then you

37:04

layer that on with rules and bug bot and

37:07

skills.

37:09

If you do all of that and you also spend

37:11

some time you know thinking about your

37:14

code quality

37:16

in in terms of of skills

37:18

you get to a place where you trust the

37:22

you trust the environment so much that

37:25

your agents can just be free

37:27

right and personally I have spent a lot

37:31

of time for this for growth bots code

37:33

base for example and

37:35

this is really the secret right well

37:38

it's not really a secret it's it's a lot

37:39

of hard work but I hope you found this

37:42

talk useful

37:45

and please reach out to me on X my

37:48

handle is potato with an e

37:51

and

37:52

I hope that you'll have

37:55

a lot of fun and successfully in your

37:57

own mission in kitchen.

37:59

Thanks for watching.

Interactive Summary

Lauren, from SpaceX AI and the 'Grok Bot' team, discusses how to build a high-trust environment for AI agents to achieve high-quality, high-velocity software engineering. She compares the setup to a Michelin-starred kitchen rather than an assembly line, emphasizing the need to treat the codebase as a form of memory for agents. She outlines a hierarchy for building agent trust: starting with architectural constraints and static analysis, followed by rules, 'bug bots', and specialized skills, ultimately aiming for a system where agents can operate autonomously and reliably.

Suggested questions

4 ready-made prompts