HomeVideos

Lauren Tan XAi Grokbot

Now Playing

Lauren Tan XAi Grokbot

Transcript

1474 segments

0:00

Everyone, I am Lauren Lauren Tan. I

0:03

guess not many people know my last name.

0:05

Uh, I am Potato on Twitter. Uh, potato

0:10

with spelled with an E. Um, and I have

0:13

been at Cursor for about 5 months. Uh,

0:18

previously I was at Meta where I worked

0:20

on the React team. Uh, specifically

0:22

working on the React compiler, uh, which

0:25

was a whole lot of fun. Uh I'm still on

0:27

the on the core team and and uh

0:29

contributing to open source here and

0:30

there. Uh so that that's really nice

0:33

that they still let me do that. Uh and

0:36

before Meta, I was at Netflix uh where I

0:38

was uh both a tech lead and uh I

0:43

transitioned to be be an engineering

0:45

manager um for about two years. So I've

0:50

had a I've have had a lot of experience

0:52

going between engineering management and

0:55

being an individual contributor. Uh and

0:58

I think something I've noticed actually

0:59

which is quite interesting is that there

1:02

are so many parallels with you know

1:04

management skills and how to like manage

1:06

agents. Uh and that's actually a big

1:09

part about what I wanted to chat with

1:10

you and everybody else about today. Um

1:14

but yeah that's that's me. Uh I do have

1:17

some like very light slides but uh it's

1:20

not going to be uh just rambling. So let

1:23

me just share my screen

1:25

and hope that I don't leak anything. Uh

1:33

oh no I need to allow permissions.

1:36

>> No worries. Take your time. There's

1:38

always tech tech trouble. Uh give me one

1:41

second to rejoin.

1:42

>> Yeah, go for it.

1:46

I see many of you already know Lauren

1:48

from from the looks of the chat here.

1:50

Um, so yeah, it's exciting to to get a

1:53

chance to chat with her uh and and go

1:55

through some of her her recent work. As

1:57

you guys heard, you know, a lot of

1:58

recent experience from from Netflix to

2:01

to Meta and then now over at Cursor. Uh,

2:04

we're going to chat a little bit about

2:05

Grockbot as well. So, uh, that'll be

2:08

exciting. I don't know if you guys saw

2:09

that was a recent release. I think

2:10

literally maybe yesterday or the day

2:12

before uh from the cursor team which is

2:14

kind of like um let's call it like

2:16

agents for everyone. You can go check it

2:18

out if you want and learn a little bit

2:19

more about the product but um but yeah

2:22

we'll we'll explore that a little bit

2:23

today as well.

2:26

All righty. Welcome back.

2:40

And Lauren, you're just on mute there if

2:42

uh you want to hop off of mute if you're

2:43

chatting.

2:44

>> Yeah, sorry.

2:45

>> No worries.

2:47

>> It's 2026 and I still don't know how to

2:49

use Zoom. [laughter]

2:50

>> That's all good.

2:51

>> Uh okay, so I assume you can see my

2:53

screen.

2:54

>> Yes. Yeah, we're good. Um, so yeah,

2:56

today, yeah, I think I think the big

2:58

theme for me as I've been using agents

3:01

to write code, and I'm sure a lot of you

3:03

have had the same experience as well, is

3:06

how do you trust it? You know,

3:07

especially if you are an engineer that's

3:10

been writing code for a very long time,

3:12

you have a lot of opinions and lessons

3:15

that you've learned about doing good

3:17

engineering. And when you see agents

3:20

just, you know, winging it and, you

3:22

know, guessing, hallucinating,

3:24

uh, you know, confidently stating that

3:27

they found the smoking gun, uh, for the

3:29

hundth time, uh, but it's actually not

3:32

the real problem. You lose a lot of

3:34

trust. And when you lose when you don't

3:36

have much trust in your agents,

3:39

I feel like you you really can't get the

3:41

most out of them. And for me, the

3:43

parallel is like with management. Uh so

3:46

if I'm an a manager an engineering

3:48

manager of a team and I have a bunch of

3:51

you know I have a team of engineers uh

3:54

on my team and I don't trust them then

3:57

the mode of operation I'm going to be in

3:59

is going to be like micromanagement

4:00

right I'll have to spend a lot of time

4:03

looking over my reports shoulders and

4:06

checking that they're doing their work

4:08

well you know that they're not shipping

4:09

bugs to production

4:13

and so I drew this chart cuz uh it's not

4:16

it's not a very scientific chart but

4:18

like this is how I imagine myself and my

4:22

journey through using agents. So you

4:25

know like fast forward or back forward

4:29

uh or fast back uh fast backwards like a

4:33

year or so when you know nobody was or

4:35

not many people were using agents to

4:37

code. Uh I think you

4:41

uh you know get into this mode where you

4:43

are

4:45

in very heavily in the loop with one or

4:48

several like a handful of agents and you

4:51

find yourself just constantly fig uh you

4:53

know trying to understand what your

4:55

agents are doing uh and you're very very

4:57

in the loop. You're watching every

4:59

single output. you are sitting there

5:01

prompting um and you really can't

5:04

parallelize beyond that because you

5:07

don't again you don't have that trust

5:09

right you can't go to a 100 agents uh

5:11

like spawn 100 agents when you don't

5:13

even trust the output of one agent

5:17

so over the past 5 months I feel like

5:20

I've really been able to uh like ascend

5:23

this trust curve and now I'm at the

5:26

point where uh I actually have this

5:29

Sounds kind of scary to say this and it

5:31

it makes me sound like a slop artist,

5:33

but I I promise I'm not, but I actually

5:35

have my agents now um automerging PRs

5:38

for me. Uh which is like a wild thing to

5:41

say, but um like I woke up today and

5:44

there were like 20 PRs landed and I just

5:47

reviewed them on Maine like they were

5:49

already landed and they were good. Uh so

5:52

how did I get to that point is basically

5:55

what I wanted to talk about today.

5:58

Uh, and again like yeah, feel free to

6:00

jump in if you have questions, Colin.

6:02

Um, but uh, oh yeah, of course I got to

6:06

show this this chart. Uh, where uh,

6:12

no do not trust to someone requested to

6:16

control my computer. Uh, probably won't

6:18

do that. Uh but yeah, so this chart I

6:22

think I I'm sharing this chart not to

6:24

kind of like flex but to kind of show

6:27

like the journey like so you can see

6:29

like the curve like it sort of like

6:31

inversely matches the contributions I've

6:34

been able to land at cursor. So I joined

6:37

five months ago and five months ago like

6:39

I you know my first month I was like not

6:41

very productive because I was ob you

6:43

know I was learning the codebase didn't

6:45

know what the heck was going on and as I

6:48

got more confident in in my agents uh

6:51

I've really been able to kind of ramp up

6:53

my productivity. Uh and again like yeah

6:56

like last month I shipped a thousand PRs

7:00

which is ridiculous. Uh, and then this

7:02

month we're only on the 12th. I'm

7:04

already at like almost 800 PRs landed.

7:09

Uh, so the velocity is definitely high

7:12

and you you I'm sure a lot of you will

7:14

definitely be questioning like how how

7:16

much of this code is actually good. Um,

7:18

and I think yeah, like that's definitely

7:21

fair to question.

7:23

Um, but uh, yeah, I think I think if you

7:28

set up your agents well, you can

7:30

definitely get to a very similar level.

7:34

Um, and so I'm going to talk about how

7:35

we do that.

7:38

Uh so for me I think I'm curious like I

7:42

guess call in your experience as well

7:44

but uh for me I think the most important

7:48

skill that you should have in your

7:51

toolbox when you work with agents is

7:53

verification.

7:55

Uh and by verification I mean the

7:57

ability for an agent to actually run the

8:00

code uh or take CPU traces or heap

8:05

snapshots or uh you know open an iOS

8:09

simulator whatever you know however your

8:12

application is exposed to your users it

8:15

can do the same thing and uh run it for

8:19

real and actually test and verify it

8:21

don't work because that's the thing that

8:23

really closes the loop. Uh it doesn't

8:26

guarantee your agent writes good code uh

8:29

but it allows them to at least write

8:31

correct code uh which is a big a really

8:34

big step forward for being able to trust

8:37

your agents. Um

8:40

I will I can share one example that we

8:43

have uh within cursor. Uh

8:48

oops

8:50

where let me open this.

9:00

Let's make this uh make me full screen.

9:03

There you go.

9:05

Uh so for the for cursor's agent window

9:09

uh so this is actually an interesting

9:11

story but uh when I joined cursor 5

9:14

months ago uh they're actually uh well I

9:18

was supposed to join a different team I

9:20

was supposed to join like the cloud

9:21

agents team uh but then since I have a

9:24

lot of experience working on react and

9:26

agents window is a react application uh

9:29

I was

9:31

uh I was asked to basically help out

9:34

with uh the agent window work. Uh but um

9:40

there wasn't really a lot of like skills

9:43

to help me. So I just found myself like

9:45

okay uh agent window is going to launch

9:47

in like a week, right? we have a really

9:49

tight deadline and um there was uh you

9:53

know I was just sitting there like okay

9:55

I'm going to open up the performant the

9:57

the Chrome dev tools and just like take

9:59

a trace look at it myself and try to

10:01

make sense of this flame graph and keep

10:03

in mind I was just like in my first week

10:05

so I had no idea what I was looking at

10:07

no idea what you know I mean I had some

10:09

idea but you know the codebase was

10:11

completely fresh to me uh and I realized

10:14

like my agent had no idea either you

10:17

know like I would take a screenshot of

10:19

the tree. I would download trades. I

10:20

would send it to it and it be like,

10:21

"Yeah, it kind of looks like this, you

10:23

know, uh, and it would like confidently

10:25

state like it's this thing." And then

10:27

I'd try to fix that. And turns out

10:29

that's not the actual thing.

10:32

So, this was very very slow process. And

10:34

if you've ever done any like performance

10:36

work yourself or, you know, just even

10:38

development with an agent where you

10:40

don't have a verification skill, you are

10:43

the verifier, right? You you're the

10:45

bottleneck. you you you tell your agent

10:47

to do something and then it goes off and

10:49

write some code. Then you open up your

10:51

you know local dev build and then you

10:53

start to say oh you know doesn't work.

10:54

Then you got to copy paste screen uh you

10:56

know screenshots or console errors or

10:59

whatever. Uh and then your agent like

11:01

slowly kind of like uh you know works

11:04

with that and then tries to understand

11:06

it and um fix the thing but then you're

11:10

constantly just in the loop and and

11:12

being a bottleneck. So there's really no

11:13

way to parallelize. So the control glass

11:16

skills are like one of the first skills

11:18

I built uh for cursor. Uh and glass by

11:21

the way is the code name for agents

11:23

window that we use internally but it's

11:26

just cursor I guess. Um and so this

11:30

skill uh is I guess the the the code

11:33

itself is not super interesting. Your

11:35

agent can very easily make one for you.

11:37

uh where if if you're building an

11:39

Electron app or a web app or even iOS uh

11:43

applications uh you can teach your agent

11:46

how to use like the ChromeDev tools

11:48

protocol or through uh Apple has some

11:53

utilities as well for running the

11:54

simulator and taking traces and

11:56

controlling programmatic control as

11:59

well. Uh so that's really useful. Uh but

12:03

one thing I want to talk about is uh the

12:08

this thing

12:10

uh where is the read me? Uh so this

12:14

skill comes with this very unique

12:16

feature called or not feature uh unique

12:19

file called a feature map. And so the

12:23

story then is like I built this skill

12:25

and so now the agent was able to uh

12:28

actually run the agent window and take

12:30

traces and whatnot. Uh but it had no

12:33

idea what what the agents window was. So

12:38

um you know like someone would say like

12:40

oh the the left sidebar is like laggy or

12:44

something like that or you know the

12:46

right side the the PR tab is not working

12:48

and the agent would just be like kind of

12:50

flailing around. It would spend a lot of

12:51

time trying to like look up the code and

12:53

you know where is this feature? How do I

12:55

actually get to it on the UI which made

12:58

it basically completely useless. uh you

13:00

know like we would I would run this

13:03

skill locally and you know it would

13:05

spawn a dev build uh but then it' just

13:08

be churning like I just try to click

13:11

here it wouldn't know how to get to

13:13

things um and it was just an awful

13:15

experience uh so is putting arrows on my

13:20

screen um so uh yeah this this feature

13:24

map has been really useful uh because it

13:27

teaches the agent how to get to all of

13:29

the features that you have. Um, and in

13:32

PAC, the plugin that I I've made, uh, if

13:36

you search for PAC cursor on Google, you

13:39

you'll find it. Uh, but there is a

13:42

create verification skill in that plugin

13:44

where it actually helps you set up

13:46

something like this for yourself. Um,

13:49

including the feature map. So, it will

13:50

actually explore the code and build up

13:52

this initial feature map that tells your

13:54

agent how to get to all of the different

13:57

features that you have. Uh and this is

13:59

extremely powerful because now like you

14:02

have these user reports that come in uh

14:04

you you can actually map even like a

14:07

vague report or even a screenshot. So we

14:10

have this uh internally at cursor where

14:14

uh we have a slack channel with you know

14:16

lots of people giving us feedback on the

14:18

agents window and rockbot and whatnot.

14:21

Uh, and often times the report is very

14:24

bad like very low quality like someone

14:27

will just put very often we get like a

14:29

screenshot like and then someone just

14:30

says question mark question mark

14:32

question mark like what is this and you

14:34

know like without this your agent is

14:36

like I have no clue right but with a

14:39

feature map like this it has a lot more

14:42

context and understanding of how to

14:45

actually navigate how to get to all the

14:47

different features. Uh so like you know

14:49

example like I guess like the sidebar

14:51

like what is the sidebar uh you know

14:54

like all the different sub features that

14:56

are present in it um like from the user

14:59

point of view here's where how do I get

15:01

to it all the different keyboard

15:03

shortcuts

15:05

uh even like the the what do you call it

15:07

the DOM elements or yeah like the

15:10

attributes that you use for selecting

15:12

things through the CDP uh are all there.

15:16

So uh again yeah this is like really

15:19

really powerful uh for for agents

15:23

>> uh and pstack ships uh that create

15:26

verification skill but also a maintain

15:28

verification skill uh so you can keep

15:30

this up to date.

15:32

>> Cool. Yeah, I was just going to ask how

15:34

you created that. So do you mind sharing

15:35

a little bit more about um that that

15:37

process in the context of Pstack and

15:39

maybe just what Pstack is for the folks

15:40

who aren't familiar?

15:42

Yeah. So, P stack is pretty interesting

15:44

because uh well, first of all, the name

15:47

is kind of goofy. like the P the P and P

15:50

sack is like potato potato snack because

15:54

I um so uh uh there is a pretty uh

16:00

famous person Gary Tan who is the CEO of

16:02

Y Combinator and he's come up with this

16:05

plugin called Gstack uh Gary Stack and

16:10

uh funnily enough we share the last name

16:12

we have no relations uh but I thought it

16:15

would be funny to kind of you know poke

16:16

fun at Gary and make peace stack my

16:19

version of of of of [laughter] his

16:21

plugin. Uh but kind of tailor it to my

16:25

own set of p engineering practices.

16:29

Uh but I honestly actually never set out

16:31

to build PAC. Uh it just started with a

16:33

bunch of skills, right? Like I started

16:35

with that control glass skill and then I

16:37

started with another skill like called

16:39

how which I also noticed through like

16:42

observing agents. Um, so like you know

16:46

in the early days of me, you know,

16:47

trying to climb this ladder, I was like

16:49

super in the loop and I was basically

16:51

nitpicking my agents to an extreme

16:53

degree. I was like I would tell it um,

16:57

you know, this feature has stopped

16:59

working. Here's a bug report. Like why

17:01

isn't it working?

17:04

And very often the agent would just like

17:06

confidently state like, "Oh, it has to

17:08

be this, right? It has to be this

17:10

thing." And I noticed like when I looked

17:12

at the actual tool calls, I noticed it

17:15

wasn't actually reading the code that I

17:18

thought should be affected. And that

17:20

made me just extremely suspicious. And

17:22

at that point, I was like, I'm not going

17:23

to I can't trust any this agent anymore

17:26

cuz it's just it's just completely

17:27

hallucinating. And I think

17:31

I think it's very easy to just you know

17:33

like build up that distrust and not and

17:36

kind of feel helpless like you know you

17:38

don't know how to help your agents

17:40

succeed. But like again I think the the

17:43

the management analogy is super helpful

17:45

because like imagine if you were a

17:46

manager of an engineering team and you

17:49

had an engineer on your team who was a

17:52

really good coder. No business context

17:54

whatsoever. you know they they just you

17:56

just hired them and they they onboarded

17:58

you know like five seconds ago. Uh and

18:02

so how do you actually teach that person

18:03

to be effective? So how you do that is

18:06

through a skill. uh a skill being just

18:09

you know it's just markdown right but

18:11

you know it encodes a lot of information

18:13

instructions a lot of uh you can really

18:17

draw out a lot of intelligence from an

18:20

agent by well some people on Twitter

18:23

call it like you know pull the agent to

18:25

a different latent space which is kind

18:27

of like a fancy way of just saying like

18:29

since uh you LLMs are sort of like they

18:31

predict the next token uh when you give

18:34

it some high quality tokens uh to to

18:37

begin with then you know it it can kind

18:39

of pattern match on like a higher space

18:42

that's you know smarter

18:45

um so that's like a very interesting

18:48

model there but yeah I built Pstack very

18:50

very incrementally uh so uh started with

18:54

just really observing how agents you

18:56

know all the fail different failure

18:58

modes of of that agents were having and

19:01

every time I saw that I just okay I'm

19:02

just going to make that a skill right

19:04

like stop hallucinating actually go and

19:07

search up, look up the code, use a lot

19:09

of sub agents, uh, and yeah, stop

19:12

guessing.

19:16

>> Yeah, that makes sense. One, one kind of

19:18

followup question here, uh, both from

19:19

myself and from a bunch of people in the

19:20

chat. So,

19:22

>> I guess it's two two parts. So, one is

19:23

like how do you maintain these skills?

19:25

So, like the product changes over time.

19:27

Obviously, there's a lot of people who

19:28

are shipping against the codebase. So,

19:30

how do these skills get maintained? Uh

19:32

and then second to that is like how do

19:34

you know when your verification is is

19:35

good enough? Uh like and you know you

19:38

can trust that the ver verification

19:40

loops that you've built are going to I

19:42

guess you trust that the outputs uh when

19:44

they're done.

19:47

>> Uh yeah maybe I'll talk about um I think

19:50

there are some related maybe I'll start

19:52

with this one first. So like how do I

19:56

maintain these skills?

19:59

So, um, if you're not familiar with this

20:01

concept, an eval [clears throat] is

20:03

essentially like a way to, uh, well, I

20:06

the mental model I have is like it's

20:07

like a unit test for an agent. Um, and,

20:11

uh, you can actually make your own

20:13

evals. You don't need like a special

20:15

framework for them. You can build you

20:17

can you can build one depending on like,

20:20

you know, how scientific and how

20:22

rigorous you want to be. Uh, my screen

20:25

is red.

20:26

>> Yeah, there's a little button. Um,

20:28

sorry. Do

20:29

>> you mind like disabling the drawing or

20:31

something? I I can't see my screen.

20:33

>> Yeah, sorry. If you guys could not draw

20:35

on the screen, that'd be great. But, um,

20:36

there's a little button in the

20:38

>> troll.

20:39

>> Yeah, the little drop down.

20:42

>> How do I clear?

20:44

>> Yeah, you got it. Perfect.

20:49

>> Yeah. So, eval

20:52

test your skills basically. And actually

20:55

in Pstack we ship uh under potato mode

20:59

there's a playbook if you search for it

21:00

called eval playbook. Um and it's uh

21:05

uh it's like not it's actually pretty

21:07

pretty rigorous the way it's done. Uh

21:10

but um essentially what I do is I spawn

21:13

a lot of different sub agents. I have

21:15

like my main coordinator agent uh come

21:18

up with a rubric for uh what I want the

21:22

skill to do. Um, and then it spawns all

21:25

these sub aents and it's it creates

21:28

individual directories for them, uh,

21:31

which are cleverly named to not let the

21:34

sub agent know that it's being evaluated

21:37

because, uh, agents can actually tell

21:40

and when they do, they change their

21:41

behavior. Uh, but it does a bunch of

21:43

stuff like that to um essentially yeah

21:48

like test whether or not the skill I'm

21:50

making or changing is actually doing

21:53

what I think it does. Um, and one of the

21:55

really nice things about cursor is that

21:57

we are we we support so many different

21:59

models. So you can actually eval your

22:02

skill across all sorts of different

22:04

models um and you know get a sense of

22:07

how well it performs across that

22:10

different matrix. um especially for the

22:12

models that you use.

22:15

Uh so I do this a lot. Every time I I

22:17

modify a skill, I will run one of these

22:20

uh like the EBAL playbook uh and make

22:22

sure that you know it's actually leading

22:25

to a result I want. Uh but I will say

22:28

like

22:29

maintaining skills is actually pretty

22:32

hard. Uh it requires I think a lot of

22:34

taste and observation. So, you kind of

22:37

need to be very good at being a backseat

22:40

driver, you know what I mean? Like if

22:42

you do pair, if you ever done pair

22:44

programming, for example, uh, and you

22:46

watch a co-orker code and you just like

22:48

you, you could probably do this better,

22:50

you know, you could do, you know, like

22:51

why did you not do this, right? You you

22:53

ask a lot of questions to your coworker.

22:55

And it's kind of a similar thing here.

22:56

You like you don't want to just be a

22:58

passive observer of your agent. You want

23:00

to be very in the driver seat in the

23:03

initial stages when you're building up

23:04

your own set of skills. uh you know

23:06

obviously you can use something like PAC

23:08

but if you're building your own set of

23:10

skills it's very I think you know

23:13

opening up the all the tool calls and

23:15

like reading the code and reading all

23:17

the uh the agent behavior and their

23:20

thinking blocks is a really great way to

23:23

see where they they fail right like what

23:26

what you know where are they being done

23:29

and then you can go and build a skill

23:30

for that and then with verification how

23:33

you trust it is it's I think it's also a

23:36

very similar iteration loop uh where you

23:38

know like I actually did the same

23:40

process for verifying the verification

23:43

skill where I actually get um so one

23:47

thing that's interesting about eval is

23:49

that you can sort of hill climb them

23:51

meaning that uh your eval can produce a

23:54

score right uh a score that you can get

23:57

your coordinator to produce uh but also

24:00

comp uh you can have a judge agent of a

24:02

different model to uh kind cross

24:06

reference and make sure that the first

24:08

model is not being biased, right? The

24:10

model that's judging all of the sub

24:12

aents that are running the thing. Uh but

24:14

you can also like hill climb. So meaning

24:16

that you can you can use like /loop in

24:19

cursor and you can say okay keep looping

24:22

on this eval right until everything is

24:25

10 out of 10 as an example. Uh and I did

24:28

the same the basically the same approach

24:30

with the control skill. And so I kind of

24:32

it was very it was very hands-off

24:33

actually. Uh so you know I uh I kind of

24:37

built I built that skill that way like

24:39

the CLI in that skill. Um and over time

24:43

it's gotten really good. Uh but yeah it

24:46

was definitely not super smooth at the

24:48

beginning. It required a lot of

24:50

iteration and I think there's an analogy

24:53

here for me which is um well I make this

24:57

analogy later in a different slide on my

24:59

drawing here. Uh but I think of it like

25:03

uh you know as a as a engineer now

25:06

you're sort of more like you uh like

25:09

maybe a manager or the analogy I like is

25:12

like you're like a a chef in a

25:14

restaurant. Uh you you're the head chef.

25:17

Uh you're not cooking all the food

25:18

yourself anymore. You have a team of

25:20

cooks, right? You have line cooks, you

25:22

have a sue chef, you have, you know, all

25:24

these different stations.

25:26

Um and it's your job to really design

25:28

the environment. you know, you you're in

25:31

charge of setting up the kitchen. You're

25:32

in charge of, you know, like giving

25:35

tasks to different people. So,

25:39

um yeah, it's a very interesting way of

25:42

working. Uh but yeah, that's that's how

25:45

I've basically built uh these

25:47

verification skills.

25:48

>> Yeah, just just one followup there on

25:50

like to go try to go one layer deeper.

25:52

So, are you let's say we wanted to build

25:56

um an eval or a skill for for something

25:59

and we wanted to kind of get better on

26:01

its own, which is is what I think you're

26:03

suggesting. Uh are you doing that in

26:05

like a work tree kind of isolated with

26:08

like the sub agents and and then the

26:09

reviewer agent and and all that? Is it

26:11

happening like in some type of cloud

26:13

hosted environment? Like what's the the

26:15

more the practical steps? If I wanted to

26:17

go do this uh and like set up a

26:19

verification system for something, what

26:20

would I what would I do or where would I

26:21

start?

26:24

Um I think that uh the best place to

26:27

start is local because you can observe

26:31

you can definitely observe what your

26:32

agents are doing. So, uh, if you're

26:34

building a verification skill for

26:36

yourself, uh, I would definitely start

26:38

local and just have your agent bring up

26:40

the application, whether it's like a CLI

26:43

or, uh, desktop app or whatever. And so,

26:46

you can actually observe, right? You can

26:47

see how the agent is interacting with

26:50

the the application. You can see it, you

26:53

know, how it calls like the different

26:56

APIs that that allow it to interact with

26:58

the uh the application.

27:02

Um but uh for me personally uh I have

27:06

basically been kind of all in mostly all

27:08

in on cloud agents because they're

27:10

extremely powerful. Uh and the really

27:13

powerful thing about cursor is the the

27:15

cloud agents actually where if you spend

27:18

a little bit of time setting up your

27:19

environment

27:21

these control skills these verification

27:23

skills pay a huge amount of dividends

27:26

because it's not just something that

27:28

makes you as a single engineer better.

27:31

It actually levels up your whole team uh

27:33

and even your whole company because uh

27:36

you can actually start thinking about

27:38

cloud agents. start thinking about

27:39

automations that automatically do things

27:42

like uh I I get I I kind of talk about

27:46

this a bit later, but I'll just kind of

27:49

get into it. Uh where where you know,

27:51

for example, like I talk a lot about

27:53

this agent we have called Benny, right?

27:55

who who uh you know takes all of the bug

27:59

reports that we get and it automatically

28:02

goes off in the cloud, opens up a cloud

28:04

uh it's you know its desktop. It runs

28:07

cursor in its own computer and it uses

28:10

the same control skills to interact with

28:13

the application and try to reproduce the

28:15

bug uh or the user report, right? And

28:17

this is so so powerful because at once I

28:20

can immediately I I get so much

28:22

information from this automatically like

28:24

here in this example you can see that uh

28:26

the Benny actually reproduced the bug uh

28:30

but it's already fixed on main. So it

28:33

actually confirms that we fixed this

28:35

problem already and all I need to do is

28:37

just release another build of of cursor.

28:40

Uh, so that's like huge information

28:42

there that I didn't have to go off and

28:44

sit with an agent, you know, and spend

28:46

an hour trying to figure like is this

28:47

fixed, is this not fixed. So you you you

28:50

gain back so much time. Uh, but you

28:52

know, everybody on my team benefits from

28:54

this. Everybody at the company benefits

28:56

from this. Uh, so definitely think that

29:00

uh, you know, keeping these uh, using

29:03

cloud agents is super powerful. Uh, but

29:06

yeah, it's like a journey. You have to

29:08

trust it first, right? before you you

29:10

get to this point. And that's it goes

29:12

back to what I was saying here where,

29:14

you know, it's very hard. It's almost

29:16

impossible. And I would definitely

29:18

encourage you not to try to jump from,

29:20

you know, like if you're still in this

29:22

zone, you don't want to jump to like I'm

29:25

going to spawn a hundred of thousand or

29:27

thousands of cloud agents right now

29:29

because you're just going to waste a lot

29:31

of tokens. Um, and it's going to be

29:33

extremely expensive.

29:35

>> Yeah. So just to kind of recap so far,

29:37

basically the if we wanted to go on the

29:39

journey that you've kind of gone on, it

29:40

would be to start with verification,

29:43

building some some skills and some some

29:46

ways of determining that the agents are

29:47

producing at least like correct code.

29:50

Whether like you said whether it's good

29:51

code or not is maybe a separate

29:52

question, but like it's it's technically

29:53

solving the problem by looking at you

29:56

know stack traces, looking at you know

29:57

the the actual behavior in the app and

29:59

so on. Um, and then once we trust it

30:01

locally, then we can start to think

30:03

about scaling into the cloud and running

30:05

more agents that are picking up signals,

30:07

I guess, on their own, right? So whether

30:09

that's like a bug report that comes in

30:10

or something, they can go and pick it up

30:12

and solve the problem and and give us

30:14

back a PR. And then maybe the last step

30:16

is like automerging the PRs, which uh is

30:19

where you're at, maybe not where

30:20

everyone is at.

30:21

>> Um, and then and reviewing them on main,

30:23

but um, is that is that about right?

30:26

>> Yeah, exactly. I think yeah, that's why

30:27

I drew this this uh this this curve,

30:29

right? because that this this basically

30:31

describes my journey of you know when I

30:34

started barely could use a couple agents

30:36

and I was just observing every single

30:38

thing. I think there's really no

30:40

shortcut for going from here to there

30:42

because this is really about your

30:44

personal level of trust in agents,

30:46

right? Um obviously, you know, as a as

30:49

engineer, you don't want to just slop

30:50

code into production. So, how do you

30:53

actually build up that trust? Takes um a

30:56

lot of uh I guess taste and judgment. Um

30:59

but, uh you know, like I think plugins

31:02

like Pstack definitely kind of help you

31:05

uh get up to speed much quicker. Uh, and

31:08

so I guess it's like if you trust me and

31:11

you trust Pstack, then in by extension

31:14

you can maybe trust your agents. But if

31:16

you don't trust me and I I definitely

31:18

would not encourage people to blindly

31:20

trust me. Uh, uh, you know, if you build

31:24

up your own set of skills that you can

31:26

obviously, you know, take a look at PAC

31:28

and kind of fork it, make it your own,

31:30

improve the skills. Definitely encourage

31:32

that. Uh but for me it's really all

31:35

about it just keeps coming back to

31:37

trust. You know every one of us here in

31:39

this chat have a different standard for

31:41

engineering. Uh and there are different

31:43

things that are important for us in our

31:46

codebase. And uh when you are able to

31:50

encode all of that into skills and you

31:51

can verify that your agent is actually

31:53

doing them that allows you to really

31:55

kind of ascend this curve and um uh you

32:00

know start automating things. Uh there's

32:02

another piece I wanted to talk about. Um

32:05

if there's more

32:06

>> Yeah, go for it. I'll I'll pick up more

32:08

questions as I go. But

32:09

>> yeah, I think there's a third part to

32:10

this which I haven't talked about yet,

32:12

which is kind of an interesting one

32:14

which is like refactoring and rewriting

32:17

like one of the uh I guess most

32:19

controversial one of the most

32:21

controversial topics in the industry I

32:23

think is like should you rewrite your

32:25

app or not? Um because I think engineers

32:30

are very prone to this where especially

32:32

when you join a company you come in and

32:34

you see like the code base and you're

32:35

like oh man this is like who wrote

32:38

this code you know it's terrible I want

32:40

to rewrite the whole thing there is a

32:42

very common inclination and I think a

32:44

lot of you know before agents um and I

32:48

guess arguably even now people will

32:50

definitely discourage you from re

32:51

rewriting stuff but I'm actually here to

32:54

make a case for why you might want to

32:56

consider

32:58

Um because

33:00

I think it really depends. Uh you know

33:03

uh brownfield applications I think are

33:05

actually in a pretty good spot

33:07

especially if they're set up well

33:09

already. Uh and like recently I've been

33:12

talking to some people but uh you know I

33:14

I was just observing I I just noticed

33:17

this parallel which is that a lot of big

33:21

tech company problems are now

33:23

everybody's problems. Um, and the big

33:26

tech company problem, you know, like

33:27

when I was working at Meta, like we had

33:29

this giant monora repo, we had like, I

33:31

don't know, tens of thousands of

33:33

engineers just, you know, like banging

33:35

on their keyboards and and shipping code

33:38

and

33:40

a lot of really great engineers at Meta.

33:42

Uh, but, uh, I'll say like, you know,

33:45

you'll be surprised that the code

33:46

quality is actually not that good.

33:48

[laughter] Um, and so I often joke that

33:51

like, you know, before AI sloth, we had

33:53

human sloth. Um and so uh you know I

33:56

think a lot of big tech infra like uh

33:58

like what Meta has or Google you know

34:01

you know really big tech companies are

34:03

actually designed for that where you

34:05

you're sort of like you're catering to

34:07

the the you know like uh this sounds so

34:10

bad to say but like the the least

34:12

capable engineer on your team right you

34:14

build you build frameworks you build

34:16

conventions you build guard rails you

34:19

know you restrict credentials so that

34:21

you know your intern doesn't wipe your

34:22

production database

34:24

Um

34:26

there's uh you know if you have that

34:28

level of infra already I think your

34:31

agents can actually already do a very

34:33

solid job right because they have the

34:35

the guard rails are already in place for

34:38

agents to not cause havoc or not cause

34:41

too much havoc uh in your codebase um

34:44

and you can always add more you know

34:46

guard rails. Uh but I think like green

34:49

field applications especially are you

34:50

know like the brand new applications are

34:53

like the biggest risk in my opinion. Uh

34:56

and also the greatest opportunity

34:58

because you know if you vibe code a

35:00

project uh a prototype um like we did

35:04

for Grockbot you know Grockbot was spun

35:05

up very very very quickly. Um and if you

35:08

if you haven't heard of of Grockbot it's

35:10

like a a new application we just

35:12

launched yesterday. Uh it's it's really

35:14

cool. uh lets you orchestrate your

35:18

create like individual agents that have

35:20

their own identity and you can kind of

35:21

orchestrate them. It's super cool.

35:23

Definitely check it out. Um but yeah,

35:25

that was it's like a very it was a very

35:26

green field application like most

35:29

prototypes are so like vibe coded very

35:32

quickly. Humans were not reading the

35:34

code at all. And uh I had this tweet

35:38

recently uh where I said something about

35:41

organic architecture. Um

35:45

maybe I'll find it. Uh but the idea is

35:49

that

35:50

uh when you have a completely vibe coded

35:52

application, you essentially have no

35:54

guard rails whatsoever. So uh your

35:57

agents

35:59

when you give them a task, they will

36:01

just solve it in whatever method is the

36:03

most convenient. And over time you get

36:06

into this uh situation where you have a

36:09

codebase that is spiraling out of

36:11

control because you don't understand it.

36:14

Uh your agents understand it I guess in

36:16

a way but like they've built something

36:18

that is you know optimized for short for

36:21

shortcuts. Uh and uh you know it will

36:24

you will suffer you'll have a lot of of

36:27

issues with that application.

36:30

Uh so I think starting your codebase

36:33

with uh like very strong constraints is

36:37

very much needed uh because like when

36:40

you have a codebase that you can trust,

36:43

right? when you have guardrails that

36:44

actually help you uh uh help your agents

36:48

write good code, you can get into the

36:50

you know like into this part of the

36:52

curve where I I where like I I said you

36:55

know I woke up today and I had like 20

36:57

PRs merged u by my agents and that's

37:00

because I invested a lot a lot of time

37:04

uh over 600 PRs I I I calculated

37:06

yesterday uh when I refactored all of

37:10

Grockbot to this new architecture that

37:12

I've been

37:13

Um,

37:15

and yeah, I've gotten to a point where I

37:19

I don't really look I really don't look

37:20

at the code anymore. And um, I say that

37:24

not just, you know, to sell you tokens,

37:25

but because I, you know, it it it took a

37:29

lot of work to get to that point. I

37:30

spent a lot of tokens to get the

37:32

codebase to this point where I no longer

37:34

have to look at it. Uh but I'm very

37:37

excited because you know of the

37:39

potential where you know it's not just

37:41

this doesn't just benefit me it benefits

37:43

everyone contributing to Grockbot and it

37:46

also empowers you know designers and

37:49

product managers and you know pe uh even

37:52

GTM people to add features to Grockbot

37:55

and I don't have to worry you know I

37:56

don't have to to wake up at night in in

37:59

the middle of the night and worry like

38:00

oh someone's just merged a perf

38:01

regression right I have a ton of

38:05

constraints and CI is like it's actually

38:07

very annoying to write code in in graph

38:09

web but like agents absorb all of that

38:11

annoyance.

38:13

Um but yeah I'm happy to talk about what

38:16

exactly that is. Um

38:18

>> yeah I think one question um

38:21

>> before we get into the this part here is

38:23

just around that element of like what

38:26

your your your CI looks like or maybe

38:28

some of the constraints and then also

38:29

like the average PR size. I saw a

38:30

question about that earlier just to give

38:32

people you know kind of a a glance. So

38:34

it doesn't have to be like

38:35

mathematically average, but just uh you

38:37

know like what generally the size of the

38:39

a PR is. Um if it's only a couple lines

38:42

of code or you know um yeah

38:45

>> um

38:47

I think it depends uh

38:51

I'm trying to do this in a way where I'm

38:53

not going to like

38:54

>> you. Yeah, you don't have to share the

38:55

actual number like an actual average. I

38:57

think this this is fine

38:58

>> benchmark

38:59

>> but like we have so okay this is not

39:02

that interesting but uh well fun fact is

39:05

that virtualization in grockbot and in

39:09

uh cursor is actually powered by uh

39:12

pretext uh which is a sort of new

39:15

library that someone's built um that's

39:19

really interesting you should you should

39:20

check it out but that's not really that

39:22

important uh I think the average PR size

39:25

I actually don't know. I pro I don't

39:27

know if I want to click on these. Uh I

39:29

probably can, but I would say like they

39:32

can range anywhere from a few hundred

39:34

lines or 50 lines to like a thousand

39:37

depending on what the thing is doing. Uh

39:40

so like here I'm actually like deleting

39:42

a bunch of files. So I expect that it's

39:44

just like mostly deletion. Uh but yeah,

39:47

it kind of varies.

39:49

>> There's no like Yeah,

39:51

>> there's no like hard cap or hard limit.

39:52

Are there all like 50 line P?

39:53

>> There's no hard cap. Yeah, there's

39:54

definitely no hard cap, but I I do

39:55

encourage my agents to split up their

39:57

work into multiple PRs. Uh I do that

40:01

mostly because uh I like I like the idea

40:05

of the I guess maybe this is much harder

40:07

to do now as in the world of agents and

40:10

you have like so many commits, but I

40:12

like the idea that you know the git

40:13

history is a very rich source of

40:15

context. Uh, and I like the I like each

40:19

PR to sort of atomically describe what

40:22

that small piece of thing is doing,

40:25

which also makes it easier for me to

40:26

revert changes and like figure out, you

40:28

know, oh, I shipped a bug and it's just

40:30

it's here, right? It's not in this

40:32

40,000 line PR where who knows what

40:36

landed in there.

40:38

Uh, but yeah, I don't have a hard cap on

40:41

PR size.

40:42

>> Cool. And then um yeah, also quick

40:45

question on like CI. So again, you don't

40:46

have to go into like uh the screen share

40:48

of like your CI does, but just generally

40:51

would you describe what the CI kind of

40:53

looks like uh or how strict it is?

40:56

>> Uh yeah, so

40:58

uh well specifically for Grockbot. So

41:01

Dune is the is the sort of cheeky code

41:04

code name for the architecture that

41:06

we've built for Grockbot. Um the CI

41:10

looks pretty annoying because there's

41:13

checks for everything. So like literally

41:16

I have um uh well if you've written any

41:19

React for example you know you know that

41:20

one of the biggest foot guns in React is

41:22

use effect. Uh so in

41:26

uh Dune and in graphbot we've banned use

41:29

effect. So Dune is just you can the the

41:32

mental model of what Dune is uh you can

41:34

kind of think of it as like Nex.js JS

41:36

for uh electron apps and it's designed

41:39

for agents to write uh and it's like

41:42

custom for you know our agent powered

41:45

applications. Um so the CI checks are

41:48

very like specific to that like you know

41:50

don't use use effect it's it's it's

41:52

banned like CI will fail uh and yell at

41:55

you. We have like some of the more

41:57

interesting ones that people might raise

41:59

eyebrows is like I actually ban code

42:01

comments as well. Uh which is very

42:04

interesting. Uh but I've noticed that

42:08

99% of the time agents just write code

42:11

comments that kind of describe some

42:13

historical thing that is actually

42:15

totally irrelevant to the code. Um like

42:18

it will often say like you know oh

42:19

Lauren said you should never do this and

42:21

it's now in in a code comment. I'm like

42:23

what? like why what that was I didn't

42:26

say that as like a durable you know

42:27

global rule I just meant like your this

42:30

PR sucks and you should change that part

42:33

but agents don't really understand us

42:36

that well surprisingly uh and or they

42:39

kind of assume too much and they kind of

42:41

do things in like very stupid ways so

42:44

like yeah we just ban everything

42:46

everything you can imagine like the

42:48

agents are bad at we ban uh so one

42:52

example that we actually suffer for a

42:53

lot in the agents window is we have uh

42:57

you know you if you've used the agents

42:58

window you've definitely seen

42:59

performance issues and you know we're

43:01

constantly trying to fix them. Uh but

43:03

it's like a it's a never- ending

43:06

struggle because there's so many pull

43:08

requests that get merged every any one

43:10

of them could just regress performance

43:12

or stability or reliability. Uh you know

43:15

the agents window doesn't have this

43:17

architecture yet. I plan to do bring

43:19

this learning back there and kind of

43:21

refactor everything there. Uh but uh it

43:25

just regresses super often uh because uh

43:29

there's just one example is like we have

43:31

very poor um isolation between

43:34

processes. So like on you know on on

43:36

Electron you have a renderer thread that

43:38

renders your UI but you also have like a

43:40

main thread that you can run other code

43:42

that you know doesn't need to block the

43:44

renderer.

43:46

Um, but we do a poor job of separating

43:49

those things. And so oftentimes you just

43:51

accidentally have code that gets pulled

43:53

into running on the renderer thread and

43:56

then all of a sudden you're competing

43:57

with the the renderer that you know that

44:00

has a very if you want like 60 fps you

44:03

have to every frame that gets drawn has

44:05

to be done in 16 milliseconds. So very

44:08

very small you know deadline per frame

44:11

uh if you want you know a very smooth

44:13

product. Uh and when you start building

44:15

bringing in accidentally bringing in you

44:17

know things that are like very

44:19

computationally heavy or they have a lot

44:21

of IO uh then you just get into like a

44:24

lot of jank right your your FPS really

44:27

drops you start uh you know losing

44:29

frames you get long tasks that take more

44:32

than 16 milliseconds and you just get

44:34

this really choppy experience.

44:37

So all of those patterns that we've

44:38

learned basically building electron apps

44:40

we've encoded into this framework and it

44:43

becomes like a hard failure. So I

44:45

literally in in grabbot we literally

44:47

have a directory called electron main

44:49

electron renderer and we have uh import

44:54

uh CI guess where we actually check the

44:58

dependency graph to make sure you're not

44:59

accidentally importing code from one

45:02

directory to another. Uh so that's

45:04

enforced by CI um as well as bug bots uh

45:09

which is our which cursors um like code

45:13

review tool that runs on CI uh you know

45:15

in our agents MD it's everywhere like so

45:18

I I I um I have this thing here where I

45:22

I talk about like um you know like there

45:25

are multiple layers I think for building

45:27

a good codebase. Uh obviously the

45:30

codebase is one where uh if you have an

45:33

architecture like this where it's

45:34

extremely strict uh you know the the the

45:38

way to build features is very

45:39

conventional that's like the strongest

45:42

strongest level of enforcement because

45:44

agents just love to copy existing

45:46

patterns. So uh one example of this in

45:49

rockbot is like we have this these

45:51

concepts called like a feature and we

45:54

have entry points and transcript cards

45:56

like oh you know the cards that you see

45:57

in the chat these are all like like

46:01

nouns I guess in in the framework and so

46:03

there's a very conventional way of

46:05

creating them and so like a feature is

46:09

all in in a single directory as an

46:10

example and so all of the code that

46:13

contributes to that feature lives in one

46:15

directory so it's all collocated in one

46:17

place. Makes it super easy. You know,

46:19

agents don't have to like uh grap around

46:21

and try to figure out like where all the

46:23

things are. It just looks at the feature

46:25

and like, oh, okay, I'm working on the

46:27

onboarding feature in Grockbot. Uh I'm

46:31

just going to work in this directory.

46:32

And for 80% of the work, it's mostly

46:35

just very encapsulated there. But, uh

46:38

like it's like very it's like designed

46:41

again for you know like the dumbest

46:43

agent like you don't have to think,

46:45

right? the the the one of the key

46:48

principles I have for this framework is

46:50

like the shortest the shortest path is

46:52

the best path.

46:55

So uh because that plays exactly to how

46:57

agents love to write code is like they

46:59

like to take shortcuts really you know

47:01

they they'll find the quickest way to

47:04

solve the problem. So why not make that

47:06

the best way to solve the problem? Uh so

47:10

I I I probably won't get into all the

47:12

specific details. Um and uh the the this

47:16

framework is really more of a collection

47:17

of ideas and principles rather than

47:19

something that will open source. Uh you

47:22

can you can you know screenshot this I

47:23

guess if you want and uh tell your agent

47:26

to uh do some build build something like

47:29

this for you too.

47:31

Um yeah, but it's really all about the

47:34

layers uh you know like the the codebase

47:36

is one part with features uh and

47:38

directories and you know import or

47:41

blocking import dependencies uh that

47:44

shouldn't be imported. Uh but and and it

47:47

all enforces that and static analysis.

47:49

So like uh there's CI checks, we have a

47:52

lot of lints for bad patterns that we

47:55

observe. uh compiler diagnostics

47:58

uh there's also rules and bugbot which

48:01

are um I think like three four five are

48:04

more soft right these two actually make

48:08

make CI red right so that you know

48:11

there's a hard constraint where the

48:13

agent can't just write crappy code

48:16

for rules and skills and buggbot your

48:20

agents can still forget right you can

48:22

still or it may not always consistently

48:25

apply them. So, I like to layer them,

48:29

but I don't I don't like to rely on them

48:32

as the only source of enforcement

48:35

because it's very very soft, right? And

48:36

if you if you only have rules and bug

48:38

ball and skills and a style guide for

48:40

your code, you will it's only a matter

48:43

of time before your code base looks like

48:45

complete trash. I'm sorry to say that,

48:48

but uh I definitely recommend yeah like

48:50

you know investing in you know things

48:52

that can be hard and forced, right?

48:55

Okay. And this is why you know maybe uh

48:57

the choice of text stack that you use is

48:59

also very important. Um like I think for

49:03

example Rust is sort of making you know

49:05

it's like getting super popular again.

49:08

Uh because the compiler is so strict

49:11

right the compiler enforces so many

49:13

different things. you know, there's a

49:14

borrow checker that you have to appease

49:16

and if as long as you make sure your

49:18

agents don't write unsafe code blocks,

49:21

uh you can more or less feel somewhat

49:23

confident that if the code compiles, it

49:25

probably works and it's good. Uh but you

49:28

see it gives you that level of trust and

49:30

confidence that you as a human engineer

49:34

no longer need to go and check it

49:36

yourself. you know, you you rely on code

49:40

and static analysis to actually make

49:43

that uh a lot smoother. Um, and I I

49:48

guess the worst part, the worst place to

49:50

be in is if you are stuck in code review

49:53

land where you actually enforce all of

49:55

the constraints, the invariance in your

49:57

codebase by literally the human person

50:01

saying, you know, reading the code and

50:03

like, okay, you should not do this,

50:04

right?

50:06

Every time you have to do that, you

50:07

should consider that as a code smell

50:08

like a anti- pattern and you should say,

50:11

"Okay, instead of me commenting on the

50:13

PR, how do I turn this into a hard

50:17

rule?" Right? How do I turn this into a

50:18

lint rule? How do I turn this into a CI

50:21

failure? Or how do I even categorically

50:23

eliminate this problem uh entirely? Uh I

50:28

I can talk about another migration I've

50:30

done, but I'll probably pause here.

50:32

>> Sure.

50:33

>> Yeah. I feel like that's that's where I

50:34

am to be honest is is what you're

50:36

describing right now which is that like

50:38

I don't have all of these rules. So I

50:40

have some things to go do after this

50:41

session in terms of being able to scale

50:43

my agents. I'm I'm definitely on like

50:45

the uh you know maybe a couple of

50:47

parallel ones locally stage like two to

50:50

three locally and I'm sure most people

50:52

here are on the same. So uh yeah I know

50:54

we only have a couple minutes left.

50:56

Lauren, was there anything else that you

50:57

wanted to to highlight? I obviously

50:59

there's lots of questions so I can grab

51:00

more but I want to give you a few

51:02

minutes if there's anything else you

51:03

want to talk about.

51:03

>> Well, I think I've been yapping for

51:04

quite a lot so maybe let's just do

51:07

questions.

51:08

>> Okay, cool. Uh, one question that had uh

51:10

a couple of uh came up a couple times

51:12

was just around like token usage.

51:15

>> So the the question is like is is what

51:17

you're describing a realistic thing for

51:19

people who are on you know uh a normal

51:23

set of token usage. They don't have you

51:25

know basically unlimited tokens uh to

51:27

work with.

51:29

I think that's a really good point. I

51:30

mean, like obviously, you know, I work

51:32

at a AI lab where we have unlimited

51:34

tokens. So, uh I definitely cannot

51:39

say that, you know, this is something

51:41

everyone should do in the exact same way

51:43

that I did it. I think it's possible to

51:45

get to this point without, you know,

51:47

breaking the bank.

51:49

But you know if you're like an

51:50

engineering leader or you know you're

51:52

you you have a startup that you lead um

51:55

I think to me it's a question of ROI um

51:58

and it's like uh yes you spend a lot of

52:03

money on tokens in the upfront stage you

52:06

know like refactoring a code base is

52:07

going to take a lot of tokens uh adding

52:09

all these things uh is going to take a

52:11

bunch of tokens but if we're heading to

52:14

a world where agents are writing all the

52:16

code and you know You want to be very

52:20

lean, right? You don't want to have to

52:22

hire, you don't want to be, you don't

52:23

want to become like meta, right? Like I

52:25

mean like in terms of you don't want to

52:26

become a 10,000 person engineering or

52:29

because I mean that's a cool problem to

52:32

have, but also you you have so much

52:34

overhead. There's like planning, you

52:37

know, like you it's it's a personally I

52:39

I wouldn't uh it it's not super fun. But

52:43

um I think you want to stay very nimble,

52:46

right? And you want to you want to be

52:47

like agents are all about allowing you

52:50

to do things that you couldn't do

52:51

before. That's really to me like the

52:53

value of agents, you know, it's not just

52:56

storing tokens on every single little

52:58

thing, but um to me like the thing I

53:00

couldn't do before is like enforce this

53:03

level of constraints in a codebase by

53:06

myself, right? Like I'm just a single

53:09

person, you know? Uh it would have taken

53:11

me years to build this framework uh and

53:15

do all the refactoring and

53:18

test everything myself and verify you

53:20

know like run imagine if there it was

53:22

just me right know in in pre- agent era

53:25

just like running you know by it would

53:27

take me so long right and my salary is

53:30

pretty high right like so you know the

53:33

the question I think an engineering

53:35

leader might have is just then you know

53:37

like what is there's a trade-off of do

53:39

Do you hire someone to do this or do you

53:42

spend the tokens to set up a codebase so

53:45

that even the the most naive, right, the

53:48

dumbest agents can do a good job? And

53:51

when you actually get to this point,

53:53

like even agents that are not, you know,

53:55

fable size do an excellent job of

53:58

writing code. And this pays a lot of

54:00

dividends as well for me personally

54:02

where I've empowered not just myself but

54:06

again like PMs, designers, engineers who

54:10

are not familiar with Grockbot to just

54:12

contribute in a way that is sustainable.

54:16

So I think yeah it's definitely like a

54:18

trade-off for sure. You know like

54:20

nothing is like free for sure. Uh and

54:22

tokens are pretty expensive. Uh but oh

54:25

actually uh I I I don't know how many of

54:27

you have seen this but we actually

54:29

announced Grock 4.6 today so very

54:32

exciting finally out um so yeah graph

54:35

4.6 would be like a great it was very

54:37

very smart uh it's really good on the on

54:40

the benchmarks uh and it's the same the

54:43

tokens uh well uh I hopefully I'm not

54:45

saying this incorrectly but uh I believe

54:48

the cost per token is the same as 4.5 so

54:52

you're actually getting more

54:53

intelligence for the same cost uh I

54:57

think this is an area that cursor tries

54:59

to cursor and spaceex AI try to really

55:02

optimize for like that heredto frontier

55:05

here of you know cost versus

55:07

intelligence. Uh you know we don't

55:09

necessarily want to build the biggest

55:10

model ever because that is extremely

55:13

expensive to run. It's really about like

55:15

how do you find that sweet spot right?

55:17

You don't you don't need a giant model

55:18

but it's just super smart right and it's

55:20

not very expensive for inference.

55:24

Uh but um yeah, I think to kind of round

55:27

it up, um I think it's like a it's it's

55:30

there's a if you do your own analysis, I

55:33

feel like it's pretty positive. It it'll

55:36

be pretty positive that the ROI you get

55:38

from investing in stuff like this, uh

55:41

just empowers not just yourself, but

55:44

your whole team to be so much more

55:46

productive, right? Like imagine if you

55:47

have an army of engineers like me who

55:50

are shipping so much improvements and

55:52

and bug fixes uh you know every day

55:56

right like that is pretty exciting.

55:59

>> Cool. Uh one last question before we

56:01

wrap up. This one is for the people in

56:03

product on the on the call.

56:05

>> So let's say we do have an army of

56:07

engineers who are shipping like Lauren.

56:09

I'm just curious like how is the product

56:11

team or other functions of your company

56:13

keeping up given that like if you're

56:15

shipping so quickly have are they using

56:18

AI more to do their jobs? Like as much

56:20

as you can speak to that obviously you

56:21

don't have like you're not in that role

56:23

but just curious about how that works.

56:25

>> Um I think this is where Grothbot has

56:28

been actually exceedingly powerful. uh

56:31

where so before graphbot like you know

56:33

uh obviously cursor only had cursor like

56:36

we only had agents window we had a CLI

56:39

we had an IDE and these are really like

56:42

power user tools right like de they're

56:44

designed for developers so it's very

56:46

very developerentric you can do

56:48

knowledge work in them but it like the

56:50

UI is not really optimized for that so

56:54

we actually didn't really have uh well I

56:57

think like a lot of people like you know

56:58

GTM product like they might have used

57:01

cursor uh to do their work but it

57:04

definitely wasn't like a delightful

57:05

experience for them. Um I think now with

57:08

Grogbot

57:10

uh it's become Grogbot is basically like

57:13

the kusher moment for people who are not

57:16

in tech in my opinion like it's like

57:18

it's like a very very accessible way to

57:21

use agents in a very comfortable very

57:24

familiar interface. It looks like

57:25

iMessage. Um, and it's very fun to, you

57:29

know, you can give your agent a fun

57:30

name. Uh, you can have you can kind of

57:33

do orchestration with in a very like

57:35

natural way where you can sort of, you

57:37

know, each agent's like a person and now

57:39

you got a team of agents like working on

57:41

you have one one agent per account that

57:42

you manage as an example. Or if you're a

57:45

PM, you have, you know, you can have an

57:46

agent that summarizes all the work that

57:48

Lauren did last night and then now you

57:50

know what I did, right? So I think our

57:53

PMs are leveraging that a lot and

57:55

they're shipping code too. Uh so you

57:58

know like often times they will just say

57:59

oh here's a bug I fixed it can you look

58:01

at it and then I'll go review it and

58:03

actually it's just perfect. I'm like

58:05

okay stamp. Uh so uh that I think that

58:08

shows that you know the the Dune

58:09

architecture is holding up right the all

58:12

the the really strict constraints allow

58:15

people who are not experts in

58:17

engineering to contribute at a high

58:19

level. Uh so I'm I feel like I'm already

58:21

seeing that pay off a lot where uh you

58:24

know designers and PMs are just able to

58:27

to to ship features directly. Um and

58:31

that just makes the Grockbot team super

58:34

fast, right? Where we can ship so

58:36

quickly. Um and we have a lot planned.

58:40

So I'm very excited uh to you know uh to

58:44

to ship more ship more stuff.

58:46

>> Yeah, that's awesome. Uh well, we are at

58:49

time. So, uh I guess Lauren, if if folks

58:52

want to support you, maybe go try out

58:53

Crockbot, try out uh 46 and uh you know,

58:57

get provide some feedback. But yeah,

58:59

this was awesome. Really appreciate you

59:00

taking the time. Uh thanks everyone for

59:02

all the messages in the chat. Lots of

59:03

good questions. I know we didn't get

59:04

through everything, but as I kind of

59:06

said at the top, way more questions than

59:07

than we could get through, but uh yeah,

59:09

really really thanks thanks for for

59:11

joining. Thanks everyone for joining and

59:13

uh hopefully you enjoyed the session.

59:15

>> Yep.

59:16

>> All right.

59:16

>> Yeah, I see it. Thank thanks for having

59:18

me. And uh if you have any more

59:19

questions yet, just DM me on Twitter.

59:21

I'll I'll open them up. I guess I'll let

59:23

the let the

59:24

>> You're going to get a lot of DMs.

59:26

>> Yeah, I'll open the updates. So yeah, DM

59:28

me. Maybe I'll do like a Twitter space

59:30

at some point. That's all for more

59:32

questions. But

59:33

>> really appreciate everyone for showing

59:34

up uh you know, taking an hour out of

59:36

your day.

59:37

>> Yeah. All right. Thanks all. I'll see

59:38

you the next one.

59:39

>> Okay. Thanks everyone. And bye.

Interactive Summary

Lauren Tan shares her journey and strategies for effectively using AI agents in software development at Cursor. She emphasizes that building trust in agents through 'verification skills'—such as running code, taking traces, and using feature maps—is crucial to moving from a manual, micromanaging loop to an automated, high-velocity workflow. Lauren also discusses the importance of architectural constraints and rigorous CI practices to prevent 'sloth' and errors, allowing even non-experts to contribute safely to projects like Grockbot.

Suggested questions

4 ready-made prompts