HomeVideos

The man behind Cursor's "memory" feature

Now Playing

The man behind Cursor's "memory" feature

Transcript

340 segments

0:00

[Music]

0:12

Hey guys.

0:15

Cool. Um, yes. So, my name is Yash. Um,

0:18

I work at Kurs on the engineering team.

0:20

And, um, I'll apologize in advance

0:22

because we don't have anyone who work

0:24

who has slides. So, I had to just kind

0:26

of put these together my own. So uh you

0:28

know bear with the very poor design but

0:30

anyways um so I'm gonna talk a little

0:32

bit about how we approach context

0:34

generally.

0:34

>> Wait ying one key point here.

0:37

>> Yes and um I have been working on

0:40

memories um by myself at cursor for the

0:42

past few months just prototyping and

0:43

we're kind of making our way out of the

0:45

woods so going to talk about that.

0:46

>> Okay so Yash isn't bragging for himself

0:48

as much as he wants to ask them how many

0:49

people are working at memory and cursor

0:50

and he goes I'm the only one. So we're

0:53

talking about a $9 billion IDE here

0:55

where memory is a loadbearing process

0:56

for it and Yasha is the one that's

0:57

working on it. So very excited to see

0:59

this.

1:00

>> Thank you so much.

1:02

>> Um

1:04

thank you. Um but yeah, so I think like

1:07

very similar to um what other folks have

1:09

touched upon, uh when you try and look

1:11

at context as this big kind of word, I

1:13

think it gets really really confusing

1:14

and muddy. And so we kind of tried to

1:17

break it down into three different

1:18

categories for our agent. Um and

1:20

obviously this is very much an art and

1:22

not a science. So and it's very product

1:24

dependent as well. But for us we kind of

1:26

have three different types. So you have

1:28

directional and the idea there is uh

1:30

when you present a highle task to the

1:33

agent um usually like it will kind of

1:35

begin its search um wide and directional

1:39

context will allow it to narrow that

1:41

search early and earlier and kind of get

1:42

to the relevant set of files as quickly

1:44

as it can. Um the next thing is

1:46

operational and that's kind of runbook

1:48

related. So how do I deploy a service?

1:51

How do I make edits in this particular

1:52

file? What are the conventions? Um and

1:54

so we have already created cursor rules

1:57

for that which are written by a human,

1:59

but the model will basically fetch them

2:00

when they're relevant to context and use

2:02

that to guide its edits. Um and the last

2:05

thing is a little bit more fuzzy. It's

2:06

kind of similar to the holistic theory

2:08

of mind that Sam mentioned earlier. Um

2:10

and that's kind of like a behavioral

2:12

context where you want the model to act

2:13

in a certain way when it's um going back

2:15

and forth with you. And so we had this

2:17

concept of user rules um where you could

2:20

specify, okay, please only speak to me

2:22

in Spanish or something like that. Um

2:24

but to be honest, most of our effort so

2:27

far has been just focused on the first

2:28

category of codebased search. And

2:30

recently we've kind of been branching

2:31

out, which is what I've been working on.

2:33

Um, and so the goal of memories was is

2:36

basically to augment all three of these

2:37

types of contexts with the idea that,

2:40

you know, you can rely less and less on

2:41

cursor rules, less and less on user

2:43

rules, and maybe even less and less on

2:45

codebased search once we've learned more

2:47

and more about how you interact with the

2:48

agent.

2:51

Okay. Um, and yeah, so like most things

2:54

at Cursor, we start with prototypes. So

2:56

I've been prototyping memories for about

2:58

a month and a half now. Um, and our

3:01

general principle was to start pretty

3:02

conservative. And the reason for that is

3:04

because um cursor is a coding agent

3:07

where the user generally likes to feel

3:09

in control. And one of the most

3:11

frustrating things about using cursor is

3:12

when it kind of goes off the rails and

3:14

stops listening. Uh which I'm sure if

3:15

any of you have used it, you must have

3:17

experienced it before. Um, and so one

3:20

worry and one thing that we saw a lot

3:22

during prototyping is if the model gets

3:24

an incorrect memory about your codebase

3:26

or about how you like to run the

3:27

terminal or anything really um it's a

3:30

really frustrating experience um because

3:31

the model will refuse to do certain

3:33

things or will do things incorrectly and

3:35

then when you try and tell it like hey

3:36

this isn't the right way it will

3:38

actually like double down because it's

3:40

like oh no I have a memory like this has

3:41

got to be correct. Um and so I'll get to

3:45

that part later but basically when we

3:46

started prototyping um we began with

3:49

kind of two parallel approaches that I

3:51

tried out. So one we called like the

3:53

sidecar approach and this is where uh

3:56

the model doesn't call any tools to

3:58

generate memory but rather in the

4:00

background as you interact with the

4:01

agent we kind of pull relevant parts of

4:03

context into a smaller model and that

4:06

smaller model makes a decision on what

4:08

to save if to save anything at all and

4:10

also what to update in the existing uh

4:12

knowledge that it has. Um and then the

4:15

second approach we tried was a tool call

4:16

approach where you just give the model a

4:18

tool called update memory. Um and the

4:21

model kind of as it detects you

4:23

interacting with the agent will decide

4:24

on its own you know this memory is

4:26

incorrect. I'm going to update it or I'm

4:27

going to delete it or I'm going to add

4:28

something new.

4:32

So in the second one the mo

4:36

okay so in the sidecar approach you

4:37

basically have a small model listening

4:40

to your conversation and at some points

4:42

you basically send the conversation off

4:44

to that small model and the small model

4:46

makes the decision completely

4:48

independent of the main of the main

4:50

thread. Um and then versus the tool call

4:52

approach is all in the main thread. Um,

4:58

>> uh, the sidecar doesn't necessarily have

5:00

to be an agent. Like you can choose to

5:01

give it tools if you want to, but you

5:02

could also, we've tried, we tried both.

5:04

Um, so yeah. Um, also feel free to

5:07

interrupt me with questions at any

5:08

point. Um, no, no, no, that was great.

5:11

I'm sure other people had questions,

5:12

too. Um, and so yeah, I'll start with

5:14

tool call memories. Um, this was

5:16

definitely the simplest thing to to

5:18

implement. Um and so you just have the

5:20

model kind of reflect um and decide when

5:23

the user has expressed something to it

5:24

that it determines is like worth

5:26

remembering and in that case um it will

5:29

create a memory itself. Um and so what's

5:31

interesting is we noticed like as I was

5:33

prototyping over the past few months

5:35

like model capabilities have improved in

5:37

such a way that this is like a legit

5:39

approach where you can just give the

5:40

tool to the model and it works. Um, and

5:42

so specifically, Sonnet 4 and Opus 4 as

5:45

well are really good at instruction

5:47

following and they've also been RL

5:49

specifically on anthropic's definition

5:52

of memories. But the tricky thing is

5:54

Anthropic actually has a different

5:56

definition of memories than what we

5:57

have. And it kind of goes to the point

5:59

that everyone has you need to decide

6:01

what memories are for your product. You

6:02

can't just think of it in the abstract.

6:04

Um, and what Sonnet wants to do with

6:06

memories is kind of create a task log.

6:09

So you can see that in like their

6:10

Pokemon example where they where the

6:11

agent kind of keeps a log of all the

6:13

things that it's tried to do. Um and so

6:16

when we just gave the gave sonnet 4 a

6:18

memory tool, it would try and save like

6:19

task specific memories which is in

6:21

coding agents like at least for us it's

6:23

kind of the opposite of what you want

6:24

because you don't want to remember like

6:26

the things that were specific to one

6:27

particular conversation. You want to

6:29

remember the things that were

6:30

generalizable and will be useful in a

6:32

future like completely unrelated

6:33

generation. Um and to that end uh we've

6:36

now like kind of tried to keep sonnet 4

6:39

and keep um opus 4 from generating new

6:42

memories but they are still really good

6:44

at reflection. So uh if the model has an

6:46

incorrect memory and the user just in

6:48

natural language kind of expresses

6:50

disagreement then they're really good at

6:51

just updating their memory themselves

6:53

without having to do anything fancy in

6:54

the background.

6:58

And the second approach is the sidecar

7:00

which I was talking about. And so I

7:02

spent weeks iterating here on the

7:03

prompting. It was like a whole roller

7:05

coaster of up and down where I kind of

7:07

lost hope and got hope back and kind of

7:09

settled in somewhere where I'm happy.

7:11

Um, and the biggest thing that I was

7:13

fighting is this concept of like a task

7:15

specific memory because in an coding

7:18

agent interaction probably 90 plus% of

7:20

what you're saying to the agent is not

7:22

worth remembering. It's very specific to

7:24

the task that you've given the agent.

7:26

Um, and so probably in most

7:28

conversations there isn't even anything

7:29

worth remembering. um for this

7:31

particular type of memory. It depends on

7:32

what you want to remember. Um and so I

7:35

tried a bunch of different approaches to

7:38

basically keep the model from focusing

7:40

too much on task specific things. Um and

7:43

similarly like with the uh tool calls it

7:46

kind of changed as the models got

7:47

better. So, um, you know, back when

7:50

cloud 3.7 was the latest model, um, I

7:54

was kind of prompting it super

7:55

aggressively, like giving it a bunch of

7:57

examples. Um, and it sort of worked, but

8:00

it also wasn't super great and a lot of

8:01

things kind of slipped through the

8:02

cracks. Um, and with the newer set of

8:06

reasoning models, um, essentially you

8:08

can just kind of give them like a very

8:09

brief description of the problem. uh you

8:11

can lay it out very you have to lay it

8:13

out very precisely like every word will

8:15

matter but you don't actually need that

8:17

much text you don't need that many

8:18

examples and they'll do a pretty good

8:20

job at the task um so eventually like we

8:23

didn't end up with a super complicated

8:25

system for the sidecar model it was very

8:26

simple um and the last thing was

8:29

evaluation so um evaluating memories in

8:32

our experience has been really tough

8:33

because it's the type of feature that

8:35

you notice when it's taken away not when

8:37

it's like necessarily there and So um

8:41

you can try and think of evals where

8:43

like the memories would help you get to

8:45

the solution faster but in some sense

8:47

it's kind of cheating because you come

8:49

up with the examples such that the

8:50

memories are useful and so when we

8:52

evaluated the memories sidecar model

8:55

like I focused it basically on the

8:57

quality of the me memory generated not

8:59

on the retrieval side and specifically

9:01

mostly just filtering out these task

9:02

specific memories. Um and so we ended up

9:05

in a place where we're pretty happy with

9:06

um the quality of memories that were

9:08

getting generated. And then the next big

9:10

question was kind of UX which is um how

9:13

much do you want to expose your memory

9:14

bank to your users and I think one big

9:17

learning has been what users think the

9:20

model should remember are not

9:21

necessarily what are useful and

9:23

especially like users get really frantic

9:26

and think that their memory bank has to

9:27

be perfect which is like so far from the

9:29

truth because in reality like even if

9:31

you give a model like 50% memories 50%

9:34

of them are junk 50% of them aren't

9:36

applicable like the models are smart

9:38

enough now to kind of filter about the

9:39

noise largely. Um, and so that was like

9:42

a really interesting thing. And so as a

9:43

result, we've kind of kept the

9:44

generation a bit hidden. And like

9:46

obviously you can go in and change it if

9:48

you don't like it. Um, but most of the

9:50

editing of memories happens like when a

9:52

model will choose to site a memory in

9:54

its generation. And then you can kind of

9:55

hover over and delete. But on the actual

9:57

generation side, we don't use it. Cool.

10:00

Um, but yeah, so it's kind of what I've

10:02

been talking about, but what's next? Um,

10:03

so so far now we can understand the

10:05

users. Um, but we want to learn a lot

10:07

more about your codebase. Um, and in

10:09

particular like we want to have an

10:11

understanding of the things that happen

10:13

in your codebase outside of just the

10:14

code because like you know we can turn

10:17

through and uh embed your entire

10:19

codebase and probably get a decent

10:20

representation of just what's happening

10:22

but there's a lot of things that you do

10:23

with your codebase that you know you

10:25

express in the sidebar like when you're

10:26

chatting with the agent but they don't

10:28

live anywhere. Um, and so we're trying

10:30

to figure that out basically. Um, going

10:32

to be a lot more prototyping and just

10:34

like you know shooting in the dark but

10:36

we'll I'm confident we'll get there. Um,

10:38

and then with that kind of new set of

10:41

memories where they can't just be

10:43

applied all the time in context, um, we

10:45

need new approaches to including

10:46

memories in context. And so, um, there's

10:49

some things we've been experimenting

10:50

there with changing the way that our

10:51

search pipeline works to include

10:52

memories. Um, but it's all just

10:54

prototypes so far. Um, and then kind of

10:57

the north star we're shooting for is

10:58

teamwide knowledge where um the if if

11:01

someone can learn from my mistakes um

11:05

that I made with the agent and they

11:07

don't their agent won't make that same

11:08

mistake, like that's the north star. But

11:10

it's also like really really important

11:11

to get it right because if you start

11:13

sharing bad memories teamwide, it's like

11:16

it's a huge degradation of quality. Um,

11:19

so this is what we're working towards,

11:20

but you know, we've still got some ways

11:22

to go. Uh, but yeah, I had a Q&A slide,

11:25

but I think we're gonna save the night.

11:27

Thank you guys.

11:31

[Applause]

Interactive Summary

Yash from the engineering team at Cursor discusses the development of 'memories' for their AI coding agent. The goal is to augment three types of context—directional, operational, and behavioral—to improve the agent's performance by learning from user interactions. Yash details the prototyping process, which involved testing 'sidecar' models and tool-calling approaches, while emphasizing the importance of avoiding task-specific, irrelevant memories. The team aims to move towards team-wide knowledge sharing, provided they can ensure high quality and prevent the propagation of bad information.

Suggested questions

4 ready-made prompts