HomeVideos

The Man Behind LangChain Memory | Will Fu-Hinthorn

Now Playing

The Man Behind LangChain Memory | Will Fu-Hinthorn

Transcript

324 segments

0:00

[Music]

0:09

[Applause]

0:14

Thank you. Thank you. Thank you, Nicole.

0:16

And thank you for Greg for organizing

0:18

this and Matt and thank you all for

0:20

being here. It's kind of crazy to see

0:21

how much excitement there is for this

0:24

ambiguous and and broad topics we call

0:26

memory. Um, so I'm Will, uh, engineer at

0:29

Langchain. One of the honors that I get,

0:32

uh, at Langchain is to work with a lot

0:33

of our customers and users in building

0:35

successful memory and context systems.

0:38

Um, and so this set of slides is just a

0:42

few takeaways, a few lessons that I've

0:44

learned, um, or a few lessons that I've

0:46

had reinforced over the past year about

0:49

sort of what it takes in order to build

0:50

a successful memory system. Um, and if

0:53

you take nothing else from this, it's

0:54

that it's really hard to make a general

0:56

purpose one. They're best if they're

0:58

focused on your specific application and

1:01

to always look at your data and have an

1:03

experimental approach.

1:05

So, first we're going to roll it back. A

1:07

little over a year ago, we launched our

1:09

first memory server. Um, it was a server

1:12

designed to take in interactions between

1:14

a user and an agent or a chatbot, and it

1:16

would extract derived uh insights about

1:19

users that you could then later use. Um,

1:22

we launched Lang Friend, which is this

1:23

journaling app, a demo in order to show

1:25

how to use it. And we had a bunch of

1:27

enterprise design partners.

1:29

Over the subsequent months, we worked

1:31

really closely with these teams and

1:34

learned that, you know, first of all,

1:36

memory is very hard. Um, and and had a

1:38

bunch of lessons that I'm going to try

1:41

to share to you today. Um, that caused

1:43

us to sort of rethink how we wanted to

1:45

share a lot of these tooling and a lot

1:47

of these approaches for memory. The

1:49

first being that there's no really real

1:51

one-sizefits-all memory solution. Um,

1:55

and if you think about it for more than

1:56

a few seconds, this this kind of makes

1:58

sense. Um, you know, memory that we call

2:00

touches on a lot of downstream jobs to

2:03

be done or skills. Um, we need to be

2:05

remembering facts, information

2:07

relationships, this deeply

2:08

interconnected web of information that

2:10

we need um to know. And an agent also

2:13

expects to know all this information in

2:14

order to perform well. Um, but that's

2:16

not all. It needs to remember all of

2:18

that situated in a temporal context. It

2:20

needs to have more like episodic memory

2:22

um in order to remember these

2:24

experiences. And it also needs to be

2:25

understanding and able to learn, you

2:27

know, procedures and new information in

2:29

order to execute on particular tasks.

2:31

Not all of these memory types are

2:33

necessary for every application. Um this

2:35

is also non-exhaustive. Um but these are

2:38

some of the different types of ways um

2:40

that we see people wanting to be

2:41

learning new types of things in their

2:43

application. Um and another

2:45

reinforcement of this is you can see the

2:47

different types of things that people

2:48

want to be learning in real

2:50

applications. Uh so Nicole mentioned SHA

2:52

GBT u launched their updated memory

2:55

implementation recently. Um and they

2:57

really catered this around their

2:59

particular UX and use case. They have

3:01

assistant response preferences so it can

3:03

try to learn sort of your preferred way

3:04

of it learning uh of responding to you.

3:06

Uh in general it has some topics that

3:09

it's trying to model out. It's got

3:11

useful insights about the user um that

3:13

may or may not be useful depending on

3:14

the context. It has summaries of recent

3:17

conversations to sort of extend that

3:19

conversational buffer into a sort of

3:21

compressed representation. So it has

3:23

more temporal context there. And then

3:24

has a bunch of random interaction

3:26

metadata that it considers to be

3:27

relevant depending on the use case. Um

3:30

contrast that with something like an

3:31

email assistant where you really want to

3:33

be focused on a particular task. You

3:34

want to be focusing on the style and

3:36

content that you're going to be putting

3:37

into it. Whenever it's scheduling, you

3:39

really need to know what you prefer,

3:40

what you don't where you're available,

3:41

all that kind of stuff. Um, you need to

3:43

be having a deep understanding of

3:45

different sort of procedures. If someone

3:47

is emailing you and reporting a security

3:49

event in event, um, and you need to

3:51

follow all of that versus um, something

3:53

where it's just trying to meet up for

3:54

coffee.

3:56

All these things you can try to program

3:57

ahead of time, but if you want your

3:58

agent to be able to learn this, you

4:00

might not get too far if you're just

4:01

reaching right off the shelf with them.

4:05

Uh the second lesson is that you know

4:07

updates in in validation are especially

4:08

errorprone. You know it's one thing to

4:10

extract what you think is particular

4:12

interesting about a uh bit of a

4:15

conversation or a document. You know LMS

4:16

are good as summarizers. It's even

4:18

harder to then synthesize that and

4:20

connect that to all the existing

4:21

knowledge and domain that you have in

4:23

order to create a consistent world model

4:25

that improves your prediction

4:26

downstream. Um and so a lot of this ends

4:29

up being application specific. This

4:30

reinforces our lesson too. Uh we've

4:32

released a couple of tools um for this.

4:33

There's things like trust call to help

4:34

with like updates so your LMS aren't

4:36

willy-nilly deleting things. Um we've

4:38

released some like other tooling and

4:40

example around it. But at the end of the

4:41

day, a lot of this is about how you're

4:44

presenting and mapping your memory

4:46

context to the information architecture

4:48

or the needs of your application.

4:50

Um lesson three is that it's part of a

4:52

broader context system. Memory isn't

4:55

just this one isolated thing where

4:56

you're learning about the user

4:58

preferences. this is a part of you know

5:00

the situation in which the agent finds

5:01

itself be that your codebase be that the

5:04

um you know recent events the user's

5:06

gone through and other sorts of things.

5:08

Um, and when we had originally built our

5:10

memory server, we had really focused and

5:13

leaned in on to what we could

5:14

differentiate. But a lot of the people

5:16

we worked with wanted to then integrate

5:17

this with all the other existing

5:19

predictions, preferences, and other data

5:21

that they have for their model. And

5:23

while you can synthesize this all at

5:24

prediction time or retrieval time, um,

5:26

it's hard to create a sort of coherent

5:28

worldview over everything user engages

5:31

with or everything the a engages with if

5:33

you're not treating this more

5:34

holistically.

5:36

Uh lesson four and this is one thing

5:37

that we already sort of believe but we

5:40

especially believe now is that memory is

5:41

really software not necessarily

5:42

hardware. Um we often have people come

5:44

to us and talk about oh I need a graph

5:47

DB or I need a very particular

5:49

instantiation of memory for this when

5:51

really it's backwards. You need to start

5:52

from what you're trying to solve. And we

5:54

sort of made this mistake in that we had

5:56

a very particular memory server with a

5:58

particular instance. We got great

5:59

feedback on that. We found that wasn't

6:00

sufficiently flexible for people um

6:02

depending on the deployment context.

6:05

So all of these insights drove our

6:08

creation of lang SDK

6:10

and we organized that around three main

6:13

principles. One is we wanted to support

6:15

easy customization and experimentation.

6:17

All the actual extraction runs in the

6:19

code that you define. Um we have a

6:21

number of primitives to try to address

6:23

some of these memory types that we've

6:24

talked about before. Um but we really

6:26

want you to be taking ownership over the

6:29

type of information how it's updated. We

6:30

want to support flexible storage and

6:32

organization of memories. Um, so you can

6:33

sort of define what makes sense for you

6:35

and we aren't going to lock you into a

6:36

particular like vector database or

6:37

anything like that. You can swap it out.

6:39

Um, and we wanted flexible processing as

6:41

well. A lot of people talk about

6:42

different sort of, you know, batch

6:43

versus online, all these types of

6:44

different ingestion methods. Wanted to

6:46

make sure that we weren't really only

6:48

speaking to one particular thing in this

6:50

SDK.

6:52

Um so to go a little bit deeper on the

6:55

first topic of customization

6:56

experimentation we have some primitives

6:58

that we included in the library um or

7:00

the memory types that we found that

7:01

people often lean towards. So one is

7:03

this learning of knowledge and facts. We

7:04

have this background memory manager

7:06

where you can really customize in terms

7:07

of instructions the steps they're taken

7:09

how to reconcile um information and all

7:11

this is sort of orchestrated in your own

7:13

code. Um, another way that we have it

7:16

that's not shown here is we just have

7:17

some simple tools and you can be

7:18

defining it with an agent. And so as

7:20

models get better, if you want to treat

7:22

this as purely reasoning over things,

7:23

you can be putting that.

7:26

Um, we have learning instructions or

7:28

code books or however you'd like to call

7:30

that complex workflows. This is very

7:32

similar to prompt optimization. It's

7:34

typically not the whole prompt. Um, but

7:36

you have different sections of the

7:37

prompt where you're going to have

7:38

instructions that are custom to the

7:40

user. Uh and this is especially like

7:42

data dri driven where you want to be

7:43

going over batches of of conversations

7:45

uh extracting insights from them looking

7:48

at explicit user feedback and then

7:49

incorporating all that all into updates

7:52

that you can then measure.

7:54

I think if you put all of this into a

7:57

very generic graph and everything kind

7:58

of lose out on a lot of the ability to

8:02

elicit the proper responses that you

8:04

want. if you want to look at things like

8:05

tone, if you want to look at learning

8:06

new capabilities or looking at more

8:08

complex multi-step interactions that you

8:10

actually um need your agent to be

8:12

learning.

8:14

Uh and finally, there's also plenty of

8:15

episodes. This is probably the least

8:16

supported in the me in in the the

8:18

library. Um but we let you define

8:20

synthetic fshots to be extracted and

8:22

then you can incorporate it down.

8:24

The second principle we wanted to

8:26

support was about flexible storage and

8:27

organization. Um so we allow you to

8:29

store it in pretty much any back end

8:30

where you have a get input method um and

8:32

a search method. And so we have all

8:33

these integrations as well. One thing as

8:35

a tangent that we're kind of excited to

8:37

test out a little bit more is this file

8:38

system as a back end or at least

8:40

exposing a file system like API. Um and

8:42

the reason for that is a lot of these

8:44

coding or a lot of these LLMs and these

8:46

foundation model companies are really

8:48

optimizing for software engineering. And

8:50

so one hypothesis that we're testing out

8:52

um we'll get back on more results in a

8:54

bit is uh that perhaps they'll be more

8:56

lending them to um managing memories as

8:59

if it were a file system. Uh and so you

9:00

know that that's the flexibility of this

9:03

uh library allows you to be

9:04

experimenting with a lot of these things

9:06

as LLM tend to lean into different

9:07

directions over time. Um you can

9:11

organize the memories by user agent role

9:14

um organization and all those types of

9:15

things. So you're not locked into a

9:16

particular like only organizing things

9:18

around a user. You can have agents learn

9:19

things just for themselves and share it

9:21

across u and this is all sort of

9:22

orthogonal to the actual process of

9:24

information processing. Um and you can

9:27

process how and when you like. We have

9:29

some instruction around being able to

9:31

execute this on a deferred basis. So you

9:32

can be batching all of this later on if

9:35

you have a really rapidly occurring

9:37

interaction. You can have it actually

9:39

process in real time or you can delay it

9:41

so that there's dduplication of

9:42

information. Um and all this can be

9:44

managed either through a local executor

9:45

or um online with line graph platform.

9:48

So it can be scaled in a very

9:49

horizontally scalable way. Um and that

9:53

concludes my talk. I guess here's a

9:54

summary slide of all the things that I

9:55

wanted to share. Again, if you take

9:57

nothing else, it's that really no one

9:59

sizefits-all and you really want to

10:00

start with what you hope to accomplish,

10:02

what you hope your agents learn and then

10:04

work your way backwards to pick the

10:06

right solution for you. You know, test

10:08

things out, but don't just be willing to

10:10

offload all of this to one particular

10:12

solution that claims to be a holy grail.

10:14

Um, we've put Langmama SDK out there.

10:16

Encourage you to experiment on it,

10:17

encourage you to give us feedback. Um,

10:19

we're always looking to improve it. Um,

10:21

but I think, yeah, would encourage you

10:23

to take ownership of that. So, thank

10:24

you.

Interactive Summary

Will from Langchain explains the challenges and principles of building successful memory systems for AI agents. He emphasizes that there is no 'one-size-fits-all' solution; developers should focus on their specific application requirements, adopt an experimental approach, and treat memory as software rather than hardware. The talk introduces the Langchain memory SDK, which is designed to support flexible customization, storage, and processing, allowing developers to tailor memory systems to their agents' specific needs.

Suggested questions

3 ready-made prompts