HomeVideos

Continual Learning: How AI Agents Get Better With Every Use | Arjun Karanam, Trajectory

Now Playing

Continual Learning: How AI Agents Get Better With Every Use | Arjun Karanam, Trajectory

Transcript

671 segments

0:00

Okay, last talk to bring us home, we

0:01

have Arjun and Ronak, co-founders of

0:03

Trajectory. They're doing a bunch of

0:05

interesting research around what I think

0:07

is one of the hotter topics, continual

0:09

learning.

0:10

Please welcome Arjun to the stage.

0:12

>> Cool. What's up? How's it going

0:14

everyone?

0:15

Thank you so much. I'm really excited. I

0:18

know I'm close to the last one, so

0:19

hopefully there's like remaining

0:20

attention span.

0:22

I'm here. Ronak is also back there. So

0:25

we're together.

0:26

But let's get started.

0:29

We're Trajectory and we're building the

0:31

platform for continual learning.

0:35

Who are we? I think we're working on

0:36

like a really really cool mission and

0:39

that's the most fun part. The second

0:40

most fun part is I get to do this with

0:42

my two best friends. Ronak over here

0:44

worked at One Surf and Train Speed 1

0:45

before this and and Michael who worked

0:47

on really cool robotics stuff at at Deep

0:49

Mind. And what I wanted to start off

0:52

with was the world view that we have

0:56

and how that shapes

0:58

why we think this problem is important.

1:00

And so the world view that we have is

1:02

that and it's not a hot take, we are

1:04

living in an incredible time in human

1:06

history here. Every single week it seems

1:09

like another model is coming out that is

1:11

leapfrogging the last one.

1:14

It's undeniable that models are getting

1:15

better and better.

1:16

But we think that they're getting better

1:18

on one axis and that is

1:22

IQ. These models are smarter and smarter

1:24

but always feels like when you're

1:25

talking to them, it's their first day on

1:27

the job. And so you take this analogy if

1:30

you have a Terrence Tao in your pocket,

1:32

that's great.

1:33

But Terrence Tao day one at an

1:35

accounting firm is probably not the best

1:37

accountant there.

1:39

But give him a few years or honestly

1:40

maybe even like a few days,

1:42

he'd probably be really good.

1:44

And so we have this orthogonal axis of

1:47

experience that is much more important

1:49

than we think

1:51

in conjunction with IQ. And that's what

1:53

we're calling the experience gap and

1:54

that is what we want to close as a

1:56

company.

1:57

Um

1:58

And then but the question is like how do

2:00

you how do you close that? Where does

2:02

experience come from? Well, it's already

2:04

out there right now. There's so many

2:06

hundred million is a number, but so many

2:09

tokens out there that are being

2:10

generated by these agents. You view them

2:13

or sometimes you don't even view them

2:14

and they get thrown away. But this is

2:16

all real work that they're doing

2:18

and then people are acting upon them and

2:20

then it's being thrown away. And so our

2:22

take as a company as Gabe mentioned like

2:24

our opinionated take is that uh this is

2:27

a signal that we should be learning

2:28

from.

2:30

And this is also how humans get better.

2:32

And in classic AI fashion like if the if

2:34

humans do it probably a good way to

2:36

mental model off of. Uh and so today's

2:38

agents they're still slow, expensive,

2:40

error prone over time. You implement

2:42

them, they probably don't get better.

2:44

We want to imagine what learning agents

2:46

look like where as they are used by

2:49

people, they get better and better over

2:51

time.

2:52

Uh and tangibly this has two good

2:53

benefits. Right away it gets you much

2:56

faster, better, cheaper models. Uh but

2:58

more excitingly it gets you to this goal

3:00

of systems that compound with use.

3:03

And so I want to start off by talking

3:05

about how we're approaching this problem

3:08

uh and then

3:10

uh some interesting stuff for you guys

3:11

at the end.

3:13

And so where it starts off and this kind

3:15

of dovetails off of what Harrison was

3:16

talking about before is it starts off

3:18

with traceability. We need to first

3:20

capture these interactions, capture this

3:22

experience that's being thrown away.

3:24

So that's step number one.

3:26

Uh

3:27

And then step number two and this is

3:28

where where we start getting to our

3:29

research. The way we view the company

3:31

we're building is we're building a

3:33

product to allow for continual learning,

3:35

but we're doing cutting edge research

3:36

under every single one of them

3:38

to make this possible.

3:40

And so

3:41

you have all these interactions and the

3:43

next thing we're building is this idea

3:45

of a like a model spec. This idea that

3:47

like you need to have a way to define

3:50

what do I want my agent to do and have

3:52

the agent learn against that. So we're

3:54

doing really cool research here on how

3:55

to extract user interactions, turn that

3:58

into reward, extract traces, and turn

4:00

that into exact specs of what you want

4:03

your agent to do.

4:04

And then okay, cool. You have what you

4:06

want your agent to do. What do you do

4:08

with that? Well, there's two surfaces.

4:09

One is the models, right? You want your

4:11

models to to learn from actual

4:14

interactions here. And so we're doing

4:16

really cool research on algorithms like

4:18

SDPO,

4:19

and and and and things along those lines

4:21

where we're using RL to uh

4:25

take these full long traces and improve

4:28

models over time.

4:29

But models are not the only thing. You

4:30

have harnesses as well. And we're doing

4:33

a lot of really cool research on

4:35

when you have good feedback from people,

4:38

does that go to a harness, or does that

4:39

go to a model? An example here is

4:42

if it's like a fact, like this person

4:45

this like this company has been

4:46

delisted, right? You probably don't want

4:48

to train that knowledge into the model.

4:49

It's probably context that should be

4:50

available to the harness.

4:52

And so that is how kind of how we're

4:54

splitting things, and that is the

4:55

product that we're building. Allowing

4:57

you to go from real interactions all the

4:59

way to specs of what you want your agent

5:02

to get better at, to better models, and

5:04

then better harnesses, and then being

5:06

able to deploy those those models right

5:08

away,

5:09

and then start using our product.

5:11

And so that's the world we want to get

5:13

to.

5:14

If you have built agents, which I'm sure

5:17

most of you guys have,

5:18

you probably know that there are so many

5:20

things that go wrong with every single

5:23

part of these, and then how every single

5:24

one of these is an unsolved problem. And

5:27

so the way I wanted to frame the rest of

5:29

this is

5:31

I'm going to imagine that I am like

5:33

straight from Aladdin, and I have like

5:35

four wishes.

5:37

And what are those wishes that I would

5:38

ask to like make the agent ecosystem

5:41

better. And so like we have agents

5:42

today,

5:43

we have this like gap,

5:45

and there's a level where agents get to

5:48

and after that

5:50

we can really see the exponential

5:51

effects of continual learning. And what

5:54

are those in

5:56

I'll walk through them step by step with

5:57

the same four categories I talked about.

5:58

What we would like to see on the

6:00

traceability side.

6:02

What we would like to see on the eval

6:04

side, what we'd like to see on the

6:06

harness side, what we'd like to see on

6:07

the model side. So let's start with wish

6:08

number one.

6:10

Uh and how can companies get there

6:11

faster? Uh traceability.

6:14

So two two sub wishes here. Number one

6:17

is around tracing the entire tree of

6:20

what happens, uh including sub agents.

6:23

Uh what a lot of companies do is they

6:25

trace the main action that is happening,

6:28

but they throw away tool calls or sub

6:29

agents that are made. Um and that makes

6:31

it very hard to learn from the whole

6:33

thing that you've done. So that's wish

6:34

number one. Uh sub wish number two uh is

6:37

probably the most important in here and

6:39

that

6:40

you should build your product around how

6:43

you can both capture this interaction

6:46

data, but also elicit the right amount

6:48

of feedback uh so that you can then

6:50

capture it. Um and a really interesting

6:53

thing here is I think the level one way

6:55

of thinking about this is uh okay, let's

6:57

have like a thumbs up, thumbs down and

6:59

like let's capture that. And that sounds

7:01

amazing in theory uh but it's incredibly

7:03

noisy and if you've used any coding

7:06

agent, you know you kind of just like

7:07

accept everything that the agent does.

7:09

And it's only like five commits later

7:11

that you're like, oh crap, like

7:13

this broke everything. Let me go and

7:14

undo that. And so it's the corrective

7:17

behavior, the like the edits, the undos

7:19

and the retries that both need to be

7:22

like elicited from the user, but also

7:24

captured. So that's my wish on

7:27

traceability.

7:29

The second wish is around evals. And the

7:32

mental model to use here is

7:35

your the product that your user uses

7:39

should be as similar as

7:41

as close to

7:43

what the eval is done in in for wise

7:45

should be as close to where the training

7:47

is done in.

7:49

In an ideal world, these are all the

7:50

same. There's no difference between

7:52

these. And so this means that evals are,

7:54

you know, drawn from traffic, how people

7:56

are actually using your product, both

7:58

how they're using it now and the things

8:00

that they're requesting on the frontier

8:01

that might not be possible, all super

8:03

helpful.

8:05

Second is making every task roll

8:07

outable. This is a pretty big infra

8:10

challenge, uh but it's this idea that

8:12

hey, if a user has done one thing, can I

8:15

somehow replay what they what the user

8:17

has done? And that is an incredibly

8:20

helpful boon to to moving forward in

8:22

this agentic era. And then lastly, this

8:24

goes with number one, grading through

8:25

like the real harness, uh not making a

8:27

variation of the harness, but like

8:28

grading the real harnesses that people

8:30

use in production.

8:32

So that is wish number two.

8:35

Uh wish number three, getting into the

8:36

harness,

8:37

uh

8:38

pretty important. Uh number one is and

8:41

and I think the mental model to use here

8:43

is I think a lot of products

8:46

have built their harnesses around models

8:49

that existed a year a year and a half

8:51

ago. Where the primary function of the

8:54

harness was to

8:56

prevent the agent from doing bad things.

8:59

Which was really good when agents would

9:01

randomly break or have misformatted

9:03

outputs, but now we're very much in a

9:05

let the agents cook world.

9:07

Uh and so

9:09

I think I would view building the

9:10

harness as not enforcing specific flows

9:13

that need to be done, but more what are

9:15

the primitives that your product has,

9:18

whether that's search tools or private

9:20

information, and then viewing the agent

9:22

as orchestrating those primitives and

9:24

then letting it do that.

9:26

So that's sub wish one. Uh sub wish two,

9:28

and this is how we're building our

9:28

product as well, is making the agent

9:31

interface as close to the user interface

9:32

as possible.

9:34

Uh in an ideal world, every single thing

9:37

that I can do on your UI, your agent can

9:39

do via tool call as well.

9:41

Uh I think that is a world that makes

9:43

training and continual learning so so

9:45

much easier.

9:46

Uh and then lastly, this is a small one,

9:48

uh making tool responses informative. I

9:50

think it's easy to be like, "Okay, I

9:52

just want to know the state." Like let's

9:54

say I'm like calling a search tool. I

9:55

want to know the state. I want to see if

9:56

it worked. Uh and if it's writing to a a

9:59

DB or something, it just writes to it.

10:01

And what we see often is like the

10:03

response from that tool call will be uh

10:05

done. Finished.

10:08

That sounds like good in theory, but

10:09

from an agent's perspective it's like

10:11

incredibly confusing. And if you're

10:13

trying to train or learn off of that,

10:15

there's like no signal to learn off of.

10:16

Right? You have no idea what actually

10:17

was written. You have no idea what was

10:19

actually was read. And so, that's the

10:21

last sub wish under the harness is to

10:22

make your tool responses really

10:24

informative.

10:25

So, that's wish number three. I think

10:27

like canonically I only get three

10:28

wishes, but let's imagine I have four.

10:30

Uh

10:31

and that's around the models.

10:34

Uh and we've talked about that a lot

10:35

today. Uh and that is

10:38

uh

10:39

I wish it was as easy as like just

10:41

switching to an open weight model, but

10:42

if you've tried that, you've probably

10:44

seen there's 50 other considerations

10:46

from security to safety to

10:50

uh you know, different access

10:51

provisioning stuff. So, start getting

10:53

comfortable running on open weights cuz

10:56

that is what unlocks the door to owning

10:59

owning your weights and then continually

11:00

improving on top of them.

11:02

Um and then Gabe also talked about this

11:04

this idea that like model routers are

11:07

uh going to probably play a really large

11:10

role in routing intelligence to the

11:14

exact capability of the task they're at.

11:15

And so, experimenting with with routers

11:17

is another wish I would wish cast into

11:19

the world.

11:20

And so,

11:22

that is the high level here. And so, I

11:24

think like

11:25

when we work with companies, the first

11:27

thing we often do is do an audit of of

11:29

these and like where things are. And

11:32

for the most part there's a lot of work

11:34

to be done to get here and so we work

11:35

with them to build it out and it would

11:37

be my dream if we go to a company and

11:41

they have a lot of this built out.

11:42

Granted a lot of the work we're doing is

11:44

also knowing that a lot of this we can't

11:45

expect and so we're doing a lot of cool

11:47

research on, you know, assuming that

11:49

nobody has Evals, somebody doesn't have

11:50

Evals, how can we overcome that? Or

11:53

assuming that the traceability is off,

11:56

like how can we learn from the signals

11:58

that are there? And so that is

12:02

that that is my my goals for the for the

12:03

universe here.

12:05

Uh just want to quickly go over

12:07

uh I think this is like a very very core

12:09

research and product problem and so

12:12

we're really excited with the team we've

12:13

put together but I think more

12:15

importantly and then two of the folks

12:16

who are here today

12:18

uh we're working with extremely frontier

12:22

customers. The folks who are really

12:23

really pushing the boundaries of what

12:26

agents can do in production

12:29

both today and tomorrow.

12:31

Uh and so really excited to be working

12:32

with them and another thing I want to

12:35

show is

12:36

we have

12:38

Oh, that's my email.

12:39

Uh

12:40

we have a beta and so

12:43

a really really core belief we have is

12:45

that

12:46

this capability, this ability to own

12:48

your own intelligence, should not be

12:51

something that you need to consult away

12:53

a lot of the times. It should be

12:54

expertise you build on your own because

12:56

of how important it is for your product.

12:59

Um and what that means for us is we want

13:01

to build a product that empowers any

13:03

company to

13:05

own their own models and mainly get to

13:07

continual learning from real

13:08

interactions and so here's some some

13:11

cool screenshots from our product. I'm

13:13

I don't know if Nico is here but I'm

13:14

supposed to show this Nico later today

13:15

so this is like a sneak peek. Uh

13:18

but here's Harvey. Here's the lab

13:19

benchmark that that Gabe was talking

13:20

about. Really really easy to import. Um

13:23

training a model is incredibly easy.

13:25

We're extremely proud of the interface

13:27

we've put together here cuz if you've

13:28

post trained, you know, there's like 50

13:30

million things that go wrong, 50 million

13:32

knobs to turn, and it's all researcher

13:34

intuition, trademark asterisk, whatever

13:37

that means. And what we've tried to do

13:40

is

13:41

take all of how our researchers post

13:43

train their models and

13:45

use an agent behind the scenes to take

13:48

care of most of those knobs and just

13:49

only expose the ones you need to know to

13:51

make post training as easy as possible.

13:54

So, really easy to train a model, see

13:55

how the model's doing, evaluate,

13:58

compare it,

13:59

uh see if it's better than the model you

14:01

were using before, and then deploy it.

14:03

And so, in terms of actual work you put

14:05

in, and that was not waiting for the

14:06

model to train,

14:07

probably like 15 minutes. And so, that

14:10

is the world we're we're trying to get

14:11

to to empower all these companies and

14:13

hopefully more folks in this room.

14:15

Uh

14:16

So, yeah, that's our goal. We want to

14:18

every company to own its own experience

14:19

layer, and we're excited to build that

14:21

alongside of some of y'all here,

14:22

hopefully.

14:27

>> [applause]

14:28

>> Thank you. It's a great talk. You know,

14:30

I wonder like, you know, uh when you

14:32

talk about continual learning, like uh

14:34

what do you think about their trainable

14:36

object

14:38

um between model weights, harness,

14:41

tools, application layers?

14:43

>> Yeah.

14:43

>> Uh how how do you think of them? How do

14:45

how do you prioritize?

14:46

>> Yeah, I think I like a fun fun phrase is

14:49

that if you ask like six researchers

14:50

what continual learning is, you're going

14:52

to get like

14:53

seven answers, probably. Uh

14:56

you can be very pure about this and be

14:58

like, "Oh, it needs to be like a human,

15:00

and it is just the weights that need to

15:02

be adopting in real time, one shot."

15:04

But, the way we view it is that

15:08

the intelligence that your product is

15:09

run off of is a system.

15:11

It's a system that has many components,

15:13

and true continual learning

15:15

is something that optimizes across that

15:17

system

15:18

and

15:20

optimizes based on what parts of the

15:22

system needs to be updated uh for the

15:24

information that you're learning. Like a

15:25

mental analogy that I have is like when

15:28

you are you know, saving things, right?

15:31

We no longer think of like where in your

15:33

RAM or where in your hard disk to save.

15:35

That's the level that's abstracted based

15:36

on what makes the most sense. We think

15:38

about models versus harnesses versus

15:40

context in the same way. It feels wrong

15:44

that we're having to make the decision

15:45

off of very little priors

15:47

what to update based on what.

15:50

This is a scientific problem that can be

15:52

solved. Let's do that and then let's

15:54

abstract that away.

16:03

Cool.

16:06

Oh, yes.

16:07

Yes.

16:07

>> Um, some of the speakers today talked

16:09

about, you know, the importance of not

16:10

training on your customer data.

16:12

>> Yeah.

16:12

>> And I imagine like a big part of

16:13

continual learning is you are learning

16:15

from all these interactions. Um, how are

16:17

you thinking about things like

16:18

differential privacy? Like how do you

16:19

actually make your system better with I

16:22

imagine most of these application

16:23

companies have customer specific data

16:25

arrangements that make that very hard.

16:27

>> 100%. I think that's an incredibly

16:29

important problem. Uh,

16:31

actually before this I was at Apple and

16:32

this is a problem we worked on quite a

16:34

bit as well.

16:35

Um, and there's a lot of clever things

16:37

you can do where you're not actually

16:40

training on customer data, you're

16:42

instead sampling distributions from your

16:44

customer data and then synthetically

16:46

generating your own and doing kind of

16:48

like a

16:49

mental models maybe like a cryptographic

16:51

thing where you're like comparing those

16:53

distributions to see is my data actually

16:55

on distribution, but you're not training

16:57

directly off of the customer data. Um,

17:00

and so that that's totally right.

17:01

There's a lot of interesting work there

17:03

and it's like very central and already

17:06

with our customers we're doing some

17:07

interesting and and fun things in order

17:08

to

17:09

uh, to get over that.

17:12

Yes.

17:13

>> Hey, I'm Sid. I'm from Finch. Um,

17:15

thank you.

17:16

Um,

17:17

I think episodic memory plays uh, a role

17:20

in continual learning to some extent.

17:22

For example, like if a user does

17:23

something on our platform that corrects

17:25

what an agent did before, we might want

17:27

to update an agent's behavior. Does

17:29

Trajectory have like an opinionated take

17:31

about that in the platform?

17:33

>> Yeah, I think

17:35

Okay, on episodic memory and I'm just

17:38

going to repeat here.

17:39

On episodic memory and like what if I'm

17:41

thinking about this correctly, how

17:43

different types of feedback impact what

17:46

is learned?

17:47

>> Yes.

17:47

>> Yes, okay. So many thoughts.

17:51

One way to look at this is there are

17:55

two different types of signals. There

17:57

are signals that are just

17:59

something went wrong. Like for example,

18:01

if somebody like flames the agent says

18:03

like, "You're really bad." But he

18:05

doesn't follow anything up with that or

18:07

just thumbs down or like drops off the

18:08

session. Those are cases where you know

18:10

something went wrong, but not what right

18:13

looks like. Then you have things like

18:15

the agent retried and then it got to the

18:17

correct solution or you corrected it by

18:19

saying this is a difference. And so that

18:21

is like a

18:22

difference we make. In the latter case,

18:23

we're very, very confident about the

18:25

reward associated with the with the

18:27

output and and what we're assigning.

18:29

With the former, we know to penalize

18:31

that behavior, but we don't necessarily

18:33

say like there's like a correct right

18:34

answer. That's one way of looking at it.

18:37

Another way of looking at it is there

18:40

are different

18:42

like there's like a hierarchy of like

18:44

how pertinent the information is

18:46

people-wise. There's some information

18:48

that's probably globally accurate,

18:50

globally true.

18:51

Things like a tool call repeatedly

18:53

failing when trying to call a certain

18:55

tool.

18:56

That's probably relevant to everybody.

18:58

And so this should be trained into the

19:00

model and the model should get better at

19:01

learning this tool.

19:02

A certain user is like, "I never want to

19:06

use the sub agent. Please, please,

19:07

please don't do it."

19:08

That's probably not something you should

19:10

train into the model and leave to the

19:12

context, and

19:14

what we were really excited about is

19:16

this is the stuff Harvey was talking

19:17

about where this is actually something

19:18

that will probably happen on a per org

19:20

basis, or per customer even further, a

19:23

per customer basis. And so, that's

19:25

another way to think about it as well.

19:26

So, the hierarchy of like what the

19:28

feedback is pertinent to.

19:30

>> Got it. Thank you.

19:31

>> Yeah.

19:35

Yes.

19:36

>> Thank you.

19:37

I think one of the themes of today has

19:39

kind of been this idea that there are

19:41

some, you know, tasks or part of your

19:43

product where you kind of want to maybe

19:45

like experiment more, go cheaper, add in

19:47

more open source stuff, right? And then

19:49

other areas where kind of being at the

19:51

frontier and using a closed model or a

19:52

closed framework is more useful. I guess

19:54

like extending that analogy to continual

19:57

learning, are there tasks or workflows,

20:00

or like what is the shape of problem

20:02

where something like this you've kind of

20:03

seen really matters a lot, right? Versus

20:06

where like just using, you know,

20:09

statically trained frontier model or

20:11

even a statically trained open source

20:12

model, and then like you know, using the

20:14

harness to inject relevant information

20:16

into the context is sufficient to kind

20:18

of achieve what the user usually wants

20:19

to achieve.

20:20

>> Yeah.

20:21

This is a great question. I think the

20:24

level one answer is I think for the most

20:25

part most tasks work for continual like,

20:27

but the ones I'm most excited about are

20:30

tasks

20:32

that are on the frontier. The mental

20:35

model I have that that we have for

20:39

AI progress in in in a way is

20:42

people will ask models, people will ask

20:44

products

20:45

with the expectation of what they think

20:47

the product can do.

20:49

Uh and maybe they'll ask at the edge of

20:51

what the product can't do, and then

20:54

they'll see it fail, or they'll see it

20:55

do something wrong, and they'll retreat

20:57

back and be like, oh, it's it wasn't

20:58

good enough. I can't use it for this.

21:01

I think like a great example is like,

21:02

would I even dream of typing in some of

21:05

the like the whack query that give the

21:06

cursor two years ago, like definitely

21:09

not. I was like way lower than this, but

21:11

then over time I started querying it

21:13

where the models were getting better and

21:14

I was like, "Okay, maybe we can do

21:15

bigger and bigger things."

21:17

What we're most excited about is

21:19

the way that has happened is like models

21:20

have generally gotten better and we've

21:22

been targeted pushing in certain

21:24

directions. Uh but what we're seeing

21:26

with some of our customers are cases

21:27

where

21:28

users ask for things, the model can

21:30

barely do it, but then during training

21:33

it learns how to do it and then the user

21:35

can now do this thing they couldn't do

21:36

before. That chain

21:39

is what's really exciting about

21:40

continual learning of like really

21:41

pushing that frontier

21:43

of what's possible based on what the

21:44

user tries and cannot do.

21:47

>> Awesome.

21:48

Thank you, Arjun. I think we're out of

21:49

time. Uh really appreciate you doing

21:51

this. It's awesome. Thank you.

21:52

>> [applause]

Interactive Summary

Arjun, co-founder of Trajectory, discusses the importance of 'continual learning' for AI agents. He argues that while current models are improving in IQ, they lack the 'experience' necessary to be truly effective assistants. Trajectory aims to bridge this experience gap by creating a platform that allows models to learn from real user interactions, effectively compounding their performance over time. The talk outlines a framework for this, covering traceability, evaluation, harness design, and model improvement, while also addressing challenges like data privacy and episodic memory.

Suggested questions

3 ready-made prompts