HomeVideos

How Harvey Built a Research Lab on a Budget | Gabe Pereyra

Now Playing

How Harvey Built a Research Lab on a Budget | Gabe Pereyra

Transcript

847 segments

0:00

Next up, we are honored to have Gabe

0:02

with us. Gabe is co-founder and

0:03

president of Harvey. Um you were a

0:06

research scientist at DeepMind ages ago,

0:09

um and then at Meta I think before you

0:11

showed your college roommates uh

0:13

what GPT-3 could do. And this duo uh

0:17

then became Harvey. Um we're really

0:19

excited to have you here. I think

0:21

everybody here in the audience is an

0:22

application company um thinking about

0:24

how to start doing their own research,

0:26

post-training their own models, um and

0:28

uh create their own labs. And so I think

0:31

Harvey has really set the example here,

0:32

and we're really delighted to have you

0:33

give a talk on how to how you guys built

0:35

Harvey Labs.

0:37

>> Awesome.

0:38

>> [applause]

0:44

>> So the alternative title to this talk is

0:48

building a research lab on a budget.

0:51

And

0:52

it's an unfair game competing with the

0:54

frontier labs if you're an application

0:57

layer company.

0:59

There are rich teams, there [snorts] are

1:01

poor teams,

1:02

and then there's us in the application

1:03

layer.

1:06

The frontier labs have more money, more

1:08

talent, more compute infrastructure, and

1:11

data.

1:12

So how do you compete?

1:16

By using the frontier ecosystem.

1:19

When we started Harvey 4 years ago,

1:22

most of these companies either didn't

1:24

exist or were just getting started.

1:27

And so we either had to build everything

1:29

ourselves

1:30

or in most cases focus on something

1:32

different. Like building our GTM org and

1:35

a great product.

1:37

But today

1:39

using the frontier ecosystem

1:41

I think you can compete with the

1:43

frontier labs and build frontier

1:45

intelligence.

1:47

This talk is going to be our high-level

1:49

playbook for doing this.

1:51

And I'm going to talk about how we build

1:53

benchmarks and training data,

1:55

how we work with the Neo Labs to do post

1:58

training

1:59

and how we serve these models in

2:01

production.

2:04

To start,

2:06

you want to build a benchmark. If you

2:07

don't have a good benchmark, you can't

2:09

train models, and if you can't train

2:11

models,

2:12

you don't need to serve them in

2:13

production.

2:14

So, this year we released three of these

2:17

data sets that we built.

2:19

We started by building Legal Agent

2:21

Bench, which is a taxonomy of tasks that

2:25

associates would do at a large law firm.

2:28

And

2:29

these cover multiple practice areas, and

2:31

they're complex tasks like doing

2:34

drafting complex fund formation

2:36

documents, doing case law research,

2:38

things like that.

2:40

We followed this up by building a

2:41

contracting data set. This teaches This

2:44

allows us to teach agents to do

2:46

negotiation like you would in an

2:47

in-house department.

2:49

And then the one I'm most excited about

2:51

that we recently released is a large

2:53

diligence data set.

2:55

This is, I think, one of the largest RL

2:57

environments that's been released. The

2:59

largest data rooms here are 80 million

3:01

tokens, and it lets us do research on

3:04

long context, very complex tasks.

3:08

And I think the most interesting thing

3:10

about these data sets is how we built

3:12

them.

3:14

One challenge we always had at Harvey

3:16

for training models is we can't train on

3:19

our customers' data. We work with the

3:21

largest law firms and enterprises,

3:24

and their legal data is incredibly

3:25

sensitive. It's privileged, and so you

3:28

can't put it in generic models. We can't

3:30

even put it in our models. And so, how

3:33

do you train models given that?

3:36

And the thing that started working

3:38

really well this year

3:39

is using domain experts to guide

3:43

synthetic data generation.

3:45

And

3:46

uh Brendan from Reka had a good analogy

3:48

that the same way that engineers now

3:51

don't write code, they vibe code and

3:53

guide these coding models. We're

3:55

starting to do the same thing. And so,

3:58

my younger brother is actually a lawyer

4:00

at Harvey, and he's gotten very good at

4:02

using the coding models. He's trained

4:04

our other lawyers to do that, and they

4:05

can generate incredibly realistic data

4:08

sets that we can use for training and

4:10

also evaluating our product.

4:14

And once you've done that,

4:16

synthetic data isn't good enough, but

4:19

it's a way to get started.

4:21

And so, we work with companies like

4:23

Mercor and Snorkel, who let you scale up

4:26

this process and build larger sets

4:29

particularly for training.

4:33

Once you've done that, you need to turn

4:35

these data sets into efficient RL

4:37

environments. It gets really expensive

4:40

as these data sets get larger and

4:42

evaluation gets very expensive. For

4:44

example, in our diligence data set, we

4:46

have over 1,000 unit tests that are

4:49

grading model outputs using LLM as a

4:52

judge. If you use the largest models and

4:55

you want to do RL rollouts and things

4:57

like this, it gets very expensive. And

4:59

so, there's a lot of work, and here's

5:01

some we did with LangChain, of making

5:03

these very efficient.

5:06

And then the last thing we did that I

5:09

think was a little controversial at the

5:11

time, was open-sourcing some of these

5:13

data sets. And the motivation for this

5:16

was

5:18

it's very hard to know your data set is

5:20

good unless a lot of people train on it.

5:23

When I used to do research at Google

5:25

Brain and DeepMind, the best data sets

5:28

were open, like ImageNet, CIFAR, MNIST,

5:31

and everyone used them, and you were

5:33

able to find all of the issues, and we

5:35

get a ton of pull requests, we get

5:37

suggestions. Um

5:39

and then I think increasingly we're

5:41

having the labs when they report new

5:43

models benchmark on our data set. And

5:46

then most importantly, we had Elon

5:48

retweet it.

5:52

So, once you've built the benchmark, now

5:54

you have something to train models

5:56

against.

5:57

And

5:59

the thing that is exciting now is

6:01

open-source models are getting

6:02

competitive. In the past, it wasn't

6:05

worth doing post training because the

6:07

models were improving so quickly from

6:09

pre-training that any post training you

6:12

did quickly got absorbed by the next

6:14

pre-trained model.

6:15

But now with models like Kimmy 3, GLM

6:18

5.2, NeMo-Megatron, Inkling, and others,

6:21

it's possible to take these very strong

6:23

open-source base models and post train

6:26

them to levels of frontier intelligence.

6:29

Maybe not general frontier intelligence,

6:31

but if you have a specific task like us,

6:34

they are competitive.

6:36

And so, the way we recommend getting

6:37

started is working with the NeMo labs.

6:41

They have a bunch of expertise and

6:43

infrastructure already in place to help

6:46

you make sure that your training data

6:48

sets are good. They have recipes. And

6:50

usually, if you work with them and

6:53

you're not able to get better results,

6:55

there's probably something you're doing

6:56

wrong with your data set, and this is a

6:58

very good way to bootstrap it.

7:01

And so, some of the interesting work we

7:03

did with these different providers, um

7:06

Fireworks, we got some very interesting

7:08

results training GLM 5.1 to use Fable or

7:12

maybe Opus 4.8 as an advisor model.

7:16

Um Base 10, we did some interesting work

7:18

on KB compaction. N gram, who I think is

7:21

here, we're doing interesting work on

7:23

enterprise search and firm knowledge.

7:26

Trajectory, um we worked with them to

7:28

train NeMo-Megatron models. And Applied

7:30

Compute, we're doing some interesting

7:32

work on our Vault product.

7:34

And I think one question we got is why

7:37

work with multiple NeMo labs. Why not

7:39

just pick one?

7:41

And for us, as we're scaling the

7:43

research lab, we have more research

7:45

projects than we have bandwidth to do

7:47

internally or just with a single Neo

7:49

lab.

7:50

And every Neo lab is taking a different

7:53

bet. They have different ways they think

7:55

about research. We have different open

7:57

source models we want to train. And the

8:00

more we work with, the more we learn.

8:04

And it's getting easier than ever to do

8:06

post training. So one, working with the

8:08

Neo labs, we're learning a lot in

8:10

partnership with them.

8:12

And then we're doing more and more post

8:13

training ourselves internally with APIs

8:16

like Tinker and the infrastructure

8:18

Fireworks and Base 10 have built. It's

8:20

never been easier to post train these

8:22

models and then serve them. And there's

8:24

increasingly more post training talent

8:26

that we're hiring and is available.

8:29

And inspired by Cursor, the goal of

8:32

these efforts is for us to build our

8:34

version of Composer one. How do we

8:37

package all of the work we've done with

8:39

synthetic data, scaling it with Mercor,

8:41

the work with the Neo labs, into a model

8:44

we can serve alongside the closed source

8:45

models.

8:49

Now, once you've post trained a model,

8:51

you need to be able to serve it in

8:53

production.

8:54

And this is non-trivial.

8:57

So I want to start first by talking

8:59

about our model serving infrastructure

9:02

because I think sometimes people still

9:03

think about application layer companies

9:05

as you're calling a single model

9:07

endpoint and it's maybe a chat product.

9:10

And so a big problem we've solved over

9:12

the past four years is

9:15

we operate in 60 countries. We have

9:17

multiple product surface areas.

9:19

Customers have different model

9:20

preferences.

9:21

And even with the closed just the closed

9:23

source models, how do you serve these at

9:26

scale?

9:27

And so

9:28

this matrix gives you a sense of all of

9:31

the things we need to handle when we're

9:32

thinking about serving these models at

9:34

scale.

9:35

And as a simple example for each of

9:37

these model families, we need to serve

9:39

multiple of these models. We need to

9:41

have fallbacks across providers to hit

9:43

our SLAs.

9:45

And now with companies like

9:47

Fireworks.ai, we can add open-source

9:49

models into this mix.

9:54

In order to think about when you serve

9:56

models in production and how you modern

9:59

them modern them in production, you need

10:01

to have this infrastructure in place

10:04

thinking of even before you think about

10:05

post-training. And so this is how we

10:07

think about when there's a new model

10:09

released, whether it's open-source,

10:11

closed-source, or a model we've

10:13

post-trained, how we make decisions

10:15

whether to put it in production. And

10:17

then once it's in production, how we

10:18

make decisions whether to keep it in

10:20

production.

10:21

>> [snorts]

10:21

>> And so pre-production, we have a set of

10:24

generic evals. So in terms of automated,

10:28

we have the lab benchmark that I talked

10:30

about where whenever a new model comes

10:32

out, this gives us a very quick sense of

10:34

is it a frontier model, how strong is

10:36

it, what areas of legal is it good.

10:39

We have human testing in the generic

10:41

case where we run side-by-sides of this

10:44

model with other models to compare them.

10:47

And then for every product surface, we

10:49

have critical user journeys and

10:51

automated product tests because a model

10:53

could be very good generically, but it

10:55

could not be a good fit for a specific

10:58

product surface.

10:59

And then we also have human product

11:01

testing. And together these signals

11:04

along with heuristics around cost,

11:06

latency, region availability is how we

11:08

decide whether we put a model into

11:11

production. And I think the important

11:13

point here is

11:14

this is the case for post-trained models

11:17

or non-post-trained models. And so you

11:19

can reuse this infrastructure and you

11:21

should have it in place before you think

11:22

about post-training.

11:24

Once a model's in production, same

11:26

thing, doesn't matter if the model's

11:27

post-trained or not. We do AB testing if

11:30

we're rolling out a large change, we

11:32

look at engagement to track if this

11:34

model's performing how we expect. Um

11:37

and then we look at things like uptime,

11:39

token efficiency, and then we call it

11:41

product feedback, I call it angry

11:43

customer emails. And so there's There's

11:45

all of these signals that tell you this

11:47

is working as expected.

11:50

And so you need that infrastructure in

11:52

place before you think about serving

11:54

models.

11:55

And then

11:56

the thing you need to do even still

11:58

before serving models is what I call the

12:01

simple open-source switches.

12:04

And so the first one is look at all the

12:07

places you're serving models and find

12:10

are there places in my product where I

12:12

can just naively swap open-source

12:14

models? And so for example, we have

12:16

parts of our product that generate

12:18

citations that don't need the largest

12:19

models and there's opportunities to swap

12:22

in GLM 5.2 and get cost or performance

12:25

benefits. And so that's the first thing

12:27

and doing this builds the muscle of

12:30

serving open-source models alongside

12:32

closed-source models.

12:34

And then the second which has gotten uh

12:37

very popular now is model routing. So

12:40

there could be places where you can't

12:41

naively swap an open-source model, but

12:44

what you can do is on certain queries

12:46

route to open-source models.

12:50

And then once you've done this, you have

12:52

everything in place to start building

12:54

the post-training flywheel.

12:56

And

12:58

once you start serving them in

12:59

production and collecting feedback, and

13:02

caveat here, you need to be very careful

13:04

about what collecting feedback means. In

13:06

our case, it does not mean training on

13:08

customer data,

13:10

but we do get feedback signals from our

13:12

user testing and other things that can

13:14

inform how we build future data sets and

13:17

improve these models.

13:20

And that is our playbook for building a

13:23

research lab.

13:24

We think in the future every AI comp-

13:27

every company is going to need to become

13:30

an AI company and figure out some

13:33

version of this playbook.

13:35

And I think despite that, most people

13:38

are still betting against this playbook

13:41

and application layer companies and the

13:43

frontier ecosystem.

13:46

But

13:47

>> But if we win

13:49

on our budget with this team

13:56

we'll change the game.

13:59

>> This is a scene from Moneyball, which

14:01

hopefully you've seen this movie,

14:03

where they're talking about we just won

14:06

20 games in a row.

14:08

And Billy Beane says, "It doesn't

14:10

matter. If we don't win the

14:12

championship, no one's going to

14:14

appreciate what we've done here."

14:16

And the quote is, "But if we win

14:20

on this budget with this team

14:23

we'll have changed the game."

14:25

And

14:26

I think now with the frontier ecosystem

14:29

all of you have the opportunity to do

14:31

the same.

14:33

Go change the game.

14:34

Thank you.

14:36

>> [applause]

14:42

>> Do you want to do Q&A?

14:44

>> You okay with that?

14:45

>> Yeah, of course.

14:46

>> [snorts]

14:48

>> Yeah.

14:49

>> I'm sorry, who's speaking? Um you talked

14:51

about the challenges with having uh

14:53

legal customers with very sensitive data

14:56

sets and talked a little bit about your

14:58

brother's lawyer and other in-house

14:59

experts doing their version of live

15:01

coding to produce the best work. What

15:03

does that look Can you give us a little

15:04

more detail about what that actually

15:06

looks like?

15:06

>> Yeah, great question. So

15:09

the biggest challenge with legal work is

15:12

when you're at a large law firm, the

15:13

type of work you're doing is, for

15:15

example, you're representing a company

15:17

doing an M&A.

15:19

And you get a data room, which is all

15:21

the contracts of the company you're

15:22

trying to acquire, plus there's all

15:24

these emails and meetings about the

15:26

negotiation.

15:27

None of that data is public, and so you

15:29

don't have this analogy of like open

15:31

source GitHub repos.

15:33

And the biggest challenge you run into

15:34

when training is you can maybe find some

15:37

of the final work product publicly, so

15:39

like a public purchase agreement, but

15:41

you don't have any of the input data.

15:43

And historically it was very awkward to

15:47

get a bunch of lawyers and say, "Hey,

15:48

make a fake data room." Because these

15:51

data rooms can be 10,000 contracts. They

15:53

all need to fit together.

15:55

And so what Julio figured out is

15:59

a very clever way to generate these data

16:01

sets. And so he started from the rubric,

16:04

and so he planted all these issues and

16:06

said, "Here's all the problems in the

16:07

data room and what I'm going to check

16:09

for

16:10

in a scenario, and then use that to

16:12

generate the data room, so you can plant

16:14

all of these issues like these contracts

16:16

don't tie together, this contract's

16:18

missing, and generate all of the data,

16:21

and then we'll use Mercor Shenorcal to

16:22

make all the contracts look realistic,

16:25

but now you have this input data set,

16:27

and you can have the model generate like

16:28

a diligence memo, and this way of

16:30

checking it to say, "Hey, did you catch

16:32

all of these issues that we know that

16:34

are in here because we planted them?"

16:36

And obviously, depending on the work

16:37

you're doing, you need to be clever of

16:39

how you create it, but that to me feels

16:42

like the really big unlock because now

16:44

we can start

16:46

doing this training that before you just

16:49

you had this chicken-and-egg problem

16:50

where we'd go to law firms and say,

16:51

"Hey, we can train you this model on

16:53

this client data." And they'd be like,

16:54

"Okay, prove it." And we were like,

16:55

"Show us the data." And they're like,

16:56

"No." And so now you can prove it here.

16:59

And now we have a lot of interest from

17:00

law firms saying, "Oh, this is really

17:02

interesting. Can we do it with our

17:03

data?"

17:05

Good question. Yeah.

17:06

>> Um

17:06

so, we're also

17:08

hiring for this.

17:09

>> We're also building an implied AI there.

17:11

Um from a hiring point of view,

17:14

it's difficult for you to compete with

17:16

researchers from Anthropic and and the

17:18

labs.

17:19

Do you try and play that game or do you

17:21

try and hire domain logic and lawyers

17:22

and experience? And where has that

17:24

worked and where has that not worked?

17:26

>> Yeah, I think this was one of the big

17:28

mistakes I made like when we first

17:29

started the company because I had worked

17:32

in these labs and so I knew a lot of the

17:34

folks and I would say, you know, come

17:35

work with us and then they were getting

17:37

these, you know, 100 million plus pay

17:39

packages and we weren't at the scale

17:41

where this made sense. And so I'd say

17:43

we're now just getting to the company

17:44

size where I would say we're not

17:47

competing with kind of the frontier

17:48

talent, but there is increasingly

17:51

folks that are doing PhDs, folks that

17:54

don't want to work at large labs and so

17:55

I think that is changing. And then I

17:57

think the second thing is

17:59

a lot of why you needed that talent

18:02

early on was it wasn't just doing the

18:05

post training, it was you need to build

18:08

the training infrastructure, the serving

18:09

infrastructure, and all of those things

18:11

combined, I think the talent to do that

18:14

was like very unique. Like my old

18:15

roommate was one of the folks who like

18:17

ran post training at OpenAI and he was

18:19

one of the best researchers I've ever

18:21

worked with, but now you can use things

18:23

like Fireworks, uh, tinker-like APIs,

18:26

and so you don't need to build the

18:27

training and serving infrastructure, and

18:29

so I think it opens up the pool of

18:31

talent and it feels very feasible to get

18:33

some of this talent now.

18:35

Yeah.

18:37

>> Uh, yeah, just another question on the

18:38

benchmark creation. So, when you're like

18:40

creating these rubrics, you still have

18:42

to create them in such a way that

18:43

they're separating them the frontier

18:45

models and you have to do that with the

18:46

experts themselves. Like how do you find

18:48

like navigating that challenge? So, the

18:50

experts kind of have to create this

18:51

rubric knowing or like trying to figure

18:53

out how the models are going to perform

18:54

on them.

18:54

>> So, the What do you mean separating the

18:57

>> Uh, I guess so, you have to create the

18:58

rubrics that can adequately separate

19:01

separate the front They they have to

19:02

challenge the the frontier, right?

19:03

>> Yep.

19:04

>> Um, I guess how do you navigate that

19:05

especially when you're like having these

19:06

in-house legal experts who might not

19:08

exactly know how to create that rubric

19:10

designed for a model.

19:12

>> Yeah, that's a good So, we think of it

19:14

more as how do we make them realistic

19:17

representations of the work our

19:18

customers are doing? And I think part of

19:21

why we picked the legal domain is if you

19:23

just build a realistic client matter

19:25

that these top firms are working on,

19:27

like the frontier model still can't do

19:30

this.

19:31

And

19:32

I think the thing over the past 4 years

19:34

we've built is So, for example, my

19:36

brother was one of our first hires. And

19:38

so, he's been working with the models

19:40

for 4 years. And so, his intuition of

19:43

how the models work, how to generate

19:44

these data sets is incredibly good. And

19:47

he's trained a bunch of lawyers to do

19:48

this. But I would say the focus is still

19:50

more of how do we build a really

19:52

realistic data room, and then rubrics

19:55

for the diligence. And then when we run

19:56

these models, we find, okay, there's

19:58

gaps and there's there's room for

19:59

improvement. But I think depends on the

20:01

domain.

20:03

Uh God.

20:04

>> Hey, uh Shadeen from Box.

20:06

>> Uh sorry, let me do that one and I'll

20:07

>> Great.

20:08

>> I mean, if you think

20:10

Great talk, by the way. Um Ross from

20:12

Trolysis. If you think of the uh your

20:15

you know, the the end-to-end motion is

20:17

like data generation, the environment

20:18

piece, then there is the algorithm piece

20:20

of like what how do we train, etc., the

20:23

research piece, and then there's the

20:24

infra piece.

20:25

If something is not working, like

20:28

anecdotally, what have you seen? Is it

20:31

like

20:31

operationally or dollar-wise or human

20:34

hours-wise, is it usually in the data

20:35

layer? Is it in the research layer? Or

20:37

is it in the infra layer? And how do you

20:39

go from there?

20:40

>> Who's orchestrating or debugging this

20:42

whole end-to-end pipeline in some ways?

20:44

>> Yeah, this is This is a great question.

20:46

I mean, this is a whole 'nother talk.

20:48

Um

20:50

I don't think there's a single

20:52

thing. So, I would say

20:55

what makes it easier for us is

20:59

we found product-market fit. We have a

21:01

product that's being used in production.

21:04

And we did this largely with the

21:05

closed-source models to start.

21:08

And so, that let us build a lot of this

21:11

kind of confidence in the thing is

21:13

working end-to-end. And so, we have a

21:16

product, people are using it, people

21:17

have been using it for many years.

21:19

The models work in this way, we can

21:21

monitor whether they're working in

21:22

production. So, that was kind of the

21:23

point of serving of you need a lot of

21:25

this in place.

21:27

And then, we have built the muscle of

21:30

new models come out, how do we put those

21:32

into the product, whether they're

21:34

closed-source or open-source models.

21:36

And so, that's you can kind of think

21:38

about the full cycle. And then,

21:39

obviously we've done a bunch of work of

21:41

like Harness engineering and context

21:43

management and all of these things. And

21:46

so, now you can think of post-training

21:47

as one small input into this broader

21:50

system. And so, we have our team

21:52

post-train a model, and then you feed it

21:54

into this broader system of, you know,

21:56

you treat it just as another new model.

21:58

And so, then you have all these signals

22:00

of, you know, what's going wrong. And

22:03

so,

22:04

I would say, but across these steps,

22:07

there's kind of easy ways to gate it.

22:08

So, if you start with the benchmarks and

22:10

you post-train a model and it doesn't

22:11

work well on those, then it's unlikely

22:13

to go farther. If it works well on that,

22:16

and then you put in a product, but

22:18

someone uses the product and they say

22:19

things feel weird because, for example,

22:22

on the benchmarks we've built, there is

22:24

still things that they definitely don't

22:26

catch. Like, the best example is you can

22:29

have a model that does very well on our

22:31

benchmarks, but they've overfit to some

22:33

degree, and then you put it in a more

22:34

generic assistant-like product, and it

22:36

kind of falls apart for when you go out

22:38

of distribution. But, yeah, you need to

22:41

have kind of all of the like stages and

22:43

gates in place, and then it kind of

22:45

becomes obvious where things are

22:46

breaking. But, good question.

22:48

>> Yeah.

22:50

>> Oh, I had a question. Hey, great to meet

22:52

you. I'm from Rocks.

22:53

Uh on the on the benchmark side, legal

22:56

agent, I think you guys had Big Law

22:58

before, Big Law bench, and then now

23:00

legal agent of

23:01

What's your philosophy on open sourcing

23:03

the benchmark? Because if you open

23:04

source, of course you're getting

23:06

traction with other people trying out

23:07

the benchmark, but the labs can help

23:08

climb as well. And then if you close

23:11

source, then there's a question of is

23:13

this a real benchmark? So how do you how

23:15

do you how do you how do you what's your

23:16

philosophy?

23:17

>> I I think that's like a good tension

23:19

that we think about. I would say even

23:21

before we open source this benchmark, we

23:23

already work closely with the labs. And

23:25

so we share data with them to help them

23:27

improve their models, which improved our

23:29

product.

23:30

And so

23:31

I would say the way we think about it is

23:34

there is a huge advantage in your

23:37

industry, particularly in legal, in

23:40

terms of helping our customers

23:42

understand how good different models are

23:44

at different things. And so I think when

23:46

we when I started Harvey

23:49

I kind of had the intuition of oh, we'll

23:50

just build the best model and then

23:53

customers will be happy because it's the

23:55

best.

23:56

And it's very clear now that every

23:57

customer has different preferences,

23:59

different models are good at different

24:01

things. And then I think the second that

24:04

I didn't anticipate is kind of how big

24:06

the frontier ecosystem would become. And

24:08

so now we have so many companies

24:10

reaching out saying, "Hey, we have this

24:12

new technique or we did this thing. Can

24:13

we try it?" And before we just didn't

24:15

have the bandwidth where we're just like

24:17

we're working on these other things. But

24:19

now with this open source data set we

24:20

could be we can say, "Hey, go try this

24:22

and if it works, then great. This is

24:24

something that's interesting to invest

24:26

in." Um and then I think the way we

24:28

think about it strategically is

24:31

the valuable data for us is going to be

24:33

helping law firms train their own

24:35

systems and how they work on private

24:38

data. And then we want to help everyone

24:41

improve these systems with synthetic

24:43

data and kind of some of the data we

24:45

open source, but obviously a balance for

24:47

the reasons that you mentioned.

24:49

Yeah.

24:51

>> You mentioned

24:52

finding a lab.

24:53

What are the remaining open questions

24:54

for for that you want help on?

24:57

>> Yeah, that's a great question.

24:58

Um

25:00

I would say one big challenge, one of

25:02

the biggest challenges

25:03

so we can generate these very realistic

25:06

synthetic data sets, we can augment them

25:08

with humans, but the distribution of

25:10

that data still doesn't match our

25:12

production distribution. And so I would

25:14

say these data sets we're generating are

25:16

somewhat future facing. And so what I

25:18

mean by that is we can generate a really

25:21

realistic

25:22

data room,

25:23

but right now our our product isn't used

25:26

purely to do a due diligence. And so if

25:28

a model does well on that, but then

25:29

someone uses it to draft an email, maybe

25:31

it doesn't work as well. And so I think

25:33

there's still that gap because we can't

25:35

look at customer data, and so how do you

25:36

like bridge that gap? So I think that's

25:39

to me that that's the biggest question

25:40

on the data set side, and that one that

25:42

one is challenging.

25:44

There is

25:46

I think there is a bunch of questions

25:48

around

25:49

for example, again with the diligence

25:51

data, like the largest data room is 80

25:53

million tokens of context. Like the

25:55

models don't do a good job of managing

25:57

this context, and how do you train these

25:59

models to operate in kind of these very

26:01

complex environments? So I think that's

26:03

another one of just there's still a big

26:05

performance gap.

26:07

Um and I would say the third is

26:11

like we're making good progress on how

26:13

do we post train our

26:14

post train models ourselves, but I think

26:16

figuring out some form of continual

26:17

learning is I think the end game here is

26:20

not for us to build the best legal

26:21

model.

26:22

It's for us to help every law firm or

26:24

enterprise customer customize it to the

26:27

type of work they're doing.

26:29

And so just thinking about how do you

26:31

operationalize that, where you have a

26:33

large law firm and every time they work

26:36

on a client matter,

26:37

their AI system gets better, but you're

26:39

also protecting the client data. I think

26:41

that's kind of one huge like both

26:44

technical, operational, and like AI

26:46

challenge that that is super

26:47

interesting.

26:49

Yeah.

26:50

Uh

26:51

We're done.

26:52

Or more? I can do more. I can

26:56

Uh yeah, maybe last one.

26:57

>> Yeah. So, you know, you talked a lot

26:59

about like, you know, the the

27:01

the foundation readiness, right, to

27:02

compete with Front Row Live, right,

27:04

including your open source model,

27:05

including your infrastructure kind of

27:07

like Ruby.

27:08

From product perspective, right, Coda as

27:10

cloud product, they're general product,

27:12

right? So, from product perspective, how

27:14

do you guys compete with the, you know,

27:17

Coda as cloud workflow, methodology,

27:20

efficiency? How do you think about that?

27:23

>> Yeah. I think the big shift that we're

27:26

thinking about is

27:28

like our original product and then

27:30

things like Coda's Coda Work are

27:33

very individual-focused products, and so

27:36

they're about individual productivity.

27:38

And increasingly, the product we're

27:39

building is about organizational

27:41

productivity. And so, if you think of a

27:43

large law firm,

27:45

the problem they're trying to solve is

27:47

not how do I make my individual lawyers

27:49

more productive. The problem they're

27:51

trying to solve is I have 10,000

27:53

clients. I'm working on client projects

27:55

for all of them. I need to make sure

27:57

that all those client projects go very

28:00

well, and then I can also do them in a

28:02

way that's profitable. And so, a lot of

28:04

the infrastructure we're building for

28:05

the law firms is how do you operate that

28:08

machine? And when you start thinking

28:10

about

28:12

an individual client project,

28:14

most of the challenge is not how do I

28:16

draft this one section. It's I'm working

28:20

on this project for 6 months. I have a

28:21

team of 20, 30 people across the firm. I

28:24

need to coordinate all of them. I need

28:25

to make sure that I'm coordinating all

28:27

the outside parties. And so, a lot of it

28:29

is starting to look like project

28:31

management that is orchestrating these

28:33

humans and these agents. And that's at

28:36

the, I would say, at the team level. And

28:38

then, if you think about it at the

28:39

organization level, now you have a

28:41

thousand of these projects, and you need

28:43

to start thinking about resource

28:45

allocation, what am I billing, what am I

28:47

pitching? And then with enterprises,

28:49

it's even more complicated where kind of

28:51

a Fortune 500, they're working with a

28:53

thousand law firms. They have a thousand

28:55

people internally. What are all the

28:57

systems to kind of start work is trading

28:59

that work? And so I would say the simple

29:01

answer is just, and this is I think

29:02

historically what enterprise has has

29:04

done is how do you just go hyper

29:06

vertical into your domain in a way that

29:08

the horizontal products won't. Yeah,

29:10

good question.

29:13

>> Thank you, Gabe. That was an iconic

29:14

talk. Thank you for joining us.

29:17

>> [applause]

Interactive Summary

Gabe, co-founder and president of Harvey, discusses how application-layer companies can compete with frontier labs by building their own research, synthetic data sets, and post-training infrastructure on a budget. He details the Harvey playbook, emphasizing the creation of domain-specific benchmarks, using expert-guided synthetic data when customer data is too sensitive to train on, and leveraging open-source models alongside model routing in production to maintain performance and cost efficiency.

Suggested questions

4 ready-made prompts