HomeVideos

Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem

Now Playing

Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem

Transcript

1638 segments

0:00

One of our big guiding principles is

0:01

just simplicity. Um so like when you

0:03

look at a model like let's say a Chai 1,

0:05

I think there are 23 distinct submodules

0:08

in Chai 1. Um and like when you're

0:10

trying to iterate on something like

0:11

that, it gets really hard cuz you're

0:12

like I kind of I need to understand each

0:14

of these submodules independently. I

0:15

need to understand all of their

0:16

behaviors, their dynamics, and like that

0:18

doesn't actually scale that well. Um so

0:20

like you can start to think, okay, how

0:22

do I simplify this? How do I identify

0:23

like what's what's really important? And

0:25

once you have that like kind of the

0:26

whole research process and like

0:28

identifying these types of scaling

0:29

directions uh becomes a lot simpler.

0:33

>> [music]

0:48

>> We're here with Josh and Matt, two of

0:51

the co-founders of Chai Discovery. Chai

0:53

is engineering molecules with AI. So

0:57

it's something of a foundation model lab

0:58

for biology. Sounds like a big idea, but

1:01

let me just start there. What is the big

1:03

idea?

1:04

>> Thanks for having us on on the show. Um

1:07

one of the exciting things we're trying

1:08

to do at Chai is to make the drug

1:11

discovery process like a little bit more

1:12

like engineering. And we've seen all the

1:14

work happening with LLMs for code

1:16

generation for instance, right? And that

1:18

works really well because code's a very

1:19

simple abstraction, right? Of like to to

1:21

kind of get across what you want to do.

1:23

Uh biology doesn't look like that today.

1:24

It's a lot of like trial and error.

1:26

Um I think if we go back to uh the early

1:29

days of uh of just like modern biotech,

1:31

right? A lot of the medicines that we

1:33

have today were kind of discovered quite

1:34

randomly. It was a bit serendipitous.

1:36

And what we're trying to do is allow us

1:38

to industrialize that process a little

1:39

bit more

1:41

uh and try to come up with the tools

1:42

that we'd need in order to have

1:45

abstraction layers the same way we have

1:46

that in code and in modern engineering

1:47

allows us to iterate really quickly and

1:49

try to bring that uh into into biology

1:51

and into drug discovery.

1:52

>> So So let me ask you a question on that

1:54

then. So there's And tell me if this is

1:56

a reasonable way to frame it. There's

1:58

almost this boundary between that which

1:59

can be engineered

2:01

and that which needs to be tested in the

2:03

real world.

2:04

And it feels like that boundary has been

2:06

moving over time. Like the the the

2:08

proportion of the drug development

2:10

process that can be engineered as

2:12

opposed to trial and error'd seems to be

2:14

increasing. Is that a reasonable way to

2:16

think about it? If so, can you talk

2:18

about like what specific developments

2:20

have pushed that boundary over time?

2:22

>> I think that's one way to think about

2:23

it. Um the way that we often uh look at

2:26

this is like the lab is is an important

2:28

part to to verify that what you're doing

2:30

is correct. And actually verification is

2:31

a very big theme in AI as well. If you

2:33

can evaluate that your model works, you

2:35

can verify it, then you can start to

2:37

hill climb that and you can make

2:38

progress on it. So I think the lab is a

2:39

very important part of this. And then

2:41

the question is how do we take drug

2:42

discovery and make it look a lot more

2:44

like drug design. So one of the reasons

2:46

why we call this field drug discovery is

2:47

we're often looking for a needle in a

2:49

haystack. We'll screen millions,

2:50

billions of molecules try to find one

2:52

that works. And if we can instead put in

2:54

what is like the dream state of the

2:56

molecule you want and then the model can

2:58

materialize that, that's going to be

2:59

really powerful. So it's not even a

3:01

question of like reducing lab testing. I

3:03

mean, maybe that happens as a result. I

3:05

actually might even take the opposite

3:06

side of the coin and we can talk about

3:07

that where maybe we'd actually do even

3:09

more lab testing cuz the ROI will

3:10

increase. The same way there's more

3:12

software engineers

3:13

or there's more demand for software

3:14

engineers now that they become more

3:15

productive, there might be more demand

3:17

for the lab.

3:18

But I think the key part is how do we

3:20

just change the paradigm here and how do

3:22

we we make it more design oriented?

3:24

>> Mhm.

3:25

Can we talk about what is state of the

3:27

art today?

3:29

And then maybe let's take a little trip

3:30

down memory lane. 5 years ago, 3 years

3:32

ago, 1 year ago. Like what have been

3:35

some of the major breakthroughs and like

3:37

how has state of the art changed over

3:39

the recent years?

3:40

>> Yeah, I think the the field has like

3:42

it's really evolved especially like over

3:44

the last decade. Um so it really wasn't

3:46

until like uh you can actually look

3:48

back. There's like this bi-annual

3:50

protein folding competition. So every 2

3:52

years uh

3:53

usually a bunch of academic groups would

3:55

compete in this protein folding

3:56

competition. What would happen is like

3:58

you kind of hold out these protein

3:59

targets that no one's ever seen before.

4:01

They never get deposited publicly. And

4:03

then all these groups compete so you can

4:04

predict the proteins the best. And like

4:07

really it wasn't until 2018 where you

4:10

started to see this like big step change

4:11

in in performance and then finally again

4:13

in like 2020 with AlphaFold 2. But

4:16

really like it was kind of like the

4:17

advent of deep learning that really like

4:19

kicked off this whole field. Um so at

4:21

first what you do

4:23

these are kind of like bring me back to

4:24

my original days in my advisor's lab. So

4:27

we worked on like one of the first

4:28

systems to do this with deep learning.

4:30

What we would try to do is predict like

4:31

the distances between amino acids in a

4:33

protein. And then we'd have some render

4:35

which then took in like these kind of

4:37

noisy incomplete looking distances and

4:39

then emitted some protein from that. Um

4:41

and what's pretty remarkable is that

4:43

with

4:44

relatively minimal information like you

4:46

could build a machine learning system

4:48

that could actually produce something

4:49

that really looks like a protein and for

4:51

the most part was was correct. Like

4:54

there are still pretty big gaps it

4:55

didn't quite have the resolution of like

4:57

what what we have today. Almost like

4:59

with the early image generation models

5:01

where it was like a little bit grainy a

5:02

little bit blurry and now it's like oh

5:04

my gosh like this is this is crazy high

5:06

definition images like in seconds. Uh so

5:08

we kind of like saw that same evolution

5:10

happen over time. So it started with

5:12

protein folding and then I guess once

5:14

that became more realizable

5:17

there were kind of these other sub

5:18

problems that people wanted to solve. So

5:20

now given a protein structure can I

5:21

design say like a sequence which might

5:23

fold to that? And that's getting closer

5:25

and closer to like the drug design

5:27

problem. So you're you're kind of like

5:29

now thinking okay

5:31

how do I start engineering proteins that

5:32

like match a certain shape or might

5:34

perform a specific function?

5:36

And so people kind of study that problem

5:38

independently and that was kind of like

5:40

2021 2022 time. And then kind of these

5:43

ideas began to merge together. It was

5:45

really like the advent of diffusion

5:46

models where we started to be able to

5:47

like, "Okay, I can now generate a

5:49

protein structure and a sequence kind of

5:51

simultaneously.

5:52

Um and I can start making these like

5:54

what we would call prompts more and more

5:56

realistic." So, I can now prompt these

5:57

models on a target that might have a

5:59

certain shape and I can say, "Hey, I

6:00

want to bind over here." Kind of like

6:02

add add more real-world constraints on

6:04

the problem.

6:04

>> What did you guys see in 2024 that made

6:06

you think that was the right time to

6:08

start the company?

6:09

>> Yeah, there were a lot of discussions

6:10

that went into it. I remember one of

6:12

these uh early discussions actually back

6:14

in Matt's house when we were uh you

6:16

know, Matt was was showing me some of

6:18

these results on like antibody antigen

6:20

like structure prediction. And we Matt

6:21

was just talking about protein folding.

6:23

Um but for for a long time people

6:24

thought that protein folding for

6:26

antibodies was just like too hard of a

6:27

problem. People were like there was not

6:29

enough data for for any in antibodies

6:31

like in the protein data bank for

6:32

instance to like solve this problem. And

6:34

people as a result thought that like

6:35

antibody design was going to be out of

6:37

reach. A lot of the work that people

6:39

even did on protein design in the early

6:40

days with deep learning, it wasn't

6:42

antibodies. It was uh these different

6:43

class of proteins called mini proteins,

6:46

um which are uh actually really

6:47

interesting in their own right, but

6:48

they're not what most of the uh drug

6:50

industry is is looking at. So, the holy

6:52

grail was whether we could design

6:54

antibody proteins on the computer,

6:55

especially ones that had all that

6:56

therapeutic function. And our thinking

6:58

was that if you couldn't predict what an

7:00

antibody looks like, how are you ever

7:02

going to design one?

7:03

>> Traditionally, people have thought like

7:05

like biology is scary. Um like there's

7:07

so much to know.

7:08

>> of biology.

7:09

>> I I am as well. Honestly, [laughter]

7:11

like

7:12

uh like my my background was was never

7:14

biology. I I studied like pure math and

7:16

started my PhD in uh theoretical

7:18

computer science. And it was only like

7:20

after my third year that I ended up

7:21

switching into like deep learning,

7:23

protein structure prediction, all of

7:24

this stuff. So, it was like totally new

7:26

field to me. Seemed insane, but like at

7:28

the end of the day, it's much simpler

7:29

and like the the problems are much more

7:31

like interconnected uh than one might

7:33

think. So, like

7:34

I think people like when they start

7:36

getting into the field, they're like,

7:37

"Oh man, what's an antibody? What's a

7:38

mini protein? What are These are all

7:40

just like sequences of amino acids." At

7:42

the end of the day, like these are just

7:43

like different types of prompts for the

7:45

model. Um but like in the same way where

7:47

you might like have a math problem that

7:48

goes into chat GBT, chat GBT can both

7:51

answer your math math problem and like

7:53

help you with your English homework. Uh

7:55

so like really we we have the same type

7:56

of thing going on with our models. Like

7:58

we just have some way of representing

8:00

these sequences of amino acids, then we

8:01

have a way of like designing, predicting

8:03

those as well. Um and I think like in

8:05

that lens, like things things become a

8:07

lot more clear.

8:09

>> And Matt, you mentioned your background

8:10

a little bit. Um Josh, talk a bit about

8:12

your background and then more broadly,

8:14

in order to pull this off, you have a

8:15

bunch of different disciplines that kind

8:17

of come together. So can you just talk a

8:19

bit about like your background and that

8:21

of some of the other core members of the

8:23

team and how these things all fit

8:24

together?

8:25

>> Yeah, I've been excited about AI and

8:27

biology uh since I was a kid. So I guess

8:29

I'm like Matt Matt, I didn't start with

8:31

like theoretical CS and get into that

8:33

way, but um I went to a high school with

8:35

a stem cell lab. So I was just always

8:36

excited about biology as a kid and I

8:38

grew up as a programmer. Um I uh I

8:41

really started my career at Open AI, so

8:43

it was on on the early team there. It

8:44

was a non-profit back then. So it was a

8:46

a pretty uh good time to be there. We

8:48

did GPT-1, GPT-2, scaling laws. Um and

8:53

it was uh the question was like if the

8:54

models can learn to speak English,

8:56

German, French, why can't they learn to

8:57

speak DNA and protein? Um and that was

9:00

kind of my research agenda since then. I

9:02

think that sort of intersects with

9:03

around the time like Matt got into the

9:04

field as well. And it was a uh I think

9:07

it was a pretty important time as well,

9:09

right? Because if you look at the kind

9:10

of methods that we were bringing in,

9:11

like there's been a lot of these, you

9:13

know, changes on the edges if you will,

9:14

right? And you know, Matt talked a bit

9:16

about the history of what's happened in

9:18

the fields here. Um but you know, it's

9:20

all about like how do we find like the

9:21

right deep learning architectures with

9:23

the right compute configuration and the

9:25

right model architectures this happen.

9:27

What are the right tasks to apply it to?

9:29

We're talking about how we even knew

9:30

that like, you know, 2024 is the right

9:32

time to start the company. As Matt was

9:34

saying, you know, these are all

9:34

different like kinds of amino acid

9:36

sequences and people thought that, you

9:38

know, antibody class of of problem was

9:40

going to be too hard and we started to

9:42

see the first signs of life that

9:43

actually this was starting to work. I

9:44

think a lot of it fueled by some of the

9:47

new architectures we're bringing to the

9:49

problems. I think it was diffusion

9:50

models back then.

9:51

>> The first time like anyone was able to

9:53

generate reasonable looking proteins

9:56

was the advent like of diffusion models.

9:59

So it was pretty crazy like there were a

10:00

bunch of generative modeling approaches

10:02

that like would kind of work if you if

10:03

you had a bunch of data. So like people

10:05

got these working for images. There were

10:07

like a bunch of tricks to make this

10:08

better and better along the way. Like

10:10

we've definitely borrowed a lot of those

10:11

ideas in our domain as well.

10:13

But it was really like once diffusion

10:15

models came around

10:16

>> And is there is there an intuition for

10:17

why diffusion models work?

10:19

>> Yeah, so diffusion is is not a one-step

10:22

process. So like I think kind of up

10:24

until this point like the the main

10:26

generative design paradigm was like

10:29

called variational autoencoders. And in

10:31

that case you're like saying I want to

10:33

just like compress my input

10:34

distribution. So like you you have some

10:36

like proteins you want to make these

10:38

look like kind of like fuzzy Gaussian

10:40

vectors. That task is just like really

10:42

hard. And maybe today if we tried like

10:45

super hard I think we'd crack it. But at

10:47

the time like I I think like it wasn't

10:49

really as developed enough to to work

10:52

for our problems. What turned out

10:53

working really well was just kind of

10:54

giving the model more time to think and

10:57

showing it like more examples like

10:58

here's like a slightly broken looking

11:00

protein. How do you make it better? And

11:02

you can kind of break that protein more

11:03

and more and more and you can make it

11:05

look more and more noisy, more and more

11:06

broken and teach the model just quick

11:08

little shortcuts. All right, here's how

11:10

I make it slightly better. You can just

11:11

keep asking over and over again. Make it

11:13

slightly better. Make it slightly

11:14

better. And kind of like breaking the

11:16

problem down to that scale worked really

11:18

really well for for biology.

11:20

>> Yeah, the make it slightly better

11:21

reminds me of a game that I like to play

11:22

with my daughters where we have chat GPT

11:25

give us a unicorn and then we make it

11:27

stronger. And we just keep telling it to

11:29

make it stronger. And by the time we're

11:30

done we have the strongest unicorn in

11:32

the world. So, about the same, right?

11:35

>> Yeah. That's uh

11:36

>> [laughter]

11:37

>> It's that easy.

11:38

>> Yeah.

11:39

>> Okay, maybe not the same.

11:40

>> Yeah. So, we feel like in this domain,

11:43

you need to assemble a kind of like a

11:44

quadrilingual group of people, like an

11:47

Avengers squad of chemistry people,

11:50

biologists, AI folks. And so, that's a

11:52

challenge. Um how have you guys gone

11:54

about finding people, convincing people

11:56

to join the team, and and who are your

11:57

superstars?

11:58

>> So, we've been really pragmatic about

12:00

this at Character. If you look at the

12:01

founding team, it was mostly AI

12:02

researchers. Um so, people who had uh

12:05

worked on either scaling models or

12:07

getting them to work in this domain.

12:09

But, really with each generation of

12:10

model, uh the kind of people we've

12:12

needed for the next milestone has has

12:13

changed. Uh or or I'd say probably has

12:16

expanded, right? Uh so, if you look at

12:18

Character 2, right? That's the point

12:19

when we started to bring in um some of

12:21

the most incredible like, you know,

12:23

antibody engineers and and scientists in

12:25

the world. Um

12:26

uh one of the scientists from Andy Andy

12:28

Young Actually, when we hired him, uh

12:30

people asked us if we had pivoted into

12:32

building a full-stack drug pipeline cuz

12:33

they're like, "You'd be crazy not

12:34

[laughter] to do that if like Andy

12:35

joined your team." Um but, uh I think

12:38

Andy has has enjoyed running more

12:40

antibody campaigns in the past couple of

12:41

months than he's probably run in his

12:42

whole whole career, uh which is very

12:44

cool to see. You have folks like Nathan

12:46

Rollins on the team. Like, Nathan was

12:47

actually homeschooled and then started

12:49

college very early on. So, he joined

12:51

like David Baker's lab who won the Nobel

12:52

Prize for protein design like when he

12:53

was 14, uh started his PhD when he was

12:56

18, and has so many creative ideas. On

12:58

our uh as the model started to get

13:00

better, we needed to build up a product

13:01

team uh because while the researchers

13:04

might get the models to point that

13:05

they're uh they're very powerful, you

13:07

need to build the right product

13:08

interfaces so that the models are

13:09

actually useful. Um and uh that's where

13:12

we started to bring on uh people who

13:14

have, you know, built some of the most

13:16

exciting products we know about today.

13:17

Uh like our co-founder Jack worked at

13:19

Stripe, Monaz who was one of the top 10

13:20

code contributors at Stripe, uh Neil who

13:23

ran his own cybersecurity company before

13:25

security started to become very

13:26

important as we deploy this to our big

13:27

partners. As as we started to scale up,

13:30

we brought in people uh who've

13:32

uh really done a lot of like the GPU

13:34

hacking, if you will, in order to like

13:35

scale up our systems that they they they

13:37

don't break when we're running them at

13:38

scale. We had an email from uh or a

13:40

Slack message from one of our

13:42

hyperscalers the other day uh where you

13:44

know we had a cluster, I think that had

13:45

some issue to it and they're like, "Oh,

13:46

like the GPUs got too hot. I think you

13:47

guys are running too many." And we're

13:48

like, "Isn't that the point?" Right?

13:50

That's probably I was like, "Good, we're

13:51

doing our job at least." Yeah, so like

13:53

how do you convince those people to

13:54

join? I think again a lot of it comes

13:55

back to uh the results and like a clear

13:57

need. We didn't hire antibody engineers

13:59

before in an antibody design model. Like

14:01

what are those folks going to do? We

14:02

were even I think worried when we

14:03

started that trend because for some of

14:05

the um next generation formats, the

14:07

complicated antibodies, like

14:09

multispecifics, like they didn't even

14:10

work with Chai too. Uh so it actually

14:12

took a couple of of uh weeks when some

14:15

of those people showed up before the

14:16

models could work at a point that they

14:17

could work on some of these interesting

14:18

case studies. Uh but fortunately

14:20

progress is fast enough to kind of bring

14:21

that online. So I think we're always

14:23

like evolving that team and going

14:24

forward you know, that next milestone.

14:26

We've got the team very small as a

14:27

result, too. Um so this way, you know,

14:29

everyone is a little bit like slightly

14:31

over capacity, I think, which means we

14:33

have to prioritize. It forces us to work

14:35

on the things that really matter most.

14:37

>> I think on the research side as well,

14:38

one of the founding engineers, Kevin Wu,

14:40

he had the first uh

14:42

I think it was the first protein

14:44

diffusion model like ever. Uh and that

14:46

speaks to Kevin's speed of execution.

14:48

Like he is a heck of an engineer and

14:50

like I think engineering has just always

14:52

been like important since day one. Um so

14:55

like really even our researchers,

14:56

they're all excellent engineers and like

14:58

we really care about building a

15:00

high-quality codebase. Like at the end

15:01

of the day, we are

15:03

technically like a software company.

15:05

We're we're AI researchers, we're we're

15:07

protein designers, we're we're a lot of

15:09

things, but uh our deliverable is some

15:12

piece of software. So we've kept that

15:13

like kept that bar really high while

15:15

also trying to like, you know, level

15:17

that with great research talent and

15:19

people that can actually like push the

15:20

frontiers of what's possible.

15:22

>> And you've had a number of it seems like

15:23

aha moments in the field. Like AlphaFold

15:26

was an aha, we can figure out how a

15:27

protein folds. And then the diffusion

15:28

models, aha, like we can generate

15:30

proteins. Um and it seems like your

15:32

latest models are a kind of another aha

15:35

moment. We're not only generating

15:37

molecules that look like proteins, but

15:39

they also have therapeutic properties.

15:41

Like they can bind really tightly.

15:44

Um they have really high hit rates. Can

15:45

you talk sort of about the quality of

15:47

the molecules that your models are

15:48

producing now and everything that sort

15:50

of went in to those models to make them

15:52

able to do that?

15:53

>> Yeah, if we look at the the quality of

15:55

the molecules that's coming out, it goes

15:57

back to, you know, one of the thesis

15:58

when we started the company that uh we

16:00

really wanted to focus in on this like

16:02

de novo generation of the molecules. If

16:05

you look at what a lot of the drug

16:06

discovery AI work was at the time, it

16:08

was about how do I take a molecule and

16:09

just make it a little bit better, right?

16:11

Which we just talked about in a sense,

16:12

but it was doing it with a lab in the

16:14

loop uh style, right? Where I take some

16:16

data, try to make it better that way.

16:18

And the the question uh

16:20

we started with is, well, can we

16:22

actually just like do all that on the

16:23

computer, right? Is there a way that the

16:25

there maybe there would be enough data

16:26

or we could collect enough data um so

16:28

that we could just zero-shot a molecule

16:30

that has a lot of these properties? So,

16:32

the first thing we needed to do to get

16:34

there um was to design molecules with

16:36

really high success rates. When we

16:37

started the company, the state of the

16:39

art for anybody design was about like a

16:40

0.1% binding rate. So, one in a thousand

16:43

of the molecules you design would

16:45

actually bind in the lab. Um so, first

16:47

of all, that means you have to screen a

16:48

lot of molecules to find some good ones.

16:49

It also means that like the gradient you

16:51

get on your process is actually quite

16:53

weak as well. So, for many targets you

16:55

won't get any hits. For the ones that,

16:57

you know, you do get some hits, you

16:58

won't have enough to actually see

16:59

whether you're having like the drug-like

17:01

properties. So, we really focused in on

17:03

how do we just make this process more

17:04

accurate? We got to uh with our chi-2

17:06

model about like a 15% success rate. Uh

17:09

so, now if you screen a a thousand

17:10

molecules, you're getting 150 back. Now,

17:12

you can get start to get some like

17:13

interesting statistics on the properties

17:14

of the molecules, right? And a lot of

17:16

allowed us to like iterate on that,

17:17

build the right evaluations around that

17:19

in the lab, and really try to hill climb

17:21

that as well. And we're getting to a

17:22

point now where we can actually bake in

17:24

a lot of these different properties from

17:26

the start into this engine. And then

17:28

maybe if you think about, you know, how

17:29

does this happen or like how are we

17:31

approaching this as well? And like why

17:32

do we think it's going to continue

17:33

improving?

17:34

>> Yeah, so like um honestly, I think Josh

17:37

and I were both surprised at how quickly

17:39

this worked. Um like when we were

17:41

originally budging this, we're like,

17:42

"Ah, maybe like 20% hit rate like 3 or 4

17:45

years." We're like really like a 1%

17:47

rate. We thought 1 in 100 would be

17:48

amazing. We're like, "This is going to

17:49

be you know, a groundbreaking thing." We

17:51

had a philosophy and like an approach

17:53

that we wanted to take and it just like

17:55

ended up working really well. And that

17:57

approach is like very similar to what's

17:58

worked in the rest of machine learning.

18:00

So, people kind of treat biology as this

18:01

like bespoke problem or like bespoke

18:03

field, but really it's like the same

18:05

principles as like self-driving LLMs.

18:07

>> Well, same principles, but one of the

18:09

things we've talked about before is how

18:10

you managed to find scaling laws.

18:14

And it's one thing to tokenize a string

18:15

of text. It's another thing to tokenize

18:18

biology.

18:19

Can you say a couple words about

18:22

without giving away any of the magic,

18:24

you know, can you just say a couple

18:25

words about that challenge and how you

18:27

guys solve that?

18:28

>> Yeah, so one thing that I really like

18:29

about Chai is like we're we're a very

18:31

bitter lesson pill company. So, like we

18:32

really believe in like scaling data,

18:34

scaling models, scaling compute. In

18:36

order to do that, obviously you you need

18:37

to identify scaling laws. Otherwise,

18:39

you're just kind of like wasting time

18:40

and resources. Um and I think like

18:43

without giving away too much, like one

18:44

of our big guiding principles is just

18:45

simplicity. Um so, like when you look at

18:47

a model like let's say a Chai 1, um so

18:50

like there are like I think there are 23

18:53

distinct submodules in Chai 1. Um and

18:55

like when you're trying to iterate on

18:56

something like that, it gets really

18:57

hard. Uh cuz you're like, I kind of need

18:59

to understand each of these submodules

19:01

independently. I need to understand all

19:02

of their behaviors, their dynamics, and

19:04

like that doesn't actually scale that

19:06

well. Um so, like you can start to

19:07

think, "Okay, how do I simplify this?

19:09

How do I identify I what's what's really

19:10

important?" Um and once you have that

19:13

like kind of the whole research process

19:14

and like identifying these types of

19:16

scaling directions uh becomes a lot

19:18

simpler.

19:18

>> One of the other things too that I think

19:19

is interesting about this is we take a

19:21

lot of these lessons from what's worked

19:22

in the rest of the deep learning space,

19:24

but as you're pointing out like the data

19:25

itself is different. Yeah. The models at

19:28

Chai are completely built from scratch.

19:29

We're not like fine-tuning GLAM or

19:31

something like that on some protein

19:32

data. We build everything from the

19:33

ground up. I think a lot of the company

19:36

building process though is is taking a

19:37

philosophy and actually sticking with it

19:39

and iterating on that and just having

19:41

some guiding principles. When you're

19:43

building like a blue sky research

19:44

company if you will, you know, Chai is

19:46

is almost like one of these neo labs,

19:47

right? Like we have this big AI problem

19:48

we're going after where as we make

19:50

progress on it, you know, that opens up

19:52

opportunity for for our customers. Um,

19:54

but if you're going to work on something

19:55

so open-ended that way, you need some

19:57

principles to guide you. And I think

19:59

we've done a very good job on like

20:00

tracking those principles in the

20:01

company, working on it. So a lot of the

20:03

things that Matt

20:03

>> Yeah, can you share them? What what are

20:04

the guiding principles?

20:05

>> So I think simplicity was one of the

20:07

ones that Matt mentioned. It's this like

20:08

bitter lesson pillness of like, you

20:10

know, scaling compute and data and and

20:12

models. Uh, it's being really rigorous.

20:14

Uh, that's something that's so important

20:15

in this space. You can fool yourself so

20:17

easily in biology. Like the error bars

20:19

in the wet lab are actually quite large

20:21

as well. Um, so it's actually a little

20:23

bit different than uh if you look at

20:25

like code generation for instance. If

20:27

you look at SWE bench, people are like,

20:28

oh, I like this is a maybe like a year

20:30

ago people like, oh, I got 16% and 17%

20:32

and 18%. I mean, in biology if if you're

20:35

like plus or minus 5% in your lab, like

20:37

that's all that might all be the same.

20:39

So it actually just means uh the bar is

20:41

really high in terms of the step changes

20:42

that you want to see with the models,

20:44

but you also need to be really honest

20:45

with yourself about whether you're

20:46

making progress or not. Um, so you could

20:48

come up with some fancy model that looks

20:50

like it works well on like one or two

20:51

new tasks, but it's very important to

20:54

show that that works more generally. Uh,

20:55

if you're actually trying to build a

20:56

product that can like bring the field

20:58

forward.

20:59

>> Yeah, and that's I think pretty

21:00

interesting because biology is one of

21:01

those inherently not so verifiable

21:03

domains. And you guys have been really

21:06

good at um, sort of showing your

21:08

progress to customers,

21:10

um, and to people like us who know very

21:12

little about biology. And so, can we

21:13

talk a little bit about the evals and

21:16

the verifiable part of of the model

21:18

progress? Like, how do you guys know

21:20

that your models are getting better?

21:21

>> Well, I would say first of all that, uh,

21:23

I actually think this is one of the more

21:24

verifiable domains. It's actually very

21:26

objective readout. If you look at

21:28

something even like CodeGen, right?

21:30

Like, maybe it's a verifiable task like,

21:32

did my code compile? Did it solve these

21:34

unit tests? But, how do you think about

21:35

the taste, right? Like, did I write some

21:37

really sloppy code that can't be

21:38

maintained? Like, what does that look

21:39

like? When we think about designing a

21:42

molecule in the lab, uh, we can actually

21:44

be quite specific about many of these

21:46

properties, right? So, maybe we get a

21:47

molecule that binds the target, but can

21:48

I manufacture it, right? That might be

21:50

your, you know, version of like some

21:51

tech debt, but you can measure that. Uh,

21:53

and I think those evals, again, actually

21:55

make this domain more verifiable. Maybe

21:57

it takes a little bit longer to validate

21:58

it, right? It's not like five seconds to

22:00

like get a readout and run a unit test.

22:02

You might have to spend a couple days, a

22:03

couple weeks in the lab to get that

22:04

readout, but at least you can be honest

22:06

with yourself.

22:07

>> Yeah. When we talk about progress and

22:09

how good the models are, there's like a

22:10

domain of targets in biology,

22:13

um, that you can, as you mentioned, sort

22:15

of address with traditional screening

22:16

methods. They take a long time. They're

22:18

very slow and rudimentary. Um, and then

22:20

there are targets that just aren't

22:22

addressable with the with existing

22:23

methods. They're not druggable for

22:25

whatever reason.

22:26

Um, and so, where are we in terms of

22:28

model progress in terms of, you know,

22:30

working on existing targets and

22:32

generating molecules faster? That's one

22:34

end of the spectrum. And on the other

22:35

end of the spectrum is unlocking novel

22:37

biology, new targets, and things that we

22:38

couldn't drug before.

22:40

>> This kind of goes back to like, uh, kind

22:42

of like our core modeling philosophy.

22:44

Um, so like, there there have actually

22:45

been like plenty of times where like,

22:47

man, if we had a 24th module, like we

22:49

can actually like unlock that new

22:51

target. Um, and we're like, is that

22:53

really something that we want to

22:54

maintain long term? Is this incremental

22:55

or is this like actually a compounding

22:57

improvement? Will this actually like

22:58

help us generalize to the the class of

23:00

like these this whole class of targets

23:02

that we we really can't hit. And so like

23:04

our our philosophy has been like, all

23:05

right, let's just like continue to focus

23:07

like identify your scaling laws. Like

23:09

there there comes a point where like if

23:11

the model is able to push loss down even

23:13

further, it has to understand like

23:15

something like very intrinsic about the

23:18

target that it's operating on. So like

23:20

one example we were talking about the

23:21

other day, Paul and I, maybe the way

23:23

that we're looking at like certain

23:24

glycosylation sites on proteins we're

23:26

like, "Oh, we might want to represent

23:28

them differently or something like

23:29

this." And we're like, "Well, even if we

23:30

didn't represent them, like there are

23:32

certain sequence motifs that will tell

23:33

the model like there should be a

23:35

glycosylation site here. And in order to

23:36

drive like loss down further, the model

23:38

should just like have to learn that." Uh

23:40

so there are like all these hidden

23:42

features of targets where like if you

23:44

really believe in scaling laws you can

23:45

believe the models will get there. Like

23:47

these types of targets like should just

23:49

unlock with better models. And of course

23:51

like you still have to take this very

23:52

seriously and like you still need all

23:53

the proper validation. Um you need to

23:55

like really challenge yourself and like

23:57

make sure that this is truly working.

23:59

But I think our approach has always been

24:01

like, you know, we with better models

24:03

like we we should be able to unlock a

24:05

lot of these targets.

24:06

>> One of the other interesting things is

24:07

if we look at we talked about like the

24:08

interdisciplinary nature of this. If we

24:10

look at the different teams at Chai,

24:12

what people will call hard target is

24:14

actually different in literally every

24:15

team. So on the science team it might

24:17

be, you know, like a undruggable GPCR

24:20

target no one's gotten something that

24:21

has like, you know, modulated that in a

24:23

functional way. On the ML research team

24:25

it'll be something like, "Oh, there's

24:26

like the the model just can't fold this

24:28

thing up. It doesn't know what it looks

24:29

like."

24:30

And then on the product team it might

24:31

look like, "Oh, I've got all these like,

24:32

you know, modifications like my

24:33

glycosylations and it's a membrane

24:35

protein. How do I represent that to the

24:36

user?"

24:38

And actually the fact that it's

24:39

different for each of these groups I

24:40

think is a feature rather than a bug.

24:42

And it means that if we want to make

24:44

like broad progress over here, everyone

24:46

is kind of pushing in parallel on these

24:47

different ways and and that means that

24:49

there's very smooth progress that we can

24:50

make all the time. And there's usually

24:52

not one bottleneck at Chai. It's not

24:53

like, "Oh, if we only had that one extra

24:55

module things would work." Or if we only

24:58

you know, like try to push this into the

24:59

product in some way, we can unlock it.

25:01

We're trying to build this unified

25:03

solution because at the end of the day,

25:04

the goal of the company is to build a

25:05

computer-aided design suite for

25:06

molecules. It's not to make one or two

25:08

molecules. It's not to get a pipeline of

25:10

like five interesting therapies that we

25:12

bring to market. It's to change the way

25:13

that medicines are discovered. And we're

25:15

if we're going to do that, we need to

25:17

work on all the hard targets regardless

25:19

of how you define hard.

25:20

>> Yeah. Say more about this idea of

25:21

computer-aided design suite for

25:23

molecules. What does that mean?

25:24

>> So at its core, it goes back to this

25:26

point about making biology look more

25:28

like an engineering discipline. So we're

25:29

not going and fishing something out of a

25:31

large library or doing a ton of trial

25:33

and error. You want to be able to

25:35

specify up front the principles that go

25:37

into designing your molecule and then

25:39

have an engine that can actually realize

25:41

that into some molecule that we're going

25:42

to go and create in the lab. And look,

25:44

you still might do some iteration on the

25:46

lab and and on your model because maybe

25:48

your hypothesis was wrong. But we want

25:50

to speed up is is actually again have

25:52

that computer-aided design suite so that

25:54

you can go from idea to testable

25:56

hypothesis very quickly. And if that

25:58

loop right now takes something like 9

26:00

months, I don't know, to go and like

26:02

discover your molecule versus if it

26:03

takes 9 weeks or it takes 9 days, you

26:05

know, each order of magnitude just

26:07

scales in a very big way the number of

26:09

ideas you can really sort through. And I

26:11

think that's ultimately how the field is

26:13

going to converge on on better

26:14

medicines. It really comes back to like

26:16

people sometimes talk about do we care

26:18

and you kind of noted on it's not like

26:19

do we care about like speed or do we

26:21

care about like the difficulty of the

26:22

targets? At some point they converge in

26:25

this way as well because a hard target

26:27

if we can like iterate through

26:28

hypothesis a lot faster, then maybe

26:30

it'll be easier to crack it.

26:32

>> If so if we can maybe detach ourselves

26:35

from the present reality and go far

26:37

enough into the future that we're not

26:39

thinking present forward, we're actually

26:41

thinking future back. 2035, 2040, 2100,

26:45

whatever whatever you want. What does

26:47

the industry look like? Like let's

26:49

imagine that computer-aided design suite

26:51

for molecules has become a standard.

26:53

Let's imagine a lot of innovation has

26:55

flowed downstream into some of the wet

26:56

lab parts of the process. Like, can you

26:59

paint the picture of what the industry

27:00

might look like 10 years from now?

27:03

>> Yeah, I think it's going to be a a

27:04

really exciting time. And we can look at

27:06

this on on a couple of angles. So, first

27:08

of all, the quality of the medicines

27:10

that we develop will hopefully go up.

27:12

There's a lot of molecules we put into

27:13

the clinic today that there is really

27:16

hard to discover a molecule. You find

27:17

something that's like 80% of the way

27:19

there. Maybe advance it anyways to like

27:21

hit my timelines. It's probably going to

27:22

benefit some patients. But then you get

27:24

beat, you know, a year later by someone

27:26

else. And it's it's really just not the

27:27

most efficient spend of resources in the

27:29

industry. You've got the kind of

27:30

diseases that are just too hard to go

27:32

after today. People have been trying to

27:33

drug Alzheimer's forever. And

27:35

unfortunately, you know, haven't made as

27:37

much progress as as we'd like. You have

27:39

things that just aren't economical to go

27:40

after today. Think about like

27:41

personalized medicines, rare diseases,

27:43

things where maybe the patient

27:44

population is going to be smaller. But

27:46

again, if we can iterate through these

27:48

ideas faster, if we can launch something

27:50

faster, if we can do it in a less

27:52

expensive way,

27:53

then those probably come into reach as

27:55

well. So, I think there's just so many

27:57

different axes

27:58

that that we're able to to push on. And

28:01

I think that means that the the future

28:03

is is really bright.

28:04

>> It used to be like, you either want to

28:05

be first in class or best in class.

28:07

Now it's like, you want to be last in

28:08

class cuz like you actually just want to

28:10

like be the final answer.

28:12

So, I think like a lot of what you'll

28:14

see is just like way more intentionality

28:16

in the types of drugs that you're

28:17

designing. Like, this drug will be super

28:19

specific to the disease of interest. It

28:21

won't have like certain interactions,

28:22

like certain negative interactions that

28:24

a lot of drugs today do. A lot of this

28:26

stuff is actually like you're able to

28:28

model most of this computationally.

28:31

Like, maybe not today, but there's

28:33

definitely a path towards getting there.

28:34

And I think that's like one of the most

28:35

exciting things for me.

28:37

>> Very cool. Can we talk about a business

28:39

model decision that you guys made?

28:41

Cuz I think a lot of times folks think

28:43

about Chai and Isomorphic in in same

28:45

neighborhood. Isomorphic is developing

28:47

drugs. You guys are enabling the

28:49

existing industry to develop drugs more

28:52

efficiently, better, faster, cheaper

28:54

than they have before. Why did you

28:55

decide to go down that path versus,

28:58

you know, the Isomorphic path?

29:00

>> Yeah, first of all, I think both of

29:01

these paths are are great and

29:03

they can they can create tremendous

29:04

value. We've always been really excited

29:06

about building infrastructure for the

29:08

industry. Our bet when we started to go

29:10

back to those results in Matt's house,

29:12

you know, a few years ago,

29:14

was that this is how most future drugs

29:16

are going to be discovered. And if

29:18

that's the case, somebody needs to go

29:19

and like build that infrastructure and

29:20

make it happen. I think part of this too

29:22

is a lot of our founding team,

29:24

like including Matt and I, had worked in

29:26

companies before where we had built

29:27

these full-stack drug pipelines, right?

29:29

And we were building AI models. We're

29:31

putting those drugs in the clinic.

29:32

Again, we're we're pretty bitter lesson

29:34

pill as well and our our thinking was as

29:36

the models get better, we want to be

29:38

spending more of the money making better

29:39

models as opposed to diverting those

29:41

resources into clinical trials and

29:43

things like that. And one of the

29:44

interesting things about our business

29:45

model is as the models get better,

29:47

actually wins us the right to continue

29:48

investing more in them, right? And you

29:50

have, you know, partnerships with with

29:52

these pharma companies that are that are

29:53

paying off today and allows us to

29:54

continue invest in it. So, it's it's a

29:56

much more, you know, scalable business

29:57

in that way. And I just think about the

29:59

ultimate impact that we can create for

30:01

the world is is a lot larger. We've

30:03

We've always wanted to just partner

30:04

broadly with the ecosystem. It goes back

30:06

to this point about fooling yourself in

30:08

biology. If you work on a small number

30:09

of drug targets, you might come up with

30:11

the most exquisite molecules, creating a

30:13

ton of value for the world by doing

30:14

that.

30:15

But, you might miss the forest for the

30:17

trees because maybe your model doesn't

30:18

generalize to the other 500 molecules

30:20

that people are going to make that year.

30:22

And if you go and and partner with

30:23

people, you just you just can't fool

30:25

yourself. Like, you look at our

30:26

partners, Eli Lilly, Novartis, Organics,

30:29

Pfizer, like these are not companies

30:30

that are, you know,

30:32

they they take this stuff for granted.

30:34

Like, you have to really deliver on

30:36

these partnerships for them to take you

30:37

seriously. And that means it can't just

30:39

work like some some time. Like, when we

30:40

ship models at Chai, they really have to

30:43

work, they have to deliver value to our

30:44

partners, and it's not like, oh, we made

30:46

some molecule, doesn't fully work, we're

30:48

going to have our our chemists like

30:49

clean it up a little bit. So, it's it's

30:51

almost a harder business to pull off. I

30:52

think that's one of the reasons why you

30:54

haven't seen it pursued many times. If

30:56

you work on a drug pipeline, again, like

30:57

there might be ways to fix things up

30:59

later. There will be the proof of like

31:00

what happens in the clinic, of course.

31:01

Um but uh when you have this partnering

31:03

based model, you have to be really

31:05

rigorous about your models. Your models

31:06

have to work really well because

31:08

otherwise those partners not going to

31:09

come easily. So, it's made our life

31:10

harder, I think in many ways, but I

31:12

think it's also the more rewarding path

31:14

if we can get it to work.

31:15

>> what have you learned? I mean, you know,

31:16

you're you're out of the lab, so to

31:17

speak, and in the real world, you know,

31:19

delivering real value for actual pharma

31:21

companies. What have you learned as you

31:23

start to work with these partners in

31:24

terms of any surprises in terms of how

31:26

their needs might have differed from

31:28

what you expected or how their level of

31:30

sophistication around this might be

31:31

different than what you expected. Have

31:32

you any surprises or any learnings from

31:34

working with these partners?

31:35

>> Yeah, so we went into these

31:36

partnerships. I think a lot of people

31:38

told us that like pharma doesn't know

31:40

how to use AI, these are not it's not

31:41

like a tech forward industry and things

31:43

like that. And to be honest, that hasn't

31:45

really been our experience. I think that

31:47

these are again, they're very rigorous

31:49

partners, very rigorous customers,

31:50

right? And they're they're going to to

31:52

to test every one of our claims, right,

31:54

before they start to deploy these

31:55

things. But when they see the data, they

31:57

go all in, right? Because pharma is an

31:59

innovation industry. Um and it's

32:01

interesting just to think about even the

32:02

whole economics of the pharma industry.

32:04

If you build a product in pharma, right?

32:06

Like like a drug, right? You only have

32:08

exclusivity on that drug uh for a

32:10

certain period of time, right? And you

32:11

have to continue to reinvest. Eli Lilly

32:13

is a trillion-dollar pharma company

32:15

right now. If they don't get more

32:17

blockbuster drugs, they will not be a

32:18

trillion-dollar pharma company forever.

32:20

Um and I think that forces these

32:22

companies to really be on their game of

32:25

adopting new technologies and deploying

32:26

them and trying to stay ahead. Pharma is

32:28

a very competitive arena. There's a lot

32:30

of people trying to bring these drugs to

32:31

patients with By the way, it's great

32:32

great for patients, but uh it means that

32:34

you have to be on top of your game here.

32:37

Uh and I think that means that um you

32:38

know, once you're through that door and

32:40

your models are working, we've seen like

32:42

an an upscale this adoption very quickly

32:44

and people thinking about how to use the

32:45

models in incredibly creative ways.

32:47

>> I really like the incentive structure

32:49

that it creates as well. Um like taking

32:51

more of the partnership model. As our

32:53

models get better, we get better results

32:54

to our customers and so on. Um so I

32:57

think like that's just like a nice like

32:59

nice side effect for us as well. Chai

33:01

has been like incredibly focused. Um and

33:03

like really as the models get better,

33:04

you can kind of like iterate on a better

33:06

model with better data and so on. So

33:08

like a lot of what we see and like a lot

33:10

of what our partners are using the model

33:11

for like generating, you know, binders,

33:13

antibodies, so on. Like we're also like

33:15

doing a lot of dog fooding in-house.

33:17

Like we have a whole science team that's

33:18

using the models and just trying to

33:19

understand what they can do. And a lot

33:21

of that comes back to like, okay, now

33:23

that like the model now that we've

33:24

unlocked use case X, like what type of

33:26

new data can we generate? How can we

33:28

make the model better that way? Um and

33:30

really like that's been the long-term

33:31

vision of Chai is like we never thought

33:33

of like we're going to build this one

33:34

model that's going to solve all of this.

33:36

Like we know that there just like uh in

33:38

other fields like there're going to be

33:39

multiple iterations of this model. And

33:41

as you get better and better models,

33:42

like you get this flywheel effect, I

33:44

guess.

33:45

>> Well, and you mentioned what data can we

33:46

generate? Can we touch on data for a

33:47

minute? You guys can't exactly just go

33:49

script the internet and have all the

33:50

data you need to build your models.

33:52

Where does the data come from? How does

33:54

it kind of compound over time? Can you

33:57

just say a word about

33:58

data as an input to your models?

34:00

>> Yeah, so uh I think like

34:03

the primary source of data or like the

34:05

the gold standard source of data is this

34:07

like protein database. Um so this is

34:09

like it it's actually just like ac- like

34:12

legit lab scientists like who since like

34:15

1970 have just been depositing crystal

34:18

structures of like proteins and other

34:19

molecules over time. Uh and like really

34:22

without that like structure prediction

34:24

and design wouldn't have been a thing.

34:26

Um so there are these uh so that's like

34:28

one source of structural data. When I

34:30

started in the field, I really took like

34:31

a structure-pilled approach. I was like

34:33

super interested in predicting

34:34

structure, like how do you think about a

34:36

machine learning model that can output

34:38

3D coordinates? Like that's that's not

34:40

doesn't look like an LLM, it doesn't

34:41

look like an image model. This is really

34:42

in its own class. Josh, interestingly,

34:45

was taking like the exact opposite

34:46

approach. Uh so, he he was like one of

34:48

the original authors on ESM, and that

34:50

was like a really like a seminal work in

34:52

understanding how to apply language

34:53

models to protein sequences. Uh what's

34:55

really cool about that work is if you

34:57

can train a language model to understand

34:58

protein sequences, what ends up

35:00

happening is it ends up kind of like

35:02

representing the 3D structure

35:03

internally. And there's like a really

35:06

interesting reason why this happens. So,

35:07

like in order to predict like missing

35:09

amino acids, so like same way that you'd

35:11

predict like next word in a sentence, I

35:12

want to predict next amino acid in a

35:14

protein. In order to do that

35:16

effectively, like you really need to

35:17

understand like, okay, what does the

35:18

amino acid's like immediate

35:20

microenvironment look like? Uh cuz that

35:22

kind of tells you, okay, what are like

35:24

the compatible amino acids with

35:25

everything surrounding it. And in order

35:27

to do that well, you need to understand

35:28

the protein's 3D shape. Um so, I was

35:30

taking like this this really

35:31

structure-based approach, and Josh was

35:33

even more bitter lesson-pilled than me.

35:34

He's like, "We're just going to like

35:35

look at the sequences, and this is just

35:37

going to emerge." Um so, yeah, the two

35:39

main sources of this data, again, like

35:41

protein database for structures, and

35:42

then like these massive massive, maybe

35:44

even like order of like trillions of

35:46

tokens, sequence databases. Um and what

35:48

you can do once you have really good

35:49

models is just run them on the sequence

35:51

databases to get new structures out. Um

35:53

so, again, like you have this

35:55

compounding effect. As your models get

35:56

better, they get more and more accurate

35:58

at predicting these structures, uh and

35:59

then you have more and more training

36:01

data for the next series of models.

36:02

>> So, nice people ask us at Chai about,

36:04

these days at Chai, like which which

36:06

paradigm are are you actually going

36:07

after? Like one of the things I love

36:09

about our team is that it's actually

36:11

just like neither. We're very pragmatic.

36:12

Like we want to solve the problem. We

36:13

don't really care is it a sequence

36:14

approach, is it a structure approach? In

36:16

practice, it's going to end up being

36:17

like both, of course. Um I think you'd

36:19

be surprised, but there there might even

36:21

be more like biological sequence tokens

36:22

on the internet than like English

36:24

language tokens.

36:25

Uh and a lot of that data is not that

36:27

useful. It might not be redundant. It

36:29

might be very noisy. But there's a lot

36:30

of data out there. There's a lot of art

36:32

and like bringing it together. And I

36:33

think also the exciting thing is as the

36:35

models have gotten into a point now

36:36

where we can design things in the lab,

36:38

write to new targets for instance, we

36:39

can actually use the models to generate

36:41

data as well. So there's a lot of

36:42

exhaust from like all the experiments

36:44

that we're doing at Chai, which also

36:46

helped to make the models even better.

36:47

So it's a similar kind of takeoff that

36:49

we saw with LLMs. Like when I was at

36:50

OpenAI, we worked on

36:52

reinforcement learning of like a GPT-1

36:53

architecture. Like did not work models

36:55

weren't like good enough. But once the

36:57

base model got good enough, then you

36:59

could start to do those kinds of

37:00

experiments. And And I think there's a

37:01

similar analogy that's starting to

37:02

happen in our world now as well, where

37:04

the models have reached a point where

37:05

there's actually like a renewed interest

37:07

in data and like how do we actually

37:09

bring the models into the loop on like

37:10

making that happen. I think that creates

37:12

another really like interesting cycle on

37:13

on compounding improvement of the

37:15

models.

37:15

>> Yeah. We're talking a bit about

37:17

compounding improvement in the models.

37:18

And you said something about how pharma

37:20

as an industry is an incredibly

37:22

competitive landscape.

37:24

And it's interesting cuz I think there's

37:25

been a renowned interest in using AI and

37:27

using

37:28

ML to generate proteins and molecules.

37:31

And so your arena has actually become

37:32

quite competitive.

37:34

And you guys have obviously done an

37:35

amazing job at staying at the frontier.

37:37

It's been the year of deployment for

37:38

you. You've locked up a number of pharma

37:40

partnerships that are making your models

37:41

better. But how do you guys think about

37:43

the competitive landscape

37:45

and staying at the front of the

37:46

frontier?

37:47

>> Well, first of all, I think it goes back

37:48

to like looking at the results and not

37:50

fooling yourself and being rigorous. So

37:52

one of the reasons why we do a lot of

37:53

evaluation of the models. We mainly do

37:56

it just to hill climb in the models

37:57

itself. I think if you look at many of

37:59

the capabilities that we've brought in

38:00

have kind of been like first in the

38:02

field. You look at our Chai 2 model,

38:03

right? Like getting to success rate of

38:05

the models that you didn't have to do

38:07

large library screening anymore to see

38:08

results.

38:09

Couple months later showing how we could

38:11

bring in a lot of these developability

38:12

properties we've talked about before

38:14

like manufacturability of the molecules.

38:16

So in in many cases we are pushing the

38:18

model forward and and trying to see

38:19

these these

38:20

emerge. And then we try to quickly like

38:22

lock those in. Like, how do we how do we

38:24

make those capabilities like even more

38:26

pronounced that they become production

38:27

ready and we can like ship them to our

38:29

customers. I think one of the one of the

38:31

things I like to tell the team is that

38:33

you know, it's not like we are

38:34

head-to-head like with other model

38:35

providers or something like that. We're

38:36

all of us I think are working against

38:38

nature. Like, nature's actually been a

38:40

pretty good baseline. People have done

38:41

drug discovery a certain way for a long

38:44

time. And you know, Matt talked about

38:45

how like maybe you don't want to add

38:47

like module 24 to the Chi model. But

38:49

people have added like module 240 to

38:51

like the existing wet lab protocols. And

38:53

they have been like tuned quite

38:54

considerably.

38:56

One of the scientists on our team, Andy

38:57

Young, was

38:59

he was one of the first people working

39:01

on yeast display at MIT actually like

39:03

two decades ago. And he's got 20 years

39:04

experience in like Pfizer and Genentech

39:06

like really honing in these methods. Has

39:08

a drug approval to his name and

39:09

antibodies. And I think you look at

39:11

someone like that and like, you know,

39:12

Andy knows how to make a good antibody

39:14

with existing tools. And and that is

39:16

actually the bar that we need to clear.

39:18

Now of course, I think the ceiling on AI

39:19

is going to be a lot higher than what

39:20

we've managed to do before. Otherwise,

39:22

what would be the point of of

39:24

doing this? We didn't start the company

39:25

just to make, you know, a 10 times

39:27

faster mouse, right? We started this to

39:29

make breakthrough medicines that that

39:30

weren't possible before.

39:32

But ultimately, that is the bar that we

39:33

needed to clear in order to get

39:35

adoption. I think we hit that inflection

39:37

point a couple months ago. That's why

39:38

you've seen a lot of these big pharma

39:39

announcements.

39:41

But now we just need to continue to hone

39:42

in on like making these things even

39:44

better.

39:44

>> For somebody who's listening who thinks,

39:46

"Wow, this sounds pretty cool. I wonder

39:48

what it would be like to work at Chi."

39:50

What is the best thing about working at

39:53

Chi? And what is the worst thing about

39:55

working at Chi?

39:56

>> Yeah, I can I can speak to some of this.

39:59

I'll think of this on the fly. The worst

40:00

thing

40:01

>> [laughter]

40:02

>> the So I think like the the best thing

40:04

is just like how actually like

40:06

mission-driven everybody is. Like

40:07

everyone is so dedicated to what they're

40:09

doing. I've worked at other companies.

40:11

Like the closest I've ever seen to this

40:14

is like maybe some of the guys in my PhD

40:16

lab. Um, but like everyone is just like

40:18

incredibly incredibly motivated. We all

40:20

work really hard. There's like an

40:22

obvious shared goal.

40:24

Um, and I think that's really rare to

40:25

see and I think this goes back to just

40:27

like the focus that we've had since the

40:29

beginning and like we've always had like

40:31

kind of a clear philosophy, a clear plan

40:32

on how things are going to get there,

40:34

how things are going to get better. And

40:35

really like everyone at Chai is very

40:37

bought into this. It's pretty amazing

40:38

just to see like the amount of

40:39

dedication that that everyone's putting

40:41

in. Uh, least favorite thing about Chai,

40:44

uh, not directly on top of dandelion

40:46

chocolate, maybe.

40:47

>> [laughter]

40:49

>> No.

40:50

>> Maybe the next office.

40:51

>> Yeah.

40:52

Uh, actually, so probably the least

40:55

favorite part now, uh, is just like

40:58

I guess it's getting things to work at

40:59

scale.

41:00

Um, and really I I didn't even know what

41:02

that meant. Actually, when we started

41:04

Chai, we had 128 GPUs and I was like,

41:06

this is like the most scale possible for

41:09

a lab. Like this is crazy. Uh, I came

41:11

from like, you know, my PhD group where

41:13

we there were four of us sharing eight

41:15

GPUs and I was like, I just felt GPU

41:17

rich. It was crazy. Um, so now like, you

41:20

know, at Chai like we have a lot more

41:22

infrastructure to maintain. We have a

41:23

lot more compute resources. Like

41:24

luckily, GPUs are parallelizable. Um,

41:27

but that also kind of brings up its own

41:29

set of problems. So just like how do you

41:30

keep a cluster healthy over time? How do

41:32

you get that large training run? So like

41:33

how do you keep that running for for

41:35

months on end? Um, and even when it does

41:37

die like how do you automatically resume

41:39

these things? How do you keep all of

41:40

your communication down? How do you

41:41

optimize the models and make the best

41:43

use of the resources that you have? So I

41:45

think these are a lot of problems that,

41:46

you know, they continuously pop up.

41:48

They're good problems to have, but I

41:49

think they're they're also really

41:51

difficult to solve. Um, yeah, and I I'm

41:53

just excited to work on this.

41:55

>> Yeah, we I remember one of the CEOs

41:57

that, uh, we worked with a couple times,

41:58

Frits van Slooten, had this line about

42:00

you either have the pain of failure or

42:02

the pain of growth. You'd much rather

42:03

have the pain of growth.

42:04

>> Yeah, [laughter] that's exactly right.

42:06

>> I think my favorite part is is probably

42:08

the results. Uh, and that that sounds a

42:10

bit cliché, but there's nothing like

42:13

It's It's working, right? And just like

42:14

knowing that you're, you know, uh I

42:16

think many of us in the company, right?

42:17

Like we've been working in AI for a long

42:19

time, right? There's all this experience

42:20

you built up. And And to know that

42:21

you're applying it to something that

42:22

really matters, like just even I think

42:24

just take Matt and and and I like we've

42:26

been working on this problem for like 10

42:27

years, right? And a couple of years ago

42:29

I'm looking at something like, "Yeah,

42:30

we're writing some cool papers. Like

42:31

everyone is like celebrating this. Are

42:33

we actually making the world better?

42:34

Like is this actually going to impact

42:35

some patients?" And I think now the

42:37

answer is like, "Actually, yes. Like we

42:39

have reached the point where this is

42:40

going to make a big difference in the

42:41

world." And every time you get one of

42:43

these breakthrough results, anytime

42:45

there's a new feature on the product

42:46

that makes our lives of our our

42:47

customers easier, whenever there's like

42:49

new lab results uh coming back uh from

42:51

the science team, it's just always so uh

42:54

honestly it's exhilarating to realize

42:56

like this is actually going to change

42:57

the world in a pretty profound way.

42:59

I think that's also then uh maybe it

43:01

comes to the least favorite side, right?

43:03

Like, you know, we're running the

43:04

company and it's like we have real

43:05

partners that are relying on this. And

43:07

like things have to work, right? And you

43:09

ship a new model generation. How do you

43:11

make sure that there's no bugs in that?

43:12

How do you make sure you don't have

43:13

regressions, right? This is no longer

43:14

just like a Again, the blue sky research

43:16

problem of like, "Oh, we got some cool

43:18

results and we move on." Uh we've had to

43:20

have really high priorities on like, you

43:22

know, having production-level code

43:23

bases. As the team grows, how do we make

43:25

sure that the code base is in a state

43:26

that more people can contribute to this.

43:27

So, um something that that one of our

43:30

other co-founders, Zach, likes to say is

43:31

that if you want to move fast in in the

43:33

long term, you sometimes have to just

43:35

move a little bit slower in in the short

43:36

term, right? And And make sure that you

43:38

are building something uh Again, that

43:40

goes back to that compounding idea. Goes

43:41

back to we don't add module 25 to make

43:44

the next thing work. Uh so, sometimes uh

43:46

you know, you're you're like so excited

43:48

to get to the next result and you just

43:49

want to jump into it, but uh we have

43:52

real partners, some of the biggest

43:53

companies in the world that are now

43:54

relying on us. Uh and it's important

43:56

that we we realize that, we take that

43:58

responsibility to heart, uh and we make

44:00

sure we're building systems that, you

44:01

know, continue continue to work.

44:03

>> I have a burning question. Why Why guys

44:05

called Chai Discovery?

44:07

>> Chemistry and AI. But we love chai tea

44:09

as well.

44:10

>> [laughter]

44:11

>> There's a lot of chai tea stuff in the

44:12

office. So yeah.

44:12

>> That's a good question. I didn't know

44:13

that

44:14

either.

44:14

>> Josh is the visionary. That was all him.

44:16

>> [laughter]

44:16

>> It is a very user-friendly name.

44:18

>> Yeah, we also wanted a name that like

44:19

like biotech companies have such

44:20

complicated names. We wanted something

44:22

that's going to be much simpler. We're

44:24

trying to build make this whole thing

44:25

simpler, right?

44:26

So we need a simple name to go along

44:28

with it.

44:28

>> Awesome. What are you guys most excited

44:30

about in the next 6 to 12 months?

44:33

>> I think for me it's just the deployments

44:36

that are happening. So we've announced a

44:37

couple of these these partnerships and

44:39

I'm really excited just to hear about

44:42

the results that our partners are

44:43

bringing online. It goes back to this

44:45

point of like making a real difference

44:47

and also why I'm

44:48

so happy to see how these partnerships

44:50

are going even, you know, posted the

44:52

agreement and as we're working with

44:53

these folks that models are not sitting

44:55

on a shelf somewhere. Like they're

44:56

actually being used on real programs.

44:59

People trying to approach devastating

45:01

diseases where, you know, if Chai could

45:02

give them a molecule that that works it

45:04

could really change the lives of

45:05

patients. So I'm really excited to see

45:07

how that goes. Just the pace of progress

45:10

here is is incredible but also just like

45:12

the pace of the models. Like a year ago

45:14

you could not zero shot a molecule and

45:15

like, you know, have a good sense that

45:17

your program was going to work. Like now

45:19

that's changed. Someone might zero shot

45:20

a molecule and be like I I think we're

45:22

going to bring this program to the

45:22

clinic now. And then a year later, you

45:25

know, they might even have some of those

45:26

first molecules going into patients. And

45:29

just the speed of that is is just

45:31

incredible and it's you know, sometimes

45:33

you get some shivers thinking about this

45:35

stuff like, okay, my model is going to

45:36

like Matt has some patents from our last

45:38

company about, you know, like just

45:39

generating the molecule on the computer

45:41

and these things are are now in patients

45:42

and just to think about the scale. Like

45:44

I don't know a few years from now, do we

45:46

have dozens? Do we have hundreds of like

45:47

Chai molecules going into people? It's a

45:50

bit mind-blowing to think about what

45:51

that might might look like for patients.

45:53

>> I think one of the things that like

45:55

again like what made Chai unique like

45:56

the dedication. Like people are like,

45:58

man, you work a lot. Um and like you

46:01

like aren't you burnt out or whatever.

46:02

It's like it's actually really easy and

46:04

like it's very motivating when you're

46:06

making progress. Seeing the progress

46:08

that we're making and just like thinking

46:09

man the next model is going to be even

46:10

better than the last. We identified this

46:12

new thing so on.

46:14

Um like that's so incredibly motivating.

46:16

For me it's more of just like what what

46:18

can we unlock next and like how do we

46:19

make these things more controllable and

46:21

like when someone comes to us with a

46:22

certain target instead of just like

46:24

hoping we get good affinity or something

46:26

like that, you know, can we actually

46:27

control this? Can we say like we want

46:29

exactly a 10 animal or binder or things

46:31

like that. Like there are a lot of

46:32

technical things that I think are we're

46:33

like kind of right on the brink of

46:35

solving. Um and for me it's like it's

46:37

really motivating just to like pin those

46:38

things down and just like get all of

46:40

this over the line and see kind of where

46:42

that leads to next.

46:43

>> Awesome. Matt, Josh, thank you for

46:45

Engineering Biology and thank you for

46:47

sharing your story with us today.

46:49

>> Thank you guys. Thanks for having us.

46:53

>> [music]

47:16

[music]

Interactive Summary

Chai Discovery aims to transform drug discovery into an engineering discipline using AI, moving from trial-and-error to intentional design of molecules. Their core principle is simplicity, which they believe is crucial for scaling models and research. Building on breakthroughs like AlphaFold 2 and diffusion models, Chai's models can now generate protein structures and sequences with real-world constraints. Their Chai 2 model has increased molecule binding success rates from 0.1% to 15%, enabling *de novo* design of molecules with therapeutic properties. Chai operates on principles of rigor and scaling, building models from scratch and maintaining honesty in progress evaluation. Their business model involves partnering with pharma companies to provide this infrastructure, which ensures rigorous model validation and broad industry impact. Looking ahead, they envision a future where computer-aided design leads to higher-quality, highly specific drugs, addressing currently untreatable diseases and making personalized medicine economically viable.

Suggested questions

8 ready-made prompts