HomeVideos

Post-Training Is How You Keep Your Taste | Fireworks CEO Lin Qiao

Now Playing

Post-Training Is How You Keep Your Taste | Fireworks CEO Lin Qiao

Transcript

718 segments

0:00

Okay, up next we have Lynn, uh close

0:02

close friend and collaborator of

0:03

Brendan's. Um Lynn, uh

0:06

actually show of hands, who here who

0:08

here does post training?

0:10

Okay, good amount of the room. And who

0:11

here uses Firework? Okay, good amount of

0:14

the room, too. So, Lynn, you have you

0:15

have some friendlies in the audience.

0:17

Um and so for this next talk, what we're

0:19

going to do is we're going to focus on

0:22

post training.

0:23

Um Lynn, si- similar setup to to

0:26

Brendan's talk. We're going to do 15

0:27

minutes on how to approach the problem

0:29

of post training, um and then 15 minutes

0:32

or so for Q&A. Uh and Lynn, really

0:35

delighted to have you here. Thank you

0:36

for joining us.

0:37

>> Thanks for having me. Uh hi everyone,

0:39

good morning. I'm Lynn, I'm CEO and

0:41

co-founder of Firework. So, as we bring

0:44

up the slide, today I'm going to

0:46

do a little bit uh deep dive on post

0:49

training and why post training could be

0:50

very relevant to uh you building your

0:52

own business. So, first of all, a little

0:54

bit like prior,

0:56

I think um this year we have seen So,

0:59

first of all, I'm a Firework weather

1:01

specializing intelligence platform.

1:03

There are tons and tons of application

1:05

built on top of us, all the way from

1:07

startups to digital native and

1:09

enterprises. So, we get uh

1:11

um the fun part of my job is we get to

1:12

see a lot of patterns.

1:14

Um what are the innovation people are uh

1:16

building on top of us and uh and how

1:19

they are what are challenges that are uh

1:22

they're facing and what are trend um

1:24

people being developers are um

1:27

building on top of us. So, uh one of the

1:29

things, especially in the past 1 year,

1:31

software development and application

1:33

development has been somewhat disrupted

1:36

because it goes fine in the past, um you

1:40

have good idea, you want to implement

1:42

it, scaling production, it requires a

1:45

team of tens of very strong product

1:47

engineers, PMs working together,

1:49

multiple quarters to deliver that. And

1:51

right now, with one person, a few weeks,

1:54

uh without understanding how to write a

1:57

single line of code, you can do that.

1:59

So, that collapsing of resource required

2:02

both in terms of timeline and deep

2:04

expertise is shifting how how

2:08

competitive the application space is,

2:10

and it's shifting people from thinking

2:13

about building on top of off-the-shelf

2:15

off-the-shelf black box API to build a

2:18

much deeper mode. So, that they can

2:20

build a much durable business.

2:23

As also you have seen lot of discussion,

2:25

especially in the past 1 week, between

2:28

open model and the closed model and all

2:31

the rallying and support across the open

2:33

model.

2:34

The depth of that

2:37

that alliance is because we believe I

2:40

mean the industry

2:42

is much deeper because

2:44

look at the whole entire industry, there

2:45

are so many companies, right? There are

2:47

so many company, all of you are building

2:48

your own company.

2:50

Every single company exists for a reason

2:53

because they focus on solving a unique

2:56

problem in a special way.

2:58

And that means they carry their own

3:01

judgment, taste,

3:04

and determination, conviction into that

3:06

product, and that's why company exists.

3:09

Today, if you build on top of

3:11

off-the-shelf API, then

3:14

you really need to think about how you

3:16

keep that special taste judgment

3:19

and unique part forward. And we believe

3:22

one approach for every company to build

3:25

a build durable business is to actually

3:28

bake your judgment, taste, and customer

3:31

deep understanding into the intelligence

3:33

you build on top of instead of just a

3:36

off-the-shelf API. So, that's kind of a

3:38

little bit context of

3:40

what's happening in industry, what we

3:42

are seeing, and why post training could

3:44

be very very relevant to you.

3:52

Okay. So, you probably heard a lot about

3:56

owning intelligence not rent.

3:58

And what does owning your own

4:00

intelligence mean? It actually means

4:01

many things. So, first of all, it start

4:04

from data.

4:05

Intelligence is derivative of data.

4:09

And

4:10

obviously the the foundation labs, the

4:14

all the foundation model we're using are

4:16

building on top of the public data and

4:19

the label data that

4:22

is has solved common tasks, but all of

4:25

you are solving a specific task. That's

4:28

why you you're building a business,

4:30

you're building a company, and be able

4:32

to curate production data with high

4:35

quality and even generate synthetic data

4:38

to enrich your production data is one

4:41

step of owning your own intelligence.

4:44

And then after you have data, you will

4:46

start to kind of use that data turn into

4:48

a model,

4:49

build on top of existing model and own

4:52

the weights. And there are a collection

4:54

of techniques you can use to get there.

4:57

Those techniques are tailored to solve

4:59

different kind of problem you could

5:00

possibly have. And those techniques can

5:03

also interoperate with each other for

5:05

you to build a reach your final goal.

5:08

And then after you have a great model

5:10

belong to yourself solving your specific

5:12

problem really well, and then you need

5:13

to work on serving it.

5:16

You probably first will do a some AB

5:18

testing to make sure it really move the

5:20

needle for your product metrics,

5:23

and then goes back in in this loop.

5:28

Obviously, I don't think you should jump

5:31

into post training right away.

5:33

And and there are different phases you

5:36

will go into.

5:37

First, prompt everyone start from

5:39

prompt, use the model as is. It's few

5:42

shots. If quickly you can use that to

5:44

test your ideas.

5:46

And then you use rag to ground

5:48

um,

5:49

the usage of AI into your own data and

5:53

you can start to do a lot of contest

5:54

engineering from that. Those are from

5:56

minutes of interaction to hours of

5:59

interaction. And then you first progress

6:01

into hey, I have some data. I want to

6:03

see how my data is going to reflect and

6:05

make uh, the model work better with my

6:08

product. So you will start to do

6:10

supervised fine-tuning that will take

6:12

you a few hours to um, hey, um, the the

6:16

model I want to reflect in personalized

6:19

taste. It is it is very unique uh,

6:22

choice of my products. So so therefore

6:25

you want to start to use preferences uh,

6:29

information where you collect from user

6:30

interaction and uh, whether thumbs-up,

6:34

thumbs-down, a lot of those kind of

6:35

information and uh, help the model learn

6:38

your product taste. Um, and and finally,

6:41

you want to build a model towards

6:43

understand carrying your expertise in

6:46

that domain. Whether that expertise is

6:49

across

6:50

uh, legal, finance, healthcare, customer

6:53

support, recruiting, marketing, sales,

6:56

you name it. Even in one industry

6:59

vertical, there are so many subdomains.

7:01

So all of that is unique and special

7:04

towards the product you're building. So

7:06

usually um, you probably heard a lot of

7:08

reinforcement learning and that is to to

7:11

actually build towards a specialty.

7:13

Um, so this progression is very similar

7:16

to how we human being learn knowledge

7:18

over time. Uh, for example,

7:21

uh, we actually run uh, learn a lot of

7:24

knowledge by reading uh, reading

7:26

literature, right? In the literature it

7:28

will say, hey, what is correct, what is

7:30

it not correct? So this is a very

7:31

similar to supervised fine-tuning. Um,

7:34

and uh, as we grow, we develop our own

7:37

taste and judgment of how we want to

7:39

conduct the

7:41

specific way of we want to approach a

7:43

problem, and that is preference or DPO.

7:47

Um and over course of time, we learn to

7:49

be a doing a really good job at certain

7:52

area, whether um hey, I want to be

7:54

accounting accountant, and I would

7:56

really know how to kind of um

7:59

build into the financial financial data,

8:02

or I want to be a specific set of um I

8:04

want to be a dentist, and you learn the

8:06

details of how to operate

8:08

um with dentistry. So, all this is very

8:11

similar exactly um to how we human being

8:14

acquire knowledges.

8:18

So, um and also very interesting

8:20

different techniques are there to solve

8:22

different kind of problems.

8:24

Um for example, if the model doesn't

8:26

know um the fact, and the fact actually

8:29

is very dynamic, the facts of your data

8:32

is uh in in your product is very

8:34

dynamic, and then typically you use rag

8:36

to solve that problem.

8:38

Um however, if your model output all the

8:42

behavior

8:43

uh or um or the structure is off, then

8:48

you you give you curate data and do

8:50

supervised fine-tuning to correct that.

8:53

Um if your model's answer, the quality

8:57

is is personal or kind of is specific to

9:00

the taste of your product,

9:03

um and then you use preference and

9:05

tuning.

9:06

Or the model is quite weak

9:09

on the special problem you're trying to

9:11

solve, then you use RL uh reinforcement

9:14

learning.

9:15

And the end result of the model is too

9:18

slow or too expensive to for you to

9:21

serve in production, and then you use

9:23

distillation to um go

9:26

um let the teacher teach the student

9:28

much smaller student model, so it can be

9:30

more much more performing or economical.

9:33

Obviously, distillation also means

9:35

different things for LLM, VLM, or uh

9:38

image generation models. Um if you are

9:40

interested, we can talk more about that.

9:44

So, there are many different ways um

9:47

teams start to hey, I want to kind of

9:49

try those technology, and um and they

9:52

may not be happy uh

9:55

of the experience because they can burn

9:57

money um and the time, but may not reach

10:00

their uh ideal results. So, so here are

10:02

the areas that you can possibly uh feel

10:05

frustrated. Uh for example, um

10:09

when it comes to data, and data is the

10:10

essence of tuning, it the quantity is

10:14

not the most important. Actually,

10:16

quality is the most important. But,

10:17

sometimes just by throwing uh tons and

10:20

tons of data into a training process may

10:22

not leading to a great outcome. So, you

10:25

really need to control the data quality

10:27

and have And usually, who's the best

10:30

judge of data quality? It's actually

10:32

your product team. Um so, that's where

10:34

we see the convergence of um before

10:36

GenAI,

10:38

product team and the research team or ML

10:40

team, they're separate organizations.

10:42

They're separate team. And they work

10:43

hand-in-hand to make things happen. And

10:46

a lot of time these days, when um people

10:49

post-train

10:50

uh GenAI models, we see kind of the

10:52

convergence of the product team needs to

10:54

make judgment call of the data quality

10:56

and start to be deeply involved in the

10:58

process. So, uh they can ensure uh the

11:01

best result.

11:02

Um and the second is you going to have

11:05

evals. Um I know all of you are very

11:08

busy, and uh part into launch quickly,

11:11

and a lot of evals is a vibe evaling.

11:14

>> [laughter]

11:14

>> And the founders uh really look at the

11:16

result and feel hey, it is it right or

11:19

not? Actually, this is judgment. This is

11:21

you put your judgment of uh of end

11:24

result and decide whether it's good or

11:26

not. And that judgment you convert into

11:29

systemic evaluation. So, this is no

11:31

different from traditional software

11:33

development where you have your unit

11:35

test, integration test to ensure

11:37

quality. Similarly,

11:39

if you want to if you think about doing

11:41

post-training and then have a way to

11:42

build the eval and build your own

11:44

judgment into a repeatable process, it's

11:47

extremely important.

11:48

Uh and then there could be when you do

11:50

RL, there could be sloppy RL environment

11:54

uh where you build a simulation and the

11:56

simulation is not really reflecting the

11:58

reality and then the model can heel

12:01

climb on on kind of the a bad um

12:05

simulation and I mean it it it will it

12:07

will also could possibly do reward

12:09

hacking

12:10

um and all kind of weird stuff. So, a

12:12

fun story about reward hacking is

12:15

um we have been asking a model to uh to

12:17

generate to do this is coding example uh

12:20

to create um

12:22

to to generate code that minimize

12:26

uh the the error compilation error.

12:28

Okay. So, guess what the model did? The

12:30

model generate zero line of code.

12:33

Okay, there's no compilation error, but

12:35

that's absolutely not what you want. So,

12:36

this is a kind of one example of reward

12:38

hacking, it's very common because model

12:40

was very smart, it'll try in all

12:42

different way to

12:43

uh to get um your goal, but it may not

12:46

be what you want. So, uh pay attention

12:48

to all these details and uh and try not

12:51

to let the model outsmart you.

12:53

Um obviously,

12:55

between um your experimentation, so

12:57

think about um your development process

13:00

as hey, you're actually doing a lot of

13:02

experimentation from training the model

13:04

to uh training the model is not the end

13:07

of of your experiment.

13:09

The judge is the final judge is whether

13:11

product matches is moving or not. So,

13:14

you need to bring the final model into

13:16

into serving tier and do AB testing. Um

13:19

and that transition is extremely

13:20

important because um the quality can

13:23

drop if you move from one training stack

13:26

to a serving stack without aligning

13:28

these two. Because think about the model

13:31

is tons and tons of calculation, math,

13:36

and the matrix multiplication. Um and

13:39

the way to do multi matrix

13:40

multiplication and calculation, if you

13:43

use different library, different

13:44

numerics,

13:46

um and different optimization, it will

13:47

lead different results. And therefore,

13:50

the

13:51

the end result of a training may not be

13:53

replicated or even uh you lose precision

13:56

during serving. So, that alignment is

13:58

very important. I can give you more

14:00

examples of that.

14:02

Um and there are a few other um

14:04

challenges you will run into and happy

14:06

to talk with you more about in details

14:09

offline.

14:11

So, there have been many pioneers

14:14

uh we worked with post-training. I would

14:17

say Cursor is one of the few. They have

14:20

started onboarding getting onto this

14:23

train from the beginning of last year.

14:25

Uh there are multiple reasons. One is

14:27

they really want to control their

14:29

destiny of the supply of the model. And

14:32

obviously,

14:33

um they have a lot of they have a lot of

14:36

user engagement. They understand um how

14:39

what is customer preferences, a lot of

14:41

data. So, that becomes the beginning of

14:43

their journey. And they do um pretty

14:46

deep mid-training to post-training. And

14:48

you have heard them announce um

14:50

continuously newer models every few

14:52

month. Um and uh and the result is

14:56

great. Um their inspiration is to

14:59

compete at the frontier quality. Uh very

15:01

bold aspiration. And they're getting

15:03

there. So, very impressive result from

15:06

composer two, composer 2.5, um on par

15:10

with um

15:11

They're always trying to be on par or

15:13

beat

15:14

um the frontier labs quality at the same

15:17

time. So, um they are kind of one of the

15:20

examples, a very vibrant example, is

15:23

completely doable as long as you stay

15:25

focused and have the right tool and

15:27

cursive build on top of us.

15:29

Um there's another there huge huge

15:32

variety of examples from healthcare. For

15:35

example, Doximity is one example where

15:38

they do clinic AI where

15:41

they basically let doctors ask deep

15:44

medical questions

15:45

matching symptoms to medication and side

15:48

effects and have well-rounded research

15:51

around everything

15:53

in medical space.

15:55

They actually also train on Fireworks

15:58

and they top a very important benchmark

16:01

which is Stanford Harvard clinic safety

16:03

benchmark.

16:04

And we are very proud of that result and

16:07

they they keep working on this training

16:10

loop. Um Factory, that's another Sequoia

16:14

company.

16:15

They build on top of Fireworks and

16:17

especially focus on security part of the

16:20

coding.

16:21

That is a very hard topic because

16:22

security is not

16:24

high tolerance. You need to get it

16:26

right.

16:27

They tune a model that also

16:31

top the kind of benchmarking security

16:33

area.

16:34

So you can see all these examples is a

16:36

demonstration. They solve all the

16:38

companies are solving a very unique

16:40

problem in a special way and they have

16:43

been able to kind of

16:45

build their own challenges use their own

16:46

data by building on top of open model.

16:51

And those are the details. Obviously, as

16:54

you know, last year 2025 is the year of

16:56

coding. There

16:58

Fireworks we support all the coding

17:00

companies build on top of us. A lot of

17:02

success in

17:04

mid-train to post-train

17:06

very strong models.

17:08

As you can see those benchmarks are very

17:10

impressive and all these coding

17:12

companies are inspired to

17:14

get on par or beat Frontier Labs in the

17:17

coding space, but it's not coding. And

17:20

we start This is the year we start to

17:23

see very interesting development in all

17:25

different kind of co-work space from

17:27

general purpose co-work to specialized

17:30

domain specific co-work as I mentioned.

17:33

There are like legal,

17:35

finance, marketing, recruiting, sales,

17:39

customer support co-work space. They are

17:40

all starting to post train and own their

17:44

own intelligence power their product.

17:46

Here,

17:48

Jen Spark is one of the generic co-work

17:53

application and they build deep research

17:56

for professionals and slide generation.

17:58

As you see,

18:00

this they compete with a frontier model

18:04

and it's slightly better, but the cost

18:07

is significantly lower. So here we're

18:09

talking about five to 10

18:11

10 times cost reduction. So

18:14

as a startup, you can once you hit a

18:16

problem like if it you can scale quickly

18:19

into a doable business. And I guess

18:22

another health care example and health

18:25

care data or use case is not well

18:27

captured

18:28

in

18:29

in frontier labs model. So you can see

18:32

the quality of tuned model is

18:35

significantly better than the state of

18:38

art closed model.

18:41

Well, we have many different type of

18:44

developers build on top of fireworks to

18:47

do post training. So what what are these

18:49

types?

18:50

So we have seen many frontier agent

18:53

builder.

18:54

So frontier agent builder, they usually

18:57

build bespoke customized harness.

19:00

They don't use common harnesses. And to

19:03

co-optimize their harness with their

19:06

systems and all different kind of tools

19:08

they want to build on top of

19:11

often time requires post training. Uh

19:14

so, we have we have seen a lot of

19:16

repeatable success on doing that.

19:19

And

19:20

obviously, we have seen a lot of big

19:22

companies

19:23

put to post training because

19:25

interestingly, those incumbents have

19:27

huge amount of traffic.

19:29

For them to deploy AI features across

19:31

the board

19:32

to all the all their customers is a huge

19:36

cost burden. Huge cost burden. Uh so, if

19:40

we say we we need to be careful not

19:43

scale into bankruptcy, it's not just for

19:45

startups. Uh it's actually for

19:47

incumbents. They literally their CFO is

19:50

blocking their AI feature launch because

19:52

of the cost. And the post training is

19:54

way to remediate that.

19:56

Uh and we have seen a lot of specialized

19:59

model operator, uh those could be uh new

20:03

uh

20:03

new labs and a lot of kind of

20:05

cutting-edge model develops also um do

20:09

significant post training.

20:13

So, uh this is a quote we have heard

20:15

repeatedly

20:16

uh that after part of market fit, post

20:19

training becomes the vehicle to for

20:22

many, many companies to build

20:23

specialized intelligence. We do believe

20:25

we do see the trend that in the future

20:27

there could be millions of specialized

20:30

models. One application per use case. It

20:34

it it it will be millions. That's how

20:36

uh I believe in that.

20:39

Um And obviously, there are

20:42

um

20:43

we have seen kind of different way to

20:45

convert the data into into model. Um and

20:49

the the signal from the reward and how

20:52

to build those reward function

20:55

is very important um part of the story.

20:58

Um so, our observation is

21:02

no,

21:03

start early, start early, and the start

21:06

to do experimentation a lot more

21:09

iteratively

21:11

and they get hands-on experience and we

21:14

we have seen people on board so quickly.

21:17

There there's actually not the barrier

21:19

as many people feel because this this

21:22

whole reward engineering is very similar

21:25

to software engineering.

21:26

The the the logical reasoning and

21:28

mindset is very very similar. So

21:32

so yeah so

21:34

get your hands dirty and and start to

21:37

kind of test it out.

21:42

Cool.

21:43

I think

21:46

Yeah so those are the kind of key

21:47

takeaways

21:49

and uh

21:50

I think we [snorts]

21:52

we we want to acknowledge there are

21:56

different specialties or different

21:59

knowledges you have right now in this

22:02

space. For example, we work with the

22:05

companies like Cursor and Cognition.

22:07

They have deep researchers. They want to

22:09

control every single knob as much as

22:11

possible to get extreme results.

22:15

We give them the lowest API.

22:17

That means

22:19

we just

22:20

have RL rollout that they can directly

22:23

interact with and they fully control the

22:25

trainer.

22:27

That gives them the best result.

22:29

But many of the company

22:31

you have product engineers or machine

22:34

learning engineers.

22:35

You have good enough information and

22:38

knowledge about post training. You want

22:40

to have some control but you do not want

22:42

to have the lowest level control. So

22:45

that's where we have our training SDK

22:49

iterate very closely with all of you to

22:52

get to to get you going and we also

22:57

are happy to kind of deploy our own

22:59

researchers. That's where we will see,

23:02

um, help hands-on like training your

23:05

team to to get on board. So, those are

23:07

the kind of different way we see

23:09

different teams want to engage and and

23:12

get familiar with the whole entire

23:14

process.

23:15

Cool. Uh, I'm going to pause here and

23:17

see if you have any questions.

23:19

>> I'll jump um

23:21

People are most most successful if they

23:22

have data and reward signal already,

23:25

maybe from their product. What are the

23:26

what are the best examples of reward

23:28

signals you've seen with

23:30

very successfully trained very quickly?

23:33

>> So, um, usually the rewards you can

23:35

think about rewards actually is code.

23:38

You write rewards in code. And uh,

23:41

usually think about rubrics of rewards.

23:44

Um, and think about uh, you want to

23:47

grade the result in multiple dimensions.

23:51

Uh, [snorts] for example, if you are

23:53

building a recruiting

23:55

agent,

23:56

uh, and then you can think about, "Hey,

23:58

how do we evaluate candidate selection?"

24:01

Um, and a different company, I guarantee

24:03

you you have different criterias. Uh,

24:05

for example, we want to grade a

24:07

candidate

24:09

um, um, aptitude.

24:11

Are they really hungry? They do not take

24:13

no as an answer. They they will break

24:15

down walls. And that's one matrix. The

24:18

second is they're really fast in kind of

24:21

um,

24:22

building things and making progress. And

24:25

and so on. So, then you have a blend of

24:27

score to blend to merge this. So, so

24:30

think about reward in different

24:32

dimensions of rubrics. And then

24:34

different company have a different way

24:35

to blend those. And that's that's your

24:39

um, unique part and secret sauce.

24:44

Oh, sorry.

24:45

>> Uh, great talk, Lynn. Raj from Vercel.

24:48

I'm curious about I think you mentioned

24:51

what are the pitfalls of starting too

24:53

early, but based on what you're seeing

24:55

from Vercel, Vercel cognition, or other

24:57

companies, when do they actually start

24:59

thinking about post training? Is it when

25:02

they feel ready? Is it when the cost is

25:04

now too much? Because the benchmarks

25:08

where they're beating it, that's a great

25:10

signal that yes, you can uh

25:12

um

25:13

you know, beat the frontier models, but

25:15

is that the primary motivator? Like that

25:17

is that 2% bump worth the cost or

25:19

investment into this whole framework?

25:22

And and how do they think about ROI? You

25:24

know, close models are more expensive,

25:26

but then post training has some

25:29

intercept of, you know, money and

25:31

investment and maintenance going

25:32

forward. So, I'm curious of like, are

25:35

they approaching when cost becomes a

25:36

concern or is it when the usage spiking

25:39

up and they really have to now plan for

25:41

a year or two in advance?

25:42

>> Yeah, this is this is excellent

25:44

question.

25:45

Uh so, usually so, think about uh in the

25:48

AI building

25:49

product market fit and the scaling the

25:51

business as actually two phases.

25:55

And then in the SaaS time, it's one

25:56

concept. You hit a product market fit,

25:58

just scale.

25:59

I know you got to scale as much as

26:00

possible.

26:01

Um but now we see a bifurcation.

26:05

Product market fit doesn't really mean

26:06

you have a durable scalable business.

26:08

And typically, we see uh companies

26:12

deploy the strategy of focus on product

26:13

market fit first by building on top of

26:17

um you know, uh Frontier Labs model

26:20

because you don't need to worry about

26:21

anything. So, just kind of spend your

26:23

money and hit a product market fit.

26:25

Uh and the other very important thing is

26:27

only after you hit the product market

26:28

fit, the data you collect from product

26:30

surface area are really meaningful.

26:33

A- and you will get uh also the volume

26:35

of high quality data start to collect

26:37

from um product surface area. That

26:40

became the fuel of uh you start to own

26:42

your intelligence. So, we um we see like

26:46

that as kind of the very strong

26:48

indication

26:49

um because once you hit upon market fit,

26:51

you really think about start start to

26:53

scale a business, you think about two

26:54

things. One is continue to keep your

26:57

competitive edge.

26:59

Two is

27:00

um build a build durable business, so

27:03

your um your revenue and your costs are,

27:06

you know, in a healthy state.

27:08

Um so, then um

27:12

post-training becomes a very appealing

27:14

solution because post-training allow you

27:16

to

27:17

um

27:18

basic encode codify um your uh unique

27:22

taste into into a model that no one can

27:26

uh steal from. Because it's very easy to

27:29

clone and copy application as is, as I

27:32

as you all know, right? Uh from coding

27:34

agent it's very easy. Uh so, from

27:36

screenshot, boom, generate the same or

27:38

even better app. Uh so, that's a very

27:40

scary that application itself is kind of

27:43

the a moat is is being um reduced.

27:47

Um and you bake your data uh where

27:50

reflecting the product engagement and

27:53

all the deep knowledge into your model

27:55

is a way to preserve that. And second is

27:58

you can post-train a model that bring

28:00

down the cost five to time 10 times. And

28:03

and then that means you can support five

28:06

to 10 times much higher traffic with the

28:08

same budget. And the your unit of

28:11

economics of scaling is so much better

28:13

than you avoid scaling into bankruptcy.

28:16

So, we see kind of those as as two

28:18

compelling story behind the reason and

28:21

timing.

28:22

>> Thank you, Lin.

28:23

>> Thanks a lot.

28:24

>> [applause]

Interactive Summary

In this talk, Lynn, CEO and co-founder of Firework, explores the role of post-training in building durable, intelligent businesses. She argues that as application development becomes easier, companies need deeper moats—specifically by embedding unique customer judgment, taste, and domain-specific knowledge into their own models. Lynn outlines a progression from prompting and RAG to fine-tuning and reinforcement learning, emphasizing that this process mirrors human learning. She also addresses common pitfalls, such as data quality, evaluation challenges, and reward hacking, while highlighting successful real-world applications of post-training in sectors like coding, healthcare, and finance to reduce costs and maintain a competitive edge.

Suggested questions

4 ready-made prompts