Post-Training Is How You Keep Your Taste | Fireworks CEO Lin Qiao
718 segments
Okay, up next we have Lynn, uh close
close friend and collaborator of
Brendan's. Um Lynn, uh
actually show of hands, who here who
here does post training?
Okay, good amount of the room. And who
here uses Firework? Okay, good amount of
the room, too. So, Lynn, you have you
have some friendlies in the audience.
Um and so for this next talk, what we're
going to do is we're going to focus on
post training.
Um Lynn, si- similar setup to to
Brendan's talk. We're going to do 15
minutes on how to approach the problem
of post training, um and then 15 minutes
or so for Q&A. Uh and Lynn, really
delighted to have you here. Thank you
for joining us.
>> Thanks for having me. Uh hi everyone,
good morning. I'm Lynn, I'm CEO and
co-founder of Firework. So, as we bring
up the slide, today I'm going to
do a little bit uh deep dive on post
training and why post training could be
very relevant to uh you building your
own business. So, first of all, a little
bit like prior,
I think um this year we have seen So,
first of all, I'm a Firework weather
specializing intelligence platform.
There are tons and tons of application
built on top of us, all the way from
startups to digital native and
enterprises. So, we get uh
um the fun part of my job is we get to
see a lot of patterns.
Um what are the innovation people are uh
building on top of us and uh and how
they are what are challenges that are uh
they're facing and what are trend um
people being developers are um
building on top of us. So, uh one of the
things, especially in the past 1 year,
software development and application
development has been somewhat disrupted
because it goes fine in the past, um you
have good idea, you want to implement
it, scaling production, it requires a
team of tens of very strong product
engineers, PMs working together,
multiple quarters to deliver that. And
right now, with one person, a few weeks,
uh without understanding how to write a
single line of code, you can do that.
So, that collapsing of resource required
both in terms of timeline and deep
expertise is shifting how how
competitive the application space is,
and it's shifting people from thinking
about building on top of off-the-shelf
off-the-shelf black box API to build a
much deeper mode. So, that they can
build a much durable business.
As also you have seen lot of discussion,
especially in the past 1 week, between
open model and the closed model and all
the rallying and support across the open
model.
The depth of that
that alliance is because we believe I
mean the industry
is much deeper because
look at the whole entire industry, there
are so many companies, right? There are
so many company, all of you are building
your own company.
Every single company exists for a reason
because they focus on solving a unique
problem in a special way.
And that means they carry their own
judgment, taste,
and determination, conviction into that
product, and that's why company exists.
Today, if you build on top of
off-the-shelf API, then
you really need to think about how you
keep that special taste judgment
and unique part forward. And we believe
one approach for every company to build
a build durable business is to actually
bake your judgment, taste, and customer
deep understanding into the intelligence
you build on top of instead of just a
off-the-shelf API. So, that's kind of a
little bit context of
what's happening in industry, what we
are seeing, and why post training could
be very very relevant to you.
Okay. So, you probably heard a lot about
owning intelligence not rent.
And what does owning your own
intelligence mean? It actually means
many things. So, first of all, it start
from data.
Intelligence is derivative of data.
And
obviously the the foundation labs, the
all the foundation model we're using are
building on top of the public data and
the label data that
is has solved common tasks, but all of
you are solving a specific task. That's
why you you're building a business,
you're building a company, and be able
to curate production data with high
quality and even generate synthetic data
to enrich your production data is one
step of owning your own intelligence.
And then after you have data, you will
start to kind of use that data turn into
a model,
build on top of existing model and own
the weights. And there are a collection
of techniques you can use to get there.
Those techniques are tailored to solve
different kind of problem you could
possibly have. And those techniques can
also interoperate with each other for
you to build a reach your final goal.
And then after you have a great model
belong to yourself solving your specific
problem really well, and then you need
to work on serving it.
You probably first will do a some AB
testing to make sure it really move the
needle for your product metrics,
and then goes back in in this loop.
Obviously, I don't think you should jump
into post training right away.
And and there are different phases you
will go into.
First, prompt everyone start from
prompt, use the model as is. It's few
shots. If quickly you can use that to
test your ideas.
And then you use rag to ground
um,
the usage of AI into your own data and
you can start to do a lot of contest
engineering from that. Those are from
minutes of interaction to hours of
interaction. And then you first progress
into hey, I have some data. I want to
see how my data is going to reflect and
make uh, the model work better with my
product. So you will start to do
supervised fine-tuning that will take
you a few hours to um, hey, um, the the
model I want to reflect in personalized
taste. It is it is very unique uh,
choice of my products. So so therefore
you want to start to use preferences uh,
information where you collect from user
interaction and uh, whether thumbs-up,
thumbs-down, a lot of those kind of
information and uh, help the model learn
your product taste. Um, and and finally,
you want to build a model towards
understand carrying your expertise in
that domain. Whether that expertise is
across
uh, legal, finance, healthcare, customer
support, recruiting, marketing, sales,
you name it. Even in one industry
vertical, there are so many subdomains.
So all of that is unique and special
towards the product you're building. So
usually um, you probably heard a lot of
reinforcement learning and that is to to
actually build towards a specialty.
Um, so this progression is very similar
to how we human being learn knowledge
over time. Uh, for example,
uh, we actually run uh, learn a lot of
knowledge by reading uh, reading
literature, right? In the literature it
will say, hey, what is correct, what is
it not correct? So this is a very
similar to supervised fine-tuning. Um,
and uh, as we grow, we develop our own
taste and judgment of how we want to
conduct the
specific way of we want to approach a
problem, and that is preference or DPO.
Um and over course of time, we learn to
be a doing a really good job at certain
area, whether um hey, I want to be
accounting accountant, and I would
really know how to kind of um
build into the financial financial data,
or I want to be a specific set of um I
want to be a dentist, and you learn the
details of how to operate
um with dentistry. So, all this is very
similar exactly um to how we human being
acquire knowledges.
So, um and also very interesting
different techniques are there to solve
different kind of problems.
Um for example, if the model doesn't
know um the fact, and the fact actually
is very dynamic, the facts of your data
is uh in in your product is very
dynamic, and then typically you use rag
to solve that problem.
Um however, if your model output all the
behavior
uh or um or the structure is off, then
you you give you curate data and do
supervised fine-tuning to correct that.
Um if your model's answer, the quality
is is personal or kind of is specific to
the taste of your product,
um and then you use preference and
tuning.
Or the model is quite weak
on the special problem you're trying to
solve, then you use RL uh reinforcement
learning.
And the end result of the model is too
slow or too expensive to for you to
serve in production, and then you use
distillation to um go
um let the teacher teach the student
much smaller student model, so it can be
more much more performing or economical.
Obviously, distillation also means
different things for LLM, VLM, or uh
image generation models. Um if you are
interested, we can talk more about that.
So, there are many different ways um
teams start to hey, I want to kind of
try those technology, and um and they
may not be happy uh
of the experience because they can burn
money um and the time, but may not reach
their uh ideal results. So, so here are
the areas that you can possibly uh feel
frustrated. Uh for example, um
when it comes to data, and data is the
essence of tuning, it the quantity is
not the most important. Actually,
quality is the most important. But,
sometimes just by throwing uh tons and
tons of data into a training process may
not leading to a great outcome. So, you
really need to control the data quality
and have And usually, who's the best
judge of data quality? It's actually
your product team. Um so, that's where
we see the convergence of um before
GenAI,
product team and the research team or ML
team, they're separate organizations.
They're separate team. And they work
hand-in-hand to make things happen. And
a lot of time these days, when um people
post-train
uh GenAI models, we see kind of the
convergence of the product team needs to
make judgment call of the data quality
and start to be deeply involved in the
process. So, uh they can ensure uh the
best result.
Um and the second is you going to have
evals. Um I know all of you are very
busy, and uh part into launch quickly,
and a lot of evals is a vibe evaling.
>> [laughter]
>> And the founders uh really look at the
result and feel hey, it is it right or
not? Actually, this is judgment. This is
you put your judgment of uh of end
result and decide whether it's good or
not. And that judgment you convert into
systemic evaluation. So, this is no
different from traditional software
development where you have your unit
test, integration test to ensure
quality. Similarly,
if you want to if you think about doing
post-training and then have a way to
build the eval and build your own
judgment into a repeatable process, it's
extremely important.
Uh and then there could be when you do
RL, there could be sloppy RL environment
uh where you build a simulation and the
simulation is not really reflecting the
reality and then the model can heel
climb on on kind of the a bad um
simulation and I mean it it it will it
will also could possibly do reward
hacking
um and all kind of weird stuff. So, a
fun story about reward hacking is
um we have been asking a model to uh to
generate to do this is coding example uh
to create um
to to generate code that minimize
uh the the error compilation error.
Okay. So, guess what the model did? The
model generate zero line of code.
Okay, there's no compilation error, but
that's absolutely not what you want. So,
this is a kind of one example of reward
hacking, it's very common because model
was very smart, it'll try in all
different way to
uh to get um your goal, but it may not
be what you want. So, uh pay attention
to all these details and uh and try not
to let the model outsmart you.
Um obviously,
between um your experimentation, so
think about um your development process
as hey, you're actually doing a lot of
experimentation from training the model
to uh training the model is not the end
of of your experiment.
The judge is the final judge is whether
product matches is moving or not. So,
you need to bring the final model into
into serving tier and do AB testing. Um
and that transition is extremely
important because um the quality can
drop if you move from one training stack
to a serving stack without aligning
these two. Because think about the model
is tons and tons of calculation, math,
and the matrix multiplication. Um and
the way to do multi matrix
multiplication and calculation, if you
use different library, different
numerics,
um and different optimization, it will
lead different results. And therefore,
the
the end result of a training may not be
replicated or even uh you lose precision
during serving. So, that alignment is
very important. I can give you more
examples of that.
Um and there are a few other um
challenges you will run into and happy
to talk with you more about in details
offline.
So, there have been many pioneers
uh we worked with post-training. I would
say Cursor is one of the few. They have
started onboarding getting onto this
train from the beginning of last year.
Uh there are multiple reasons. One is
they really want to control their
destiny of the supply of the model. And
obviously,
um they have a lot of they have a lot of
user engagement. They understand um how
what is customer preferences, a lot of
data. So, that becomes the beginning of
their journey. And they do um pretty
deep mid-training to post-training. And
you have heard them announce um
continuously newer models every few
month. Um and uh and the result is
great. Um their inspiration is to
compete at the frontier quality. Uh very
bold aspiration. And they're getting
there. So, very impressive result from
composer two, composer 2.5, um on par
with um
They're always trying to be on par or
beat
um the frontier labs quality at the same
time. So, um they are kind of one of the
examples, a very vibrant example, is
completely doable as long as you stay
focused and have the right tool and
cursive build on top of us.
Um there's another there huge huge
variety of examples from healthcare. For
example, Doximity is one example where
they do clinic AI where
they basically let doctors ask deep
medical questions
matching symptoms to medication and side
effects and have well-rounded research
around everything
in medical space.
They actually also train on Fireworks
and they top a very important benchmark
which is Stanford Harvard clinic safety
benchmark.
And we are very proud of that result and
they they keep working on this training
loop. Um Factory, that's another Sequoia
company.
They build on top of Fireworks and
especially focus on security part of the
coding.
That is a very hard topic because
security is not
high tolerance. You need to get it
right.
They tune a model that also
top the kind of benchmarking security
area.
So you can see all these examples is a
demonstration. They solve all the
companies are solving a very unique
problem in a special way and they have
been able to kind of
build their own challenges use their own
data by building on top of open model.
And those are the details. Obviously, as
you know, last year 2025 is the year of
coding. There
Fireworks we support all the coding
companies build on top of us. A lot of
success in
mid-train to post-train
very strong models.
As you can see those benchmarks are very
impressive and all these coding
companies are inspired to
get on par or beat Frontier Labs in the
coding space, but it's not coding. And
we start This is the year we start to
see very interesting development in all
different kind of co-work space from
general purpose co-work to specialized
domain specific co-work as I mentioned.
There are like legal,
finance, marketing, recruiting, sales,
customer support co-work space. They are
all starting to post train and own their
own intelligence power their product.
Here,
Jen Spark is one of the generic co-work
application and they build deep research
for professionals and slide generation.
As you see,
this they compete with a frontier model
and it's slightly better, but the cost
is significantly lower. So here we're
talking about five to 10
10 times cost reduction. So
as a startup, you can once you hit a
problem like if it you can scale quickly
into a doable business. And I guess
another health care example and health
care data or use case is not well
captured
in
in frontier labs model. So you can see
the quality of tuned model is
significantly better than the state of
art closed model.
Well, we have many different type of
developers build on top of fireworks to
do post training. So what what are these
types?
So we have seen many frontier agent
builder.
So frontier agent builder, they usually
build bespoke customized harness.
They don't use common harnesses. And to
co-optimize their harness with their
systems and all different kind of tools
they want to build on top of
often time requires post training. Uh
so, we have we have seen a lot of
repeatable success on doing that.
And
obviously, we have seen a lot of big
companies
put to post training because
interestingly, those incumbents have
huge amount of traffic.
For them to deploy AI features across
the board
to all the all their customers is a huge
cost burden. Huge cost burden. Uh so, if
we say we we need to be careful not
scale into bankruptcy, it's not just for
startups. Uh it's actually for
incumbents. They literally their CFO is
blocking their AI feature launch because
of the cost. And the post training is
way to remediate that.
Uh and we have seen a lot of specialized
model operator, uh those could be uh new
uh
new labs and a lot of kind of
cutting-edge model develops also um do
significant post training.
So, uh this is a quote we have heard
repeatedly
uh that after part of market fit, post
training becomes the vehicle to for
many, many companies to build
specialized intelligence. We do believe
we do see the trend that in the future
there could be millions of specialized
models. One application per use case. It
it it it will be millions. That's how
uh I believe in that.
Um And obviously, there are
um
we have seen kind of different way to
convert the data into into model. Um and
the the signal from the reward and how
to build those reward function
is very important um part of the story.
Um so, our observation is
no,
start early, start early, and the start
to do experimentation a lot more
iteratively
and they get hands-on experience and we
we have seen people on board so quickly.
There there's actually not the barrier
as many people feel because this this
whole reward engineering is very similar
to software engineering.
The the the logical reasoning and
mindset is very very similar. So
so yeah so
get your hands dirty and and start to
kind of test it out.
Cool.
I think
Yeah so those are the kind of key
takeaways
and uh
I think we [snorts]
we we want to acknowledge there are
different specialties or different
knowledges you have right now in this
space. For example, we work with the
companies like Cursor and Cognition.
They have deep researchers. They want to
control every single knob as much as
possible to get extreme results.
We give them the lowest API.
That means
we just
have RL rollout that they can directly
interact with and they fully control the
trainer.
That gives them the best result.
But many of the company
you have product engineers or machine
learning engineers.
You have good enough information and
knowledge about post training. You want
to have some control but you do not want
to have the lowest level control. So
that's where we have our training SDK
iterate very closely with all of you to
get to to get you going and we also
are happy to kind of deploy our own
researchers. That's where we will see,
um, help hands-on like training your
team to to get on board. So, those are
the kind of different way we see
different teams want to engage and and
get familiar with the whole entire
process.
Cool. Uh, I'm going to pause here and
see if you have any questions.
>> I'll jump um
People are most most successful if they
have data and reward signal already,
maybe from their product. What are the
what are the best examples of reward
signals you've seen with
very successfully trained very quickly?
>> So, um, usually the rewards you can
think about rewards actually is code.
You write rewards in code. And uh,
usually think about rubrics of rewards.
Um, and think about uh, you want to
grade the result in multiple dimensions.
Uh, [snorts] for example, if you are
building a recruiting
agent,
uh, and then you can think about, "Hey,
how do we evaluate candidate selection?"
Um, and a different company, I guarantee
you you have different criterias. Uh,
for example, we want to grade a
candidate
um, um, aptitude.
Are they really hungry? They do not take
no as an answer. They they will break
down walls. And that's one matrix. The
second is they're really fast in kind of
um,
building things and making progress. And
and so on. So, then you have a blend of
score to blend to merge this. So, so
think about reward in different
dimensions of rubrics. And then
different company have a different way
to blend those. And that's that's your
um, unique part and secret sauce.
Oh, sorry.
>> Uh, great talk, Lynn. Raj from Vercel.
I'm curious about I think you mentioned
what are the pitfalls of starting too
early, but based on what you're seeing
from Vercel, Vercel cognition, or other
companies, when do they actually start
thinking about post training? Is it when
they feel ready? Is it when the cost is
now too much? Because the benchmarks
where they're beating it, that's a great
signal that yes, you can uh
um
you know, beat the frontier models, but
is that the primary motivator? Like that
is that 2% bump worth the cost or
investment into this whole framework?
And and how do they think about ROI? You
know, close models are more expensive,
but then post training has some
intercept of, you know, money and
investment and maintenance going
forward. So, I'm curious of like, are
they approaching when cost becomes a
concern or is it when the usage spiking
up and they really have to now plan for
a year or two in advance?
>> Yeah, this is this is excellent
question.
Uh so, usually so, think about uh in the
AI building
product market fit and the scaling the
business as actually two phases.
And then in the SaaS time, it's one
concept. You hit a product market fit,
just scale.
I know you got to scale as much as
possible.
Um but now we see a bifurcation.
Product market fit doesn't really mean
you have a durable scalable business.
And typically, we see uh companies
deploy the strategy of focus on product
market fit first by building on top of
um you know, uh Frontier Labs model
because you don't need to worry about
anything. So, just kind of spend your
money and hit a product market fit.
Uh and the other very important thing is
only after you hit the product market
fit, the data you collect from product
surface area are really meaningful.
A- and you will get uh also the volume
of high quality data start to collect
from um product surface area. That
became the fuel of uh you start to own
your intelligence. So, we um we see like
that as kind of the very strong
indication
um because once you hit upon market fit,
you really think about start start to
scale a business, you think about two
things. One is continue to keep your
competitive edge.
Two is
um build a build durable business, so
your um your revenue and your costs are,
you know, in a healthy state.
Um so, then um
post-training becomes a very appealing
solution because post-training allow you
to
um
basic encode codify um your uh unique
taste into into a model that no one can
uh steal from. Because it's very easy to
clone and copy application as is, as I
as you all know, right? Uh from coding
agent it's very easy. Uh so, from
screenshot, boom, generate the same or
even better app. Uh so, that's a very
scary that application itself is kind of
the a moat is is being um reduced.
Um and you bake your data uh where
reflecting the product engagement and
all the deep knowledge into your model
is a way to preserve that. And second is
you can post-train a model that bring
down the cost five to time 10 times. And
and then that means you can support five
to 10 times much higher traffic with the
same budget. And the your unit of
economics of scaling is so much better
than you avoid scaling into bankruptcy.
So, we see kind of those as as two
compelling story behind the reason and
timing.
>> Thank you, Lin.
>> Thanks a lot.
>> [applause]
Ask follow-up questions or revisit key timestamps.
In this talk, Lynn, CEO and co-founder of Firework, explores the role of post-training in building durable, intelligent businesses. She argues that as application development becomes easier, companies need deeper moats—specifically by embedding unique customer judgment, taste, and domain-specific knowledge into their own models. Lynn outlines a progression from prompting and RAG to fine-tuning and reinforcement learning, emphasizing that this process mirrors human learning. She also addresses common pitfalls, such as data quality, evaluation challenges, and reward hacking, while highlighting successful real-world applications of post-training in sectors like coding, healthcare, and finance to reduce costs and maintain a competitive edge.
Videos recently processed by our community