RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor
704 segments
Um, Record, I think you guys grew from a
1 to a 2 billion dollar revenue run rate
in the last 4 months or so.
Um, so this company's off to the races
and I think you were just so front and
center to how companies are thinking
about uh post training their own models,
uh building their own intelligence. So,
thank you for joining us for this
conversation. Um, format-wise what we're
going to do is we've 15 minutes or so of
content from Brendan. He's going to talk
about uh RL environments in particular,
which I think is a, you know, new
frontier topic. It'll be fun to fun to
explore. And then we're going to leave
15 minutes or so at the end for Q&A
again. So, uh please keep please keep
questions back pocket. I will turn it
over to you, Brendan.
>> Sweet. So, I'll be talking about RL
environments. Starting out, I figured
it's helpful to give a little bit of the
background on the history of the data
market and how that history ties into
Record's origin story. Where things
really started in 2020 in the era of
crowdsourcing data for behavior cloning.
So, this was mainly supervised
fine-tuning data, inputs and outputs,
and RLHF data where you would have a
annotator select from a couple of model
responses which they preferred. And we
were able to make all this progress in
fine-tuning GPT-3, making progress
towards ChatGPT and GPT-4 in the
crowdsourcing era of agentic data. But,
what we saw changing in the market,
especially as we uh headed into 2024,
was this giant transition away from the
low-skilled crowdsourcing era of
behavior cloning data and moving towards
the agentic era of data. Of how do we
find the highest-skilled experts in the
world that can work collaboratively in
teams to build frontier evals and RL
environments for the next generation of
models models. All the software
engineers, lawyers, doctors, bankers, et
cetera that could measure the frontier
of intelligence and help to use that to
improve model capabilities. And so,
Mercor grew up with our first big
project being deep research. I guess the
first prominent RL agent
scaling up dramatically with all of the
frontier labs to become the primary
agentic data vendor to
all of the leading labs
and also all of the leading application
layer companies ranging from Harvey,
Cera, Cognition to Ramp. And what's been
really exciting over the last 12 months
especially is how RLVR within the
agentic data paradigm has evolved to
also include RL environments with these
rich apps and worlds that teach agents
how to use all of the tools on
our laptops that we use every day. So,
I'll be talking about that
and of course how this technology that
started in the frontier labs is now
getting disseminated to the application
layer and all of the products that all
of you are building
as you
work on your company. So, high level on
what an RL environment is is that it
includes three parts. The first part is
the worlds. So, this includes all of the
messages, slides, docs, sheets, etc.
that correspond to everything you would
have in a real project or company that
you're working on. The second part is
the apps which is high fidelity clones
of popular applications, Salesforce,
ServiceNow, Microsoft 365, etc. that
agents can interact with via MCP, CLI,
or Kua. And then the third part is the
tasks where we have prompts and
verifiers. Verifiers could be rubrics or
unit tests that can be used either for
eval or training. And the barrier for
frontier labs to automate everything
that you can do on your laptop using
Claude is how do they cover the full
distribution of all of the worlds, all
of the apps, and all of the tasks in the
economy. And so, there's been this
enormous scale out in order to do that.
Um, where humans have been really
central to how we build these
environments, obviously with models in
the loop meaningfully. And so I put a
graph here of the amount of expert hours
that um,
of throughput in from our talent network
over the last 24 months. Um,
and it's a a pretty crazy trajectory
with respect to um, 2.5 million hours
um, in Q2 alone with growth sort of
accelerating on the amount of expert
time uh, used to build out all of these
environments. The reason being of course
as I mentioned we need to scale out the
environment distribution across every
category in the economy. Many of you
might know GDP val where there's 205
domains in the Bureau of Labor
Statistics across all the different
jobs, but then you have to think through
how do we have all of the apps
corresponding to all of those jobs, all
the different scenarios, all the tasks.
Is this enormous build out. Only humans
can measure the frontier in most
domains, not every domain. There are
rare exceptions like math where you have
a really clean simulation environment
and so uh, the model's able to learn
from whether it got the right answer,
but in most domains like building a
slide deck uh, the model has an
incredibly hard time identifying
reliably where it made its own mistake.
It's as if you would be asking a human
to grade their own homework. And so
that's why it's really valuable to have
a human create a rubric similar to how a
professor would create a rubric to grade
an essay or a TA would grade that slide
deck. Similar to the way that a lot of
us learn it's in large part from the
feedback we got from those around us
rather than
uh, purely plugging things into a
calculator or clean simulation. Um,
and then building these verifiers is
hard cuz anytime you're building the
slide deck you need to understand the
full problem space of what are the 10
different slide decks that, you know,
could be a good path to go down? What
are the dozens of mistakes you could
possibly make? And how do you build a
comprehensive verifier that captures
this full solution area of what's
possible? And so, what I'll walk through
is a sample RL environment.
Excuse me, to also break this down for
all of you. Um and part of the reason
that this is so cool, which I'll get to
in a moment, is that we developed a lot
of this technology in collaboration with
the labs. These are, of course, ones
that we have open-sourced and published
to the world, but now that's all
starting to get disseminated to the
application layer companies that are
building and owning their own
intelligence. As they realize that the
three core pillars of their AI strategy
are their compute, their algorithms or
researchers, and the data sets they
build. And data's often the most
differentiating factor. And so, this is
one that we published, um
as a legal environment, where we have
lawyers from top law firms like Latham
and Watkins write out a scenario of a
real project that they worked on in
their big law job. And then they create
a full outline for a data room that
corresponds to all of the different uh
messages, emails, files, size of files.
I cut off the full data room cuz it's
it's very extensive. Um and of course,
there's a lot of model in the loop with
how they effectively populate this.
Similar to how a software engineer now
should not be coding by hand entirely
themselves, they should probably be
orchestrating agents in how to do this
very productively. Um
and then we render that data room into
the apps, the clones of Google Workspace
you can see in this scenario, and have
prompts uh to roll out model
trajectories against this. And so, in
this one, it's evaluating the maximum
total liability for Star Tanker uh
Tankers International Limited compared
to Cooper Jefferies Energy Corporation
under the Oil and Petroleum Act, uh
considering all of the context uh from
this real scenario in the data room. And
then as I mentioned, similar to how a
professor would create a rubric to grade
an essay, they have these key rubric
criteria that correspond to what are the
characteristics of a of a accurate model
response. And making sure that these
rubric criteria avoid reward hacking and
effectively align with the uh when you
roll out 100 trajectories, making sure
all of those scores are accurate is
incredibly technically challenging. And
so there's an enormous amount of
research, agenda quality control,
training on the data, etc. that goes
into how you solve that problem and then
ultimately produce these high-quality
verifiers and leaderboards um that give
you an aggregate model score across how
well all of the different models are um
doing on a particular domain. And as we
can see, one of the big changes over the
last few months is that GLM 52 and
Chimera K3 are on the leaderboard. And
so that is a huge opportunity for all of
you because that gives us the foundation
to actually achieve frontier
intelligence and all of the specific
applications uh and verticals that
you're focusing on um that's not too far
away. And so to give a little bit of
context on what that looks like, um I'll
share an example of post-training on
Apex Agents, which is the data set that
I uh or the sample I just showed before,
where we have 1,800 tasks in this
example. This was post-training run of
GLM 47, but we're redoing a lot of them
for Chimera K3, so we'll have updated
results for you all soon. Where you can
see the jumps just on 1,800 tasks with
about 500k in compute are pretty
dramatic. Um corporate law going from
4.7% to 26.6%,
um but notice that this is just Apex
Agents data set we gave it, and it
actually generalized incredibly well to
GDP valve and uh Apex V1, which doesn't
have these data rooms.
Even just seeing nominal gains and some
other benchmarks
as well. And so, we're doing a lot of
this work of working with customers like
Harvey, who I know will present on stuff
later to help build out the
environments corresponding to their
specific domain so that they can build
frontier intelligence within that. And I
think Andrew talked about how Cursor was
a great first example of how an
application layer company could build a
industry-leading model that, you know,
built an enormous amount of value for
their customers. And I believe that over
the next 12 months, there is going to be
dozens of examples just like that where
companies own their own intelligence and
that is the key source of the modes that
they're building. And actually, Josh and
I talked about this the other day as
well. Um
a couple of examples of ways to curate
high-quality data sets.
The general three that we see most that
I'm happy to
talk about and send people links to is
first by task is the most common.
Where people would say, "I really like
this data shape of environments in law
and we will pay $2,000 per task to scale
this up." And
as an example, certain frontier labs
might buy 50,000 tasks a month from us.
And so, it tends to be pretty dramatic
scale.
And these tasks would generally be very
complex. Some would even take humans up
to a month to complete that given task.
Sometimes it would take just a few
hours. And these would be sort of custom
per task pricing. Second is
off-the-shelf data
where we have we've invested hundreds of
millions of dollars in building our own
data sets that we sell to multiple
customers. All these new neo labs are
generally airing more towards
off-the-shelf data because it doesn't
make sense for 10 different labs to all
be building their own uh data sets. Um
there's a lot of value to building
something once that can uh then be
applied to everyone. And then the final
which um we see a little bit of, but is
uh less of our focus anymore, is just
providing um the experts so that
customers are able to
uh or organize the experts on their own
um
in just an hourly model. Um so we we do
a little bit of that when people like
that was how Harvey got started with us
hiring
uh some lawyers, um but it generally
moves towards more of these uh scaled
offerings of data over time. So, that's
a little bit of the background of how to
build RL environments, what they are,
and I'm really excited about all this
technology that we have uh
that has previously been limited to the
frontier labs all making its way to all
of you. And so, happy to answer any
questions um
about that.
Sweet.
Go ahead.
>> Hey, uh
um I'm Ali from Astro Guide. My question
is it's kind of open-ended, but simple.
How do you price data?
Like how do you value data?
>> So, there's a so many different ways. I
mean,
the most natural would be our customers
care about model improvement, right? And
so, our customers have a given goal of
they want to you know, be at the
frontier on a given leaderboard. And so,
we're able to work backwards from how
much is that worth to them and how much
should we charge uh per task, how many
tasks do we think would get them to that
goal. And so, when we think about uh a
company like Nvidia, they're probably
willing to pay, you know, a billion
dollars to have a frontier open-source
model. And so, there's a lot of
complexity of like how do we you know,
price all the different ingredients that
go into um making that happen.
The other way that we price
when it lens we look at it through is
also our cost structure to make them
where of course
when we have a task that takes 10 hours
of human time and
we're paying the human $150 an hour
there might be a $1500 cost basis and so
then it becomes a question of what
margin do we want to run on top of that
based on how differentiated and frontier
that specific task is.
But it's it's super wide range. Like we
have tasks that range from $50 to
$10,000. So.
>> How do you think about data quality? You
mentioned like you know utilize human
expert to label data and how do you how
do you compare the human label data and
you know the frontier lab you know
frontier
alliance the judgment data? How do you
compare them to from your opinion?
>> So first question was how do we think
about quality? Second one was sort of
how do we compare the like judgment of
preference labels to the auto graders?
>> Yeah.
>> So the core spot which is generally when
people say data quality they're
referring to two things. First is
realism and secondly is accuracy of
verifiers.
On realism it's just like they want to
automate everything in the economy that
corresponds to corporate law in this
case right? And so it's like how do we
make sure that this actually reflects
the real distribution of what we would
see in a real lawyer's environment. And
that's one of the reasons that experts
create outlines and help to guide all
the processes of the data curation.
Realism of the environment the apps the
tasks everything is incredibly important
and and also granularly understanding
the taxonomy that drives that realism
across the entire distribution that
you're looking for. The second part of
it relates to the other way people think
about quality, which is the accuracy of
the verifiers. Cuz the way that you
would train one of these models is you
might roll out 100 trajectories of Kimi
K3 and then use this rubric to score all
of those trajectories. And as you can
imagine, there's like so many different
paths that a model can go down. And so
you want to make sure that the way this
rubric is doing the scoring is the same
as if we were to just have human stack
rank those 100 trajectories. And so what
we do for that is a process called
trajectory analysis, where we roll out
10 trajectories of the model that we're
focused on improving,
um, and then score all of those, and
have some combination of agentic quality
control systems and some human review go
through to make sure that, um, all of
the scores align with, uh, the goals.
Um,
and sometimes you can also use, uh,
human feedback evals or preference
labels to as a eval for your auto
grader, um, is the other way related to
that that you're able to solve for it.
Go ahead.
>> Um, how much do you think like synthetic
data generation plays into all of this,
especially like creating these large
data rooms?
>> So the fascinating thing is I think that
there's been a lot of misinterpretation
of what people mean when they say
synthetic data because, like, RLVR is a
bet on synthetic data. It's basically
let's roll out a bunch of synthetic
model trajectories rather than having
the humans write the SFT, and then let's
score all of them, and let the models
learn from all of these like, uh,
synthetic model trajectories. So I think
that's the first way that synthetics get
used. The second way is that models play
giant role in the way that we populate
environments and create tasks in the
same way that, um, a lawyer that is
writing a legal memo should definitely
be using
Claude or ChatGPT to do that. The
experts that are building out these data
rooms should definitely be using Claude,
ChatGPT, or whatever model
to help them do that. And there's a lot
of ways that the model can make them
more efficient. But the reason that
humans are still an essential component
of the process that's incredibly
differentiated is that you need humans
almost definitionally to measure what is
beyond the frontier of the model
capabilities.
Like the models you you can't just tell
the model like come up with the legal
environment and then like tell me which
of your legal memos are like good and
bad. It's super noisy and there's not
like clear signals associated with that.
You need
something that has capabilities beyond
the frontier of that model to do so
reliably.
>> Thank you for doing this. Um my question
is RL environments seem like they're all
the rage now and maybe have been for
about a year. I had I hadn't really been
hearing about them prior to that and it
was all
human expert labeling. And so
kind of why why why is it all about RL
environments now? Is that the most
relevant thing for application companies
to be thinking about? And is there any
something after RL environments?
>> Um so I'll I'll start with why it's
become the rage and maybe some of the
differences also between the deep
research paradigm and the like envi- the
sort of environments paradigm as we saw
it in 2025. And then I'll talk about
looking forward what we see evolving in
the data landscape. Specifically, I
think that the reason the deep research
environments were the first was because
like deep research had tool use with
search. So search was the tool in the
environment that the model
would work in. But the experts would not
necessarily be populating apps. So it
was sort of a lighter version of an RL
environment where they would just like
create rubrics corresponding to this.
And again, like I only talk about this
stuff cuz it's a couple of years old at
this point, so it's no longer
uh super confidential. And then for the
like trend of apps in 2025, I think that
really became giant because people
realized that the primary bottleneck to
making the models useful was how they
started to use both all the context in
the code base and all of the tools on
our laptops, right? And so if we want
this in the user distribution of usage,
then we need to get it in the data
distribution that the models are
learning from.
Uh and so there's going to continue be
to be this giant scale up of diversity
across all three of these categories on
a going forward basis, but there's going
to be some changes
um to maybe name two of those changes
that we're thinking about the most. The
first one is ultra long horizon. Like
right now agents mostly aren't trained
to do things that are over 10 hours, and
we need to start building tasks for
things that might take a human 100 hours
or even 1,000 hours to do. And so that's
going to be a giant shift. And then the
other large shift that we're seeing is
introducing virtual co-workers, which
corresponds to that. Um like one my
favorite questions to ask people when
they're thinking about their data
distribution is what percentage of tasks
that they do in their job require
interacting with other people.
Uh and most people would say like 60% or
70%. Some people say a lot more, some
people say a little bit less. Um but
then if you map that on to what
percentage of evals measure how well the
models can interact with other people,
it's like 1%, maybe Tau bench has a
little bit of this.
Um
and so there's this giant realism gap
associated with how you actually measure
how well agents engage in social
interaction um throughout uh all of the
different people and other agents that
they need to work with uh in their jobs.
Good.
>> You talked about rubric generation,
which is like is it like a bespoke
access of verify that you put task? And
from what I understood, that's like
bottleneck by experts. Have you found
any success of like being able to scale
that up with your models giving you like
some heuristic or even like post any
models for
>> We have found that you can make it a lot
more efficient if you have a like AI
copilot that's able to work with the
expert in creating the task and the
verifier.
Um so, the expert can talk to the
trajectory and understand exactly what's
happening and where it's going wrong.
The challenge is just that if you're
trying to improve FABLE,
FABLE cannot reliably write out the like
rubric criteria for where it's making
mistakes. It might get like half of them
right and half of them wrong, and that
amount of noise is unworkable from a
training standpoint. Um and so, that's
the reason that the the process that
requires humans the most is the task
creation. Uh like a lot of the
environments, we can get
like use a lot of synthetic. Uh
it's helpful to have humans write the
outlines they're familiar with the
environment and and uh
and they're grounded in reality, a
realistic distribution, but with the
task that's those tend to really require
humans. Um
with with rare exceptions, um
in code or if you're sort of distilling
from like if you have a model that's
worse than Kimikaze 3, then it can
definitely learn from tasks Kimikaze 3
is creating. So, there are there are
some exceptions if you're doing it that
way.
>> Hey, this is uh Nikhil from uh Cribl.
So, when we think about RL environments
for certain provable domains like cyber
defense or incident response, where
there's the model is or the agent is
trying to find a flaw in an existing
system,
Do you use uh
humans for just authoring that
environment or setting it up or do you
also use that for grading? Is there Is
there a way to scale that up?
>> I I actually think cyber is one where
you don't necessarily always need humans
for the verifiers cuz you can have an
attacker and a defender agent.
Um
and uh I think you're right in saying
that for cyber you can have humans more
so or or architect what is a realistic
like environment um and sort of set up
the environment um cuz you do need a lot
of diversity.
Uh but then it's less human intensive
with respect to uh building verifiers.
>> Thanks.
>> How can you tell when you're limited by
the base model?
>> What do you mean by that?
>> Well, ostensibly you're using the same
data set for all these models here and
they somewhat land around the same final
performance on this list here. But maybe
if you try to smaller model, which is
maybe a good place to start, it would
land much lower. Um what is the cause
there? Is it just parameter count or
>> So
the parameter count will definitely play
a role in
how effectively the model does insofar
as how trainable it is. I think that um
the main thing to look at is generally
the gap between the like
pass at 16 and the pass at one. If you
have a model where you roll out 16
trajectories and it gets all of them
totally wrong, the it's sort of like
hopeless that the model is going to
learn from that um for the most part.
Maybe you roll out another 100
trajectories and it gets one of them
right. Um
versus if you have um
the ideal case is that you have
um
pass at one it fails, but then pass at
16 when you roll out 16 trajectories it
gets it right once or twice, and then
the model is able to learn very
effectively from that. So, that's
generally the heuristic we would use for
how strong the base model needs to be to
effectively learn from a given data set.
Cool. Maybe
Oh,
final question really quick.
>> What advice do you have for
companies as they partner with you
on data that they should rely on
themselves as
part of their RL post-training versus
relying
What's the complementary
>> Well, I think this is why a giant
portion of our business is custom data,
where it's like we have teams that are
siloed and fully exclusive to
critical customers to make sure that we
build out the best data sets in the
world that they own.
And that allows them to maintain their
competitive advantage associated with
this, while also benefiting from all of
the infrastructure that we've built. And
I think there are some companies that
try to like build out all of the talent
network and infrastructure in-house, but
I think if you look at what the frontier
labs do and the best models do, it's a
pretty good indication that there's so
many economies of scale from working
with a partner that has all of these
economies of scale, the platform, the
talent network, etc. So, anyways, thanks
for all having me.
>> [applause]
Ask follow-up questions or revisit key timestamps.
Brendan, from Mercor, discusses the evolution of AI training data from low-skilled crowdsourcing to the 'agentic era.' He explains that modern AI development relies heavily on building Reinforcement Learning (RL) environments that simulate real-world tasks, apps, and workflows. These environments, built by human experts, allow AI models to learn from interaction and feedback, which is crucial for achieving high-level performance in professional domains like law and research. He also covers the methodology for pricing data, ensuring quality through verifiers, and the future of long-horizon AI agents.
Videos recently processed by our community