Reasoning in the Wild - Wenting Zhao
750 segments
hi I'm witing j a PhD student from
Cornell University in this video I
wanted to share my research on reasoning
in the
while let's start by focusing on the
word reasoning what do I mean by
reasoning here is an example of
reasoning question what project my kilan
and witin collaborate on I searched this
question on Google but I couldn't really
find an
answer therefore in order to answer this
question will have to use existing
knowledge that is available and DOA
inference to drive new
knowledge here is how we can answer this
question we first apply a decomposition
strategy where we break down the
question into two parts what is g
interested in and what is one interested
in where this knowledge is pretty
available by just looking up our
personal
websites next we apply deductive
inference on identify reasoning topics
around of overlapping
interests so in this talk I'm interested
in the class of natural language
problems that requires using inference
or problem solving strategies to drive
new knowledge that was now part of the
training
data one example question that does not
fit into this category is who is the
44th President of the United States
which can be answered by just recalling
facts from the training
data recently we have seen that language
models become really impressive on a
wide range of benchmarks like they can
do competition math problems or pass law
school admission test but I want us to
forget about all these Benchmark results
and ask can language models help you to
do your reasoning tests the answer is as
a user the experience of using language
models like chat gbt has been quite a
mixture of successes and
failures let's look through a few
questions that user ask in the while in
this first question the user asked does
parsley sinking milk trt responded to it
by saying yes parsley generally sinks in
milk which is incorrect because in this
physical world if you put parsy on milk
you will
float and the reasoning chat provided
was the density and the texture
partially typically caus it to sink
rather than to float however the
language model didn't really think about
the actual density it just hallucinates
way to the
conclusion in the second question the
user asks your task is to create a web
application where a user uploads a video
and converts it to aski Art the user
should be able to choose character sets
and output
scaling to get a sense of whether
language models can do complex task like
creating web
application we create a benchmark for
similar path that's called commit zero
which challenges AI to generate Library
is from
scratch it turns out that the best C
language models we have today they only
achieve 5% pass rate on this task so it
is clear that a complex task like
creating libraries from scratch is still
far beyond what language models can
do now let's try to understand why this
user task are so hard first there is the
distributional shift from what data the
language models are train on and what
are the questions user ask at a while
for example language models are train on
data like Barack Obama is American
politician and lawyer who serve as the
44th President of the United States from
2009 to
2017 however in the while what user ask
about Barack Obama was is it fair to
call Barack Obama a fraud for failing to
address the issues he ran on in
2008 in order to answer this question
the model needs to First recall what
were the promises made by Barack Obama
when he was running for president and
what were the things he actually did
when he was the president and compare
these
two so in summary language models are
trained on data like Wikipedia documents
and books but they're used for all kinds
of questions with all kinds of
intents Beyond this distributional shift
supervision for this user questions is
also very hard to collect let me show
you how challenging it is to provide
notations for these user
questions in the first question the user
asked please write down the bowling
point formula and calculate the
theoretical bowling point of sugar and
salt where answering this question
requires knowing physics in the second
question the user ask is it safe to take
malonian AR from age s for three years
which answering require a medical
knowledge and finally we have this
example again where the user asked
language models to create a web
application from scratch which usually
takes a team of Engineers years or
months to bop so the resources required
for annotating this question are
insanely high so to Summarize each of
these question requires domain expertise
and domain experts are hard to find and
second even if you find these experts
annotating this complex task still take
a long
time so both of these reasons make
collecting human annotations not always
a viable
option seeing these challenges my
research goal is to develop methods that
use alternative supervision for real
word reasoning tasks let me give you an
overview of the
talk in the first part I'll discuss how
we can ground reasoning on natural
settings despite the fact that NLP is
datadriven field we have no data to tell
us what questions users are ask in the
language models to overcome this issue
we collect a data set called wchat which
is 1 million user chat gbt conversations
collected across the world we first
inspect the use case distribution of the
data and we use this data to improve the
reasoning capability of language
models in the second part we are going
to discuss in the absence of direct
human supervision how can we teach the
models to Reason by using alternative
supervision let's first establish what
we mean by a reasoning task given a
question X the language model needs to
produce the reasoning chain Z as well as
an answer y now the goal is we want the
model to learn to produce this ring
chain without human
annotations specifically the alternative
supervision we are going to take is to
use the latent structure of reasoning we
take a classic idea from machine
learning early 20s we build a latent
variable model through reasoning here
both questions answers are observed and
my Approach is to use these observations
to infer reasoning
chains in the third and final part we
are going to explore another type of
alternative supervision where we build
language agents that learn to Reason by
interacting with the
environments coming back to this web
application example it is really hard
for the language models to generate a
web application and it just
works instead it needs to write a code
for the application as the code code and
revise the code based on the feedback
from the compiler and we are going to do
something similar here so the goal is
that we want the model to learn to
predict Solutions without human
annotations the alternative supervision
model learn from is environmental
feedback where instead of telling the
model the correct solution it's going to
tell the model whether their generated
Solutions are correct or not with this
feedback we are going to do some
imitation learning
so to summarize this talk is about
reasoning in the wild is a challenging
problem because we often don't have
human
annotations in this talk we'll discuss
ways to build AI systems with
alternative
supervision great let's dive in in this
part we will cover how we can ground
reasoning in natural settings I'll cover
the paper B chat 1 million chatri
interactional locks in the
wild so we don't have any exist data
sources that tell us about what users
are asking the language models and
what's worse is that this data is how
privately at companies like open Ai and
Google and this really hurts the goal of
open
science PR work has explored if we can
cize user questions from language models
or collect user questions in a
crowdsourcing fashion but user questions
generated by these approaches either are
unnatural or a lack of diversity in
response we introduced a data set called
while chat to collect this data set we
first obtain explicit consent from users
we then provide the best language models
at the time to these users for free
acting as the proxy server we collect
their
conversations and by the best language
models I meant the newest open AI models
as a result we collect over 1 million
conversations from over 205,000 users
around the world please know that our
data is fully anomalous there is no way
for us to trace a conversation back to a
user and here is a snapshot of where we
are getting conversations
from to show you some general statistics
of the data set here is the language
distribution where
55% conversations are English with the
rest 45% being other languages such as
Chinese Russian Spanish and French I
wanted to know that this is pretty rare
in the existing data sources which are
dominated by English in total we
deducted 68
languages here is the use case
distribution 65% conversations are about
writing where users were asking for
writing assistance such as drafting or
editing emails 14% conversations being
decision analysis for example pros and
cons
analysis and the remaining categories
are 7% coding related questions 6%
information seeking questions and 6%
reasoning
questions to improve open source models
we distill the models on Wow chat
conversations there are two sources of
improvement the training data is more in
distribution and the responses come from
best
models we empirically evaluate how
useful it is to perform instruction
fine-tuning on while chat versus other
data sets we do evaluation on the Mt
bench Benchmark which is a set of user
return curies broken down into a few
categories in addition to Wild chat we
train on apaka which is a user Cur data
set synthesized from language models and
on Dolly which is a user cury data set
Crow Source from a few hundred
annotators with to a llama 2 based model
on each of these data sets and we
measure the performance with LM as a
judge where we ask GPT 4 to produce a
Liker
score here are the results training on
wild chat leads to significant
improvements in the domains of reasoning
math and
coding in addition to improving
reasoning while chat has been used to
study AI fairness researchers have
examined how chots treat users of
different genders when models can infer
gender from the usern
names for example when user ask for five
distinct ECE projects the answer
distributions defers significantly
depending on whether the username is
Ashley the name typically associated
with females and Anthony which is
generally considered a male
name while chat also highlighted some of
the real word challenges of AI safety
our findings indicate that 78% of
conversation
containing harmful content are related
to sexual information while 10% involve
violence to summarize the contribution
in this part we introduced a data set
called wchat and for the first time we
captured how language models are used in
the real world and training language
models on wild chat improve their
reasoning
capabilities bring this back to our main
theme we have deified what real world
user cues are like
now let's move on to the next part of
the talk collecting supervision for
rward task is hard and expensive so how
can we train models without human
annotations in this part we will cover
the paper help Union generate
explainable multihop reasoning without
rationale
supervision let's formally Define a
reasoning example given a question X the
language models needs to produce a
reasoning change Z as well as the answer
y
for example given the question does
parsly sink in milk the Lage models
needs to produce a reasoning chain Z
parsy has a density of 26 gram per cubic
centimeter when fresh milk has a density
of 1 G per cubic cm and ions think if
they are denser than the surrounding
material and so the answer is
no here is how we are going to train
models to reason with annotations on Z
because both questions and answers are
observed we start from the objective
where we minimize the netive log
likelihood of P of Y given X and we
introduce C through
marginalization so now we are switched
to this new
objective then we use a classic idea
from machine learning where we build a
latent varable model for reasoning the
generative process reeds in two
steps in the first first step we are
going to produce the reasoning chain Z
given question X and in the second step
we will produce the answer y by
conditioning on both question X and
reasoning
Z for now everything seems working like
a magic like how does summing over all
reasoning chains tell you about which
reasoning chain is correct let me give
you a bit of introduction for why
training on this objective
works you can think of this approach as
reinforcement learning where you take an
action and each action is associated
with the reward in our case an action is
to pick a reasoning Chain by looking at
the question and the reward you will
receive by taking this action is How
likely the correct answer is under the
reasoning chain you pick that is if the
reasoning chain makes the correct answer
more likely then you get a higher
reward as a result of this training
process the reasoning chains that make
the correct answer more likely will be
upweighted using this latent training
approach we have successfully remove the
need for supervision on
Z however this latent weable training
introduces a fundamental issue it is
computationally intractable to enumerate
all possible reasoning chains our
solution to this intractable problem is
to use the best of end training approach
instead of summing over all possible
reasoning chains we focus only on the
more promising ones we find this by
considering the reasoning chains with
high probability under question
X so formally instead of summing over
all possible chains we only sum over
attractable set s which is top K chains
from P of Z given X however there are
still issues remaining when we sample Z
from language models without any
constraints they can be any sequence of
of text as a result the reason chains we
sample are often factually
incorrect another issue is that a lot of
these free text reasoning chains are
useless and irrelevant it would be
really nice if we can just throw them
away so to ground reasoning chance on
the factual World Knowledge here we
redefine the space of C instead of
having Z to be any sequence of PR text
we make his sentences in Wikipedia
documents
again we want the space of Z it to be
structured so we don't want to just
select a set of sentences across random
documents so we add some hierarchical
structure in Z we break this down into
two steps we first identify the most
relevant Wikipedia
documents in the second step we only
search for sentences in the most
relevant documents identify in the first
step this way we don't just sample a set
of sentences across random documents
instead we are sampling sentences from
only the relevant
documents as a result this hierarchical
structure really helps in narrowing down
search space of
C so just to summarize the whole
approach we have built up to this point
first we have both X and Y observed and
we want to minimize the negative log
likelihood of P of Y given X
and we introduce Z through
marginalization because an unconstraint
space z will lead to many problems like
hallucinations we constrained it through
sentences in Wikipedia documents and we
wanted to leverage the natural
hierarchical structure of text where we
first identify relevant documents then
relevant
sentences and finally to overcome the
intractability issues in training we use
a b of end training approach where we
only s over the most promising documents
and most promising sentences and this
completes our
approach we name our approach hug which
stands for H unit and generate in plain
words this latent varable Training
Method treats reasoning chain as the
latent varable and it overcomes the
intractability issue using hierarchical
best of end
training we empirically evaluate hug
against the previous best on supervis
method we do evaluation on haa QA and
musical which consists of questions
requiring multiple steps of
reasoning we compare against the
previous best approach rack which have
also have access to Wikipedia documents
but can only think in one step unlike
hug which can think in multiple
steps and to make the comparison there
we reimplement rag so that both hug and
rag are using the same based model Bart
which was the best sequence to sequence
model at the time of 2022 and we
measured the reasoning performance using
sentence F1 between ground truth
sentences and predictive
sentences here is the result our
approach provides a promising way to
learn multihop reasoning in question
answering when supervision is not
available let me show you an example of
what reasoning chain our method is able
to recover the question is when copsy
was made Earl of North Andria he went to
reside in a town at the Confluence of
which two
rivers hug selected the following
sentences to be the reasoning chain
first by using copsy as an entity Bridge
hug identifies the sentence in return
William and copsy Earl of North Bria and
send him back to
York then in the second step
went from York to a new sentence York is
a historic wall City at the Confluence
of the rivers oen Falls in North
Yorkshire
England and finally by combining these
two pieces of information hug predicted
Al and false to be the answer and
therefore successfully solved this
reasoning
problem we have extended hug to other
reasoning problems such as adoptive
reasoning and reasoning about
ambiguity in both most cases using hug
sco approaches eliminates the need for
human supervision and demonstrate strong
empirical
performance to summarize this part we
first developed a principal latent
variable Training Method which remov the
need for human supervision and by
incorporating structure information we
achieve strong empirical
results and bring this back to our main
theme when human supervision is not
available in the wild we can train
reasoning models with structural
supervision where learning exploit the
latent structure of
reasoning now the second part of the
talk have discussed how to use language
models to solve reasoning problems in
reality language models alone are not
often enough I will share my work on how
we can build a gentic system that
interact with the environment to improve
their reasoning
capabilities in this part we'll cover
two papers commit the Library generation
from scratch and multi- turn code
generation through single step rewards
in this part we focus on the task of
code generation code generation have
some of the hardest reasoning problems
because these problems often require a
combination of different problem solving
strategies and a long sequence of
inference despite their complexity the
feedback is fully specified without any
ambiguity let's formally set up the
problem
here the variable X is a problem
description and a set of unit tests for
example the problem description is write
a function to re res reverse words in a
given string and corresponding unit test
are asserting after reversing Pi program
is program pi and empty string is still
empty string after reversing and the
goal is to predict why which is code
that fulfills the problem description by
passing all unit heads
We additionally have execution feedback
o where we observe the results of
executing unit test at the beginning the
observation should be no unit has pass
and the goal is to observe all unit has
passing here we Define language agents
language agents are systems that predict
an action by conditioning on state as
whereas is a trajectory that alternates
between action Y and observation o we is
a set of trainable
parameters in the case of code
generation the agent produces code y
given a state and sends the code to unit
has executor and get adds and
observation o there are two ways we
could use feedback observation o to
improve language agents the first way is
to use the feedback at test time and the
benefit of this is we don't even need to
update the
model the second way involves training
the models to use the feedback which
require updating model parameters but it
leaves more room for
improvement we will first discuss the
test time methods here are two ways we
collaborate test time feedback the first
way is to sample a ton of code
independently and then pass these
samples individually to the UN test
generator and pick the Yi that passes
most unit test
alternatively you could use execution
feedback in an iterative way first you
can sample the agent to produce an
ential solution based on the problem
description and execute the code then
sample the code based on the observation
from the previous
execution and apply this process
iteratively we could repeat this process
until Yi passes all unit test or until
the maximum number of turns allowed
we empirically evaluate these two
approaches on Commit Zero The Benchmark
we mentioned at the beginning that
challenges AI systems to generate
software packages from scratch we use
the best C language models sonant 3.5
and we use unit has pass rate to measure
the
performance the results are here from
this result we can conclude that by
using execution feedback we can improve
code generation without even updating
the
model test time feedback method are nice
but they are limited by its current
model capabilities so let's see what
more we can do when we get to train the
models test time methods have a sparsity
issue where you generate a bunch of
samples Yi but none of the Yi are
correct to overcome this sparcity issue
we replace the binary unit has execution
feedback with A continuous learn reward
model that Tak takes a problem
description and a generated solution and
produce a real value score which tell us
how close current solution is to the
correct
solution if a fully correct solution is
value one we can assign a value from 0o
to one to incorrect
Solutions and the solutions that have
the highest value is the solution that
is closest to the correct
solution here is an overview of our
training Loop first we sample and
trajectories these trajectories
alternate between sampling and test
execution then with simple random pair
of cod Solutions one passes all unit
test and the other doesn't pass
them and finally we train the reward
model on these pairs using a brary
loss and with this reward model we are
going to relabel the trajectory and
train on this relable
trajectory specifically we will have the
reward function to score every generated
solution and find a generated solution
that has the highest reward and denote
that with
Yar then for every trajectory we will
replace the solution at the last step in
this example Yi has a reward value of
05 and we are going to replace this YN
with Yar which has a higher were value
than5 after this relabeling operation we
train the agents on this trajectory so
that no matter what solution the agent
start with they can always end up with
the optimal
solution we empirically evaluate our
Training Method on two code generation
data sets mbpp and human eval we
compared to the iterative ttime feedback
approach we trained the open source
llama 2 3. 2 1 billion model and again
we measure the performance with unit has
pass rate here's the result our Training
Method consistently improved by using
feedback at test time
only in this part we have developed a
language agent that learns by
interacting with the environment and
training agents to use this feedback
leads to strong reasoning
performance now bring this back to our
main theme when human supervision is not
not available in the wild we can train
language agents using execution feedback
generated by external
tools great let me sign up the talk We
Begin by discussing two challenges in
reason real world reasoning where the
data language models are training on is
different from the user questions and
collecting human annotations for these
questions is both challenging and
expensive in the first part of the talk
we deify what user s in the while are
like in the l two part we explore two
types of alternative
supervision so where do we go from
here here are some interesting ideas for
building next generation of reasoning
models LM training consists of three
stages you first pre-train on internet
data for models to understand text then
you align them with user questions
finally you the post training aims to
further improve the reasoning capability
of language models using reinforcement
learning in post training a major
component is to perform search with
verifiers just like what we did in
training the language agent we sample
many solutions and we check the
solutions with the
verifier we and we train the models on
these correct Solutions judged by the
verifier in the first two stages because
we are training on existing data we are
also bounded by existing data however
post training relies on search and this
search can take us to
unknowns I think the next generation of
reasoning models will rely on better
post training
methods the first interesting post Trin
direction is the combination of language
models with with verification tools
right now the verifier focuses only on
the co-executors and language models
learn to orchestrate a diverse set of
verification tools like satisfiability
or mixed integer programming sers or can
we extend post training to domains that
are not strictly verifiable such as
creative
writing the second interesting post
training direction is the use of
synthetic data currently most post
training approaches tring on model
generated outputs and this can lead to
many issues like mode clap
so one idea is to develop principal
methods for learning with synthetic data
to deal with these issues another idea
is to develop scaling law to better
understand the compute and data tradeoff
when learning with synthetic
data finally I find the most exciting
reasoning problems come from the
frontier of scientific research where
data is extremely scarce and noisy
previously we buil AI system that
combine reasoning and learning to Aid
the discovery of the solar field
material now with the growing reasoning
capability of language models can we use
them for
automatically material or drug Discovery
is one interesting problem to
study with that I would like to conclude
my talk and thank my amazing
collaborators without whom this talk
wouldn't be
possible and thank you all for listening
Ask follow-up questions or revisit key timestamps.
This presentation explores the challenges of language model reasoning in real-world scenarios where human annotation is scarce or expensive. The speaker addresses this by introducing the 'WildChat' dataset to better understand user interactions, developing latent variable training methods ('HUG') to ground reasoning in factual sources without explicit supervision, and creating language agents that learn via environmental feedback in code generation tasks. The talk concludes with future directions for reasoning models, emphasizing the potential of post-training with diverse verification tools and principled synthetic data usage.
Videos recently processed by our community