From ML Engineering to AI Engineering
1340 segments
great
great
so we are ready to go so hello everyone
and uh welcome to today's uh ACM Tech
talk so this uh webcast is part of acm's
lifelong commitment um uh for Learning
and professional development um serving
This Global membership for computing
professionals and students um I will be
today your moderator my name is
Alejandra sedo uh I am currently member
at large uh for acm's uh governing board
as well as uh director of engineering
science and product at zando uh I am
also uh scientific adviser at The
Institute for ethical AI uh where I've
LED contributions to EU policy including
things like AI act data act etc etc so
for those of you who may be a little bit
unfamiliar with the ACM um or what it
has to offer here is uh some information
so ACM offers educational and
professional development resources that
bolster skills and enhance career
opportunities uh you can see some of the
highlights on your screen so these are
basically some of the uh key areas uh
that ACM is known for uh some of the
things to um highlight that are worth
noting is that ACM provides access to
what is referred to as the ACM Digital
Library this is the world's uh most
comprehensive database for computing
literature um ACM also provides leading
Publications and Global conferences that
draw top experts from a broad spectrum
of computing topics um they also provide
support for Education and Research uh
including curriculum development teacher
trading and uh perhaps also you may have
heard about the ACM uh touring and ACM
pricing Computing Awards um as well as
the ACM code of ethics uh which is a
collection of principles and guidelines
designed to help Computing professionals
make make ethically responsible
decisions so we have also published just
recently uh things like the um
generative AI principles which are of
course using this as the backbone as
well as the principles for algorithmic
responsibility uh which also served as
uh the backbone for several of other of
the ACM uh policy products uh before we
get started um let's mention a couple of
housekeeping rules rules um if you have
any questions at any time please uh type
them in the zoom's uh Q&A button I know
that you can also use the chat but the
Q&A button will allow us to collect it
more systematically uh we will then be
able to um yeah raise them as as the as
the um uh talk finishes the session is
being broadcasted uh uh and recorded uh
and it will be archived so you'll
receive an email notification when it's
made available and check learning.
acm.org for updates on upcoming
webcasts so today we actually have a
very very exciting topic uh with uh also
a very exciting uh presenter uh the
presentation is titled from ml
engineering to AI engineering and we
will be having uh chip uh who uh Works
to accelerate data analytics on gpus at
Voltron data as their VP of Open Source
uh and AI uh previously she was uh at
snorkel Ai and Nvidia and founded an AI
infrastructure startup uh that was
acquired uh she has also taught machine
Learning Systems design at Stanford and
she's the author of the book designing
machine Learning Systems which is an
Amazon uh bestseller in AI a lot of the
insights today are are going to be uh
providing a lot of depth on this so with
that uh and without further Ado um uh ch
please uh take it
away thank you so much aleandro for the
very intro and hi everyone I'm looking
at the chat right now and this is
incredible uh diversity in uh the
audience so feel free to like ask
questions when I'm speaking and I know
is that sometime going to get excited I
speak pretty fast so feel free to tell
me just slow down as well um so yes
so wait oh yeah my somehow my video was
not on I hope you can see me now hey hi
again um so yeah so today I'm talking
about from ma shielding engineer to AI
engineering as alandro mentions uh I
wrote um I'm the author of Designing
Machining systems and I'm working on a
new book that is similar like just built
on top of Designing Machining systems
but focused more on Foundation model so
we cover everything from like prom
engineering evaluations uh F tuning um
context r for actions agents and all the
things that you might have questions
about so in the process of working on
this book I do a lot of research so I
try to look at every
single AI tool out there and it's a very
time consuming things because there are
so many of them so to make my life
easier I only focus on repos at least
500 Stars so you can actually see them
on my blog uh on my website
h.com Lama poish so I think it's I find
pretty useful to like uh so like if I
wondering like who is doing inference
optimization so I can look at like I can
filter them out by the category of like
inference optimizations I want to see
who is doing Asian who is doing AI who
is doing cing um who is doing Serving so
it's been very helpful and you can see
like the number of changes like because
I notice a pattern just like I got a CO
the hype curve so you will see a lot
like sometime a repo come out and you
see a lot of excitement and a lot of
stars but then nobody ever uses this
again and it was just like flatten out
so I think it's if want to track and not
just a number of stares but like how
fast this is growing um okay so today we
talk about a engineering so engineering
is the process of building applications
with Foundation model so I use the term
Foundation models because it's a super
set of large language model um so I I do
think it's like we we seeing the
conversion of a lot of different
modalities when I started out in AI in
2014 we saw that like NLP natural
language processing was his own field uh
computer vision was his own own field
like reinforcement learning was its own
field but do see that like things are
like coming together so we see this like
a lot of models today like
gbd4 or Gemini they own incorporate not
just text but also images and we see
reinforcement learning being used train
model is like reinforcement learning
from like human preference so it's been
very interesting to see that um and
Foundation I think the term was coed by
Stanford because it also denote uh the
ideas that the concept that Foundation
model just ler foundation and it there a
need to be adapted to specific
needs um so there are a lot of changes
in um from like machine learning
engineer to engineering so one thing is
that models are getting really big big
and very few people can afford to train
them so we see they are like model
providers model developers who develop
models and make them available to use as
a service so this service you can access
through API or you can download them
like as an open weight model and you can
host them to S so this means that like
anyone even those with no AI uh
knowledge can now leverage a bu
applications so before right like always
is a company who could afford to build a
applications as a a big a company to
have data who can train their own model
and then these companies and then these
AI products are incorporated into their
main product but now like you can just
build it but then now you need to figure
out a way how to serve those
applications so one thing is very
exciting that we see a lot of interface
to serve the AI applications such at
like plugin so it can incorporate AI
into the browsers Google Docs or like um
vs code you can also see uh I'm very
excited about like getting like
embodiment of AI so for example we can
have like audio or like 3D character I
think like the ultimate goal is that we
have like AI embody in like physical
robots that can help us do things rather
well um so yeah so I think that AI has
become a common component in a software
engineering the same way the JavaScript
and databases are so uh I noticed like
engineering is coming like closer to F
speack than like traditional data
scientist so another changes uh is that
we are moving away from like uh we are
moving from close ended evaluation to
open-ended it's not moving away it's mon
need combinations of of of a lot of ways
to evaluate so Foundation models are can
gener open-ended responses and this
makes them really really powerful
because now with open-ended responses
you can adapt them to like different use
cases oh thank you yes I'm G to try to
speak slower um yeah
so yeah open-ended responses make um
being being open-ended makes partion
models useful for lot of use cases
because now can ask like say gbd4 should
do translation should write essay should
write code uh should do math however the
open-ended nature of models also make it
really really hard to evaluate so before
if you have say a classification task
right so you have the model who
maybe like one of those three categories
and say if you expect category like a
positive and the respon is not positive
then it knows that the respon is
incorrect but now it's very hard because
um there are many different ways to
express the same idea and even if some
model out put something that's like not
the same as expected sentence expected
text it's really hard to tell uh whether
is um is incorrect and also like as
models get more intelligent it also
becomes harder to evaluate them so
imagine like anyone can tell if the
solution to a first grade math problem
is wrong but it's very few people can
tell whether the solution to like a PhD
level math questions can um is wrong so
now that like less obvious mistake we
can detect from Ai and I so the
incorrectness or like inaccuracy of AI
now you have to do fact check in one
critical thinking and make it a lot
harder to evaluate
uh so we are seeing a lot of new
evaluation approaches for the
applications so one thing uh one
approach is comparative evaluations so
basically you you see um with this
approach you see two different responses
side by side and users pick which one is
better so this different from AB testing
because in AB testing you only see one
response uh at a time uh another
approach I that used to cause a lot of
controver icies but I do see that is
getting more adoptions which is like
using AI models as a judge so you can
ask um sometime you think it's like
using um evaluating whether an AI model
is good by like asking his friend so you
have a model to generate uh responses
and another another model or is the same
model to evaluate whether that respon is
good or not so there are many different
ways uh ai ai judges can be formulated
so one way is just like like a scoring
like a reward model uh so scoring is you
ask a model to Output like maybe a score
from one to five on the uh respond with
is good a reward model um or like some
others other models can also do
comparative evaluation for you so you
can ask a model to like pick which of
the two responses that humans would
prefer so I'm very excited about this
category of AI judges because if we
could uh find a reliable way to like
predict what humans prefer that could be
the goal right we wouldn't need a lot of
like humans annotation to create uh data
to train the model because like human
preference is is what like a lot of
people want for their models because for
this model to align with uh what humans
want um yeah and like other AI as a
judge is like you just ask a model like
hey does this s correct I given the
answer or like is this like relevant to
the questions but it's it's exciting but
it's like very hard to use because you
need a lot of rigor in making those
matrics standardized um that uh we are
doing okay I'm still speaking pretty
fast and I don't see
um I don't see the
audience it's VI out approach evaluation
we won't get back to this later maybe in
the Q&A I love that um so I wish had a
conversation yesterday with a few
friends and some the smart artist people
I know I was asking them like how do you
evaluate the model and they were like oh
we have about like five pet respond like
five queries five prompts that we use to
test and was like wow um things it's
like Vibe check is very important but
it's not sufficient so you do need to
like interact with the models every day
uh like to see the to menu see that to
like for for um because over time you
might have better idea like as you
direct the model you have a better idea
of whether what what consider a good
response one is the hardest challenge of
evaluations it's like creating the guid
like um so I so my friend he used to
like um like create guy life annotators
but then you you need to like it's
really hard to Define like uh some uh
you sometime need like to train people
to do so um so so yeah so like one thing
is you have to create very clear
guidelines to like what is good and bad
and only by interacting with the models
that you get have a better understanding
of like is response really good or not
because sometime a response can be good
but not not sympathetic so LinkedIn has
an interesting case study so they found
out that like um for they call is like
job assessment chatbot so when a
candidate see a job uh posting uh it can
ask AI like hey am I going to be a good
candidate for the job posting so the AI
can respond like no you're not right so
even though it's correct it's not very
sympathetic so they need to stay the
response the guy what could be a good
response so the respon is like maybe
like more encouraging like uh today
you're not but if you go to this skill
ABC then it would be a better fit and
here are the resources you can use to
like go to to make yourself um suitable
for this
role so I do think that evaluation is
the biggest challenge for JF AI
applications so one uh one col of
evaluation is like if we don't have a
way to evaluate whether model responses
are good or not just like too much like
it can make AI application too risky for
a lot of applications so today we see a
lot of Enterprises using um deploy AI
applications not because these
applications are potentially the most
impactful but they are the easiest one
to evaluate so for example a lot of
people want to do say um recommended
system because for recommended systems
you can evaluate the impact very clearly
because it will increase something like
click through rate or purchase through
rate or a lot people do coding because
with coding you can evaluate model using
a functional correctness so is a Model
gener A piece of code you can you can
execute the code and see if it runs and
whether it produce expected results so
because of that um if we can fight a
which to reliably evaluate responses it
believe that we can unblock a lot of
exciting applications
um okay so another things is like um I
also see more of the transitioning from
like future engineering to context
contractions so with classical machine
learning you need to like engineer
features like what could be uh so so
that for age of the query it would need
like creative features is important to
uh that's you think would be useful for
the model to make predictions so with J
AI model Foundation model was they use
prompt so for queries you need like uh
you need to inject useful information to
the prompts uh to the context so that
the model can respond uh to the query so
um Hallucination is an interesting uh is
an interesting uh term because uh
hallucination make it sound like a bad
thing but it's not always necessarily a
bad thing like you do one model like
sometime like be creative so it's
actually very useful for cre creative
applications um but sometime like for a
lot of use cases when you demand a
strict um uh like so some case when like
curacy is very very important like legal
or medical you you don't want models to
like make up things so we do see that
like many Studies have showed that like
models are more likely to hallucinate
when it when they don't have the right
information so uh so models have this I
think it's like a bias with actions so
even when they don't know something or
they unsure they just somehow like have
to respond and it had to like make up
something like if you if you don't know
the answer um I'm not quite sure like
how to correct that so I think part of
the training because like in the
training process right uh for on the
training us have like prompt response
every prompt isn't so there had to be a
response so I do think it's like to
correct that bias is going to be quite
challenging but anyway in the meantime
when you work with this model we have
seen that um models can perform much
better and hallucinate a lot less when
they apprach provided with the necessary
information to answer the question um so
like how do we enhance context like how
do we provide models with the correct
like with a relevant information uh
necessary for it to like respond so we
can do it to way like once it like with
retrieval so you provide access to your
databases quick documents images table
data so that they the model can retrieve
the relevant information from the
databases or you can also give model
access to tools so that it can gather
information sself so some tool like web
search is extremely useful so when um
the pattern of giving model access to
tools it also known as a gentic pattern
and we can go into more and chooses in
the Q&A if you're interested I think
it's very exciting but it can be a
little bit overhyped because you're
going to think of like um yeah we can go
more into this later so here's some like
visualizations of what enhancement uh
cont enhancement would retrieval look
like so it gives a model access to
external data databases so which be like
a collection of documents or images and
you have a retrieval system so like
given a query it will retrieve the
relevant information from the from the
databases and then the context is Joy
with a query so that the model can
generate a respon and some back to
users um so R for is old science so it's
not new I think I believe that's the
first digital
retriever systems would describe back in
the 1920s so it's like a century ago and
retrieval has been a backbone of a
search and recommended systems um so
it's it's not new so they have many
technology algorithms developed for
retrieval search so in general there are
two main approaches to a retrieval one a
turnbas so first of all like keyword
search like if you have a query like um
about say Transformer architecture it
can look into the database and search
for all the document that contain
Transformer architecture uh so there are
lot of term based Solutions um that
available like bm25 and elastic search
none of them is new but they have very
very powerful and like a lot easier to
stud it than the other approach which is
embedding based um retrial or like some
people call it a lot people call it like
uh Vector search so Vector search is
more complex um is like take longer um
is is more comp comp computationally
expensive but it can potentially provide
like much stronger performance because
you can have a lot of room for fun um
but at the same time I do things it's
like um if you're just studying out I do
things that like um I do think that a
termbase solution like bm25 is a very
robust um approach and is a very very
strong basine to beat so um when you
reach retri data right you don't just
retrieve structur data you also want to
retrieve like tabular data as well so
maybe you have like a c database with a
lot of SQ tables and I have an example
here so retrieving table data is very
different from retrieving like
unstructured data because now to aeve
retrieve data from from this a SQL table
you need to like write SQL query so the
first step would be like you have to
convert from like query like natural
language into the SQL query uh the
flavors that your database can
understand and then after that you
execute the SQL um the SQL query so this
like a bit trickier um and it can be
quite dangerous so say like what if like
people want to like execute a dangerous
SLE query like let's drop the table or
something like that so you would need to
like put in gut rail like like what kind
of queries that you can uh the model can
execute um so so yeah so in this case um
Al Texas pretty tricky because I choose
Jared a SLE query you need not just a
query but also need to understand the
table schemas so I say if you have a lot
of tables and the table schemas is too
long to fit into context length it could
be uh it could be hard uh so so and also
like like pick like what tables are
relevant so some people are trying to
build uh embeddings for tables based on
the table description and table schemas
so that for given query it can
automatically retrieve all these
relevant tables and then based on the
schemas of this relevant table jate a SQ
query for
those um yeah so so here example of like
H text queries can convert into s query
and it would need uh to know that you
are using the table
sales um so yeah you can also use retri
tool so here's what it looks like so
basically you need to give models uh you
need to define a set of that you that's
a model G take so action could be like
using a search API using a news API
using weather I think um the most
exciting tools that uh people wanted for
for AI like especially on CH stuff with
like web browsing because now with web
browsings you can allow model to stay up
to date so with our web browsing right
like um informations can be like um
models can get outdated so you ask a
model about something that happened
today and the information is not in the
database or it was not in the model
training data the model won't be able to
respond but with like web search the
model can stay a lot more relevant um I
think my website is like very
exciting uh okay so so I grouped on of
this um retrieval from databases ands
into what they call like people call
like context constructions um yeah so
like the go to enhance context so that
the US so that the models can answer
questions better and also reduce
hallucinations okay so so another change
um I know that we have been going
through like a lot of content I hope
that it's uh it's fine it's okay but now
feel free to ask questions we don't have
time to go over them um so the model's
going to get a Bigg like become bigger
so I remember in the early day uh we
talk like oh my God one billion um
parameters with like so much just crazy
like I remember when the first GPT model
came out so it was like uh 1 million
parameters and was like oh my God that
is so big and now we're like hundred of
billions and I think people are aiming
for like a trillion parameter models so
the the number of parameters uh is not
the only things that matter right it
also depends depends on how sparse the
model is because like the more SP model
is uh So the faster so the less
computational computation uh
computational resources are required to
run the model but anyway in general
bigger models require more uh more
computational resources to run so they
are more expensive to run they also have
higher higher latency and also it's just
like if you want to host those big
models it's require more expertise
because like it just don't simply put on
the machine you need to work on like
distributor system how to load dat this
you you know in um in um how how to sh
data so it's a lot harder to operate
larger model and then smaller model uh
so that also making that like um a lot
of Technologies to make this model run
faster and makes them smaller very
exciting so so one thing uh that I'm
very excited about is like inference
optimizations so inference optimization
is not new um so when I'm when I'm like
just uh so actually I used to take a
course of Stanford about um five years
ago not four years ago wow 10 fles um
but um I was like covering inference
optimizations and then I revisit the
topic
again sorry people still cannot hear me
um okay I'm going to try to speak
slower um yeah so so when I revisit the
topic of inference optimizations
recently and I realized that a lot of of
core techniques for in inference
optimizations still remain the same um
so so like the qu thing for example like
um quantizations quantization extremely
exciting it works pretty well like for
wide range of models and applications um
it's is not new right it use the
precisions from like um float uh 32 uh
to like float 16 and now8 bit and now
people even going for like four bit um
so another technique is um is low rank
low rank factorizations which is
basically like the core of like on the
Laura fing technique and also like they
like um sparcity with pruning and now
there new mixture of expert and also
with distillations where you trained a
smaller model to uh to imitate the
behavior of the larger model and of
course you are things like cash uh
parameter efficient F tuning so all of
that space is very
exciting okay so that is pretty much for
my presentations just for one last thing
I want you to cover uh so like see like
think of like for there are so many
adaptation techniques out there like to
your adap model but I would generally
put them into like three category like
one is like just prom engineering uh
like so you try to write instruct um
write clear instructions they compose
like a big task into like smaller
subtabs add think a lot of examples and
then the next one is like context
optimization so here I put rack but they
can be on so like action API as well so
the general retrieval plus actions and
the other category is like f tuning um
so fing I would say is like it's never
the first answer so I would say over
time as you progress through the
applications I think it's very important
to like get the most manage out of PR
engineering with like use very use good
examples with very clear guidelight for
like the model like good good example of
like how the model should act for
example if you want model to say give a
score from one to five Define very
clearly what is what does one mean what
what is an example of in one look like
and why it get it get a score of one and
the same thing one two three four five
um and after that you can try with like
simple like retrieval for example like
BM uh like keyword search or bm25 and
after that you can choose whether go to
more complex retrieval uh like vector
search because it would require you to
like use an embedding model and use a
vector database like sorry Vector search
which can be used as part as part of a
vector database so so yeah so like
because it's add more component to the
system it it increase the complexity or
it can also go as a f tuning rout uh and
then you can also like Joy together like
using r with a f tune model or using R
to like create um training examples like
uh to to fune the model so yeah so I
think that what what it looks like um um
so in general I would categorize so uh I
would decide what technique you use also
depends on what kind Behavior you want
the application to have or like what
kind of issues you want to correct for
the applications so I would in general
think of them as like uh failure modes
could either be information based so
like when the response are incorrect uh
or just outdated or another what I call
is a behavior based failure mode so when
the responses are like factually correct
but it's like either irrelevant uh or it
just doesn't follow the expected format
so for the information based um arrrows
I find the rack like context
optimizations helps really well but for
Behavior based uh like getting the model
should like follow the format then you
can try that with like um prompting
adding more example using vators um um
using some f technique like um
constraint samplings or using like
bigger models because as models become
more powerful they also getting better
at like following instructions and Jing
the correct format or you can fight tune
so F tun I think of like as a last L of
fans uh because fight tuning is exciting
but it also require both a higher a
front cost because you need to curate
data for f tuning and it also require
continual maintenance because you have a
model and now what right you have to
update it regularly and like what if the
newer model base model comes out do you
want to leverage a new base model which
mean you have to like maybe fight T
again so it's a lot of like cost
associated with f tuning that you really
should consider before committing into
it okay so um um so this is like a
little un ready to talk it's about like
what our work is doing so we build GPU
native CER engines um so the idea is
that gpus now are everywhere so I think
one one consequence of AI is that like
now a lot of companies have gpus and a
lot of people are using gpus for
inference and training but I do think a
big category is like data processing um
data processing um like ETL I do think
this like can be very suitable for gpus
because uh let's say if you want a model
if if you want to process a billion rows
of data you can just Farm z m to like
thousand of fours of GPU so you can
speed things up and we can also like
just make things like uh a lot cheaper
um we we also also contribute to open
source project like IIs and apach arrow
um yeah and if you want to find me you
can find me on my blog LinkedIn Discord
uh so yeah thank you so much
everyone great great let's please all
get our virtual hands together to give a
Applause to chip uh this was a fantastic
talk also big props uh chip for also
keeping an eye on the chat uh very you
know active uh uh reactions on on some
of the comments and even the questions
so we have uh lot of really great
discussion and much more to come uh we
have been able to group some of the
questions into different topics so we
will dive into what seems to be um areas
of uh hallucinations rag prompt
engineering AI judge evaluation uh gen
limitations and a couple of others but
maybe why don't we start with a more
General and overarching question so uh
sh you know how do you uh you know how
do people know what to learn about gen
given that the field is just changing so
fast so
um so gen AI a lot applications enabl by
AI gen AI are new but I a lot of
Technology surrounding it is not right
so so for example like vector like a
retrieval is not new at all um or um f
like on a lot of inference optimization
like
quantizations um distillations and also
not new the concept of language modeling
is not new like it was introduced um in
like
1951 um it's like a really really good
paper like CL shanon paper so a lot of
the concept are not new so so I would
say this like there are certain
fundamentals that I would recommend
people to get started so realiz this
learning learning approach of like going
in depth So like um I know actually try
not to read news because I find news
very distracting and usually there a
companies with the biggest marketing
budget usually like dominate the The
Narrative so so I I try to like Focus
like okay what problems I want to solve
and then I look into the solution to Sol
that problems and I fight out that like
one skills that would never go away is
problem problem solving like given a
problem like look possible solutions and
don't just like pick the like fanciest
uh or like most hyped solution right
pick the solution that is like works
best for you so yeah like pick a
solution a pick a problem you want to
solve and look at different solutions
and then in the process uh I think like
like on demand learning it's like when
you enter a new sub Challenge on that
like you fight out like learn more into
the concept like why is it the problem
Oh another thing is like when reading or
learning something um people tend to ask
like what and who so like okay what is
this doing or like who's doing this
right I think an important question is
like why uh so for example like if I
read about like Mi of expert I was like
um why do it using eight experts instead
of not like if we read about
quantization it was like why are we
doing eight bit why why can't we just
use one bit like what's the limitation
here like what what makes this hard so
like we try to understand like what uh
yeah why are people doing this um yeah
and why is this
hard great yeah know I really like that
um so ultimately on understanding that
uh it's built on foundations um you know
getting your hands dirty um you know and
also uh uh some really good tips U there
were quite a lot of questions
interesting questions on the topic of
hallucinations so maybe we before we
dive into some of the specific um
questions uh on this maybe can you also
explain a little bit more about like
what are hallucinations in the context
of gen and also what are your thoughts
about
hallucinations um so it's pretty funny
because like the term hallucinations in
AI has changed over time like originally
right everything that created by AI is
hallucinated because it's not real right
it's AI came about AI created those um
but I think it's like H hen announce
usually means that when um it's only
really a problem when AI creates
something gener something that effectual
incorrect so so hogen is not a bad thing
like say if I want to use AI to gener
artwork the hogen is actually good
because it's like you want something
like some creative some creativity
something that hasn't existed before um
so like um so what I focus on like
factual in inconsistency so like factual
inconsistency um so so you want to
measure so I could defy them like a two
types of factual inconsistency one I CL
as a local uh local factual
inconsistency so it's just like when you
have a fixed knowledge area and whether
it's um whether the respond is like
consistent with it or not so say let's
say it's like if you have an essay to
say the sky is purple and the moral J is
like oh the sky is purple then even
though like is not correct but it's
consistent with the essay so I would
still say especially like uh consistent
um but there like Global incon factual
inconsistency right is like you just
like basing only information out there
for example like um do vaccines cause
autism right so so so so you don't want
to like um so ask a question and AI has
to leverage on the knowledge it knows or
has access to to answer those questions
and sometimes the hardest part of it is
trying to determine not the
verifications but should determine what
the facts are because when you leave AI
to itself like going through the
internet and um with a lot of like
misinformation and people just like
putting for like like uh it's really
hard to differentiate um even like for
for humans right like uh it's really
hard so so so you need to make the task
easier for AI we focus on a local
factual inconsistency which means that
we provide AI with a ways or guidelines
or access to like information that's
relevant so the job AI now is like to
focus on like exract information from
that instead of trening to mean like
what is the correct what is the factual
things to do why is that why I think
like contact enhancement is like so
useful for for to to deal with uh
factual
inconsistency yeah I know that's that's
quite interesting so so I like that um
yeah kind of direction of uh in factual
contexts we need to make AI boring and
uh you know in others uh it has to be uh
creative uh so that's really interesting
and maybe also then diving into the
topic of uh rag um the interesting thing
is you know we have quite a large
audience uh and we're getting all the
spectrum of questions on very very
specialized and very Advanced also some
some more more
foundational um before diving into
questions related to rag I think it
would be uh great if you can maybe paint
a picture set some intuition of um you
know what what rag could look like in
practice in regards to prompt
engineering like what would a prompt
look like in order to suddenly make an
llm being able to talk to a database
like how can uh AI do that uh if you can
maybe just give us some of that that
intuition less architecturally and maybe
just like you know what does that look
like in a
prompt um so so
so this is uh let's go through this
right so let's say I have a query uh of
like
um tell me about the
movie Inside Out so the retri like first
query is sent to a retri systems So like
um it could just be like an API keyw
search and it's it's like going to some
documents so and it's just using example
keyword search so it identifies the
keyword here is inside that and it was
just search on the document just in
which one contained inside out and that
it just like paste a document into like
here owns the documents for inside out
and here's a query Tell me about the
movie Inside Out and the model takes
that as the context to generate a
response and say like inside out is a
movie that pix up blah blah blah and
send it back to users
yeahh yeah and I think I think then the
follow-up question to that um uh which
it answered uh it had asked basically
did rag actually replace fine
tuning so they addresses uh they address
different problems so maybe let me see
again actually have another slide so I
just give a talk on um on on on Tuesday
um that actually covering this
like yeah let me see um so I see like um
so so I do think if I to approach um to
picking the solutions like sometime it
can be solution Focus right oriented
it's like oh here's a solution and want
you use this how do I use this another
approach is like problem oriented it's
like okay what problems do I have and
then I see like what solution is best
suited for that problem so I would
categorize them as like
um usually if you want to enhance con um
if you want to fix information based
failures when the respond usually like
um oated so here's example of like an uh
an oated response right like if it's ask
like how
many booms studio and Booms has to Swift
released and if it respond like 10 11
it's likely because the cut of dat was
before the DAT the last un boom was
released or you can also see that's like
you can also check yourself like
sometime if find models are more likely
to hallucinate when um you ask them
about very very Niche information just
unlikely to appear in the training data
or um or just um like very rare like it
doesn't appear a lot so so you can do
that like by so that's when if R is very
useful but then if you see the problem
with um so there a lot of research
showing that like rack and Ral can work
quite well for like a lot of bench mark
so you can see a base model plus rack uh
on this paper and it's very interesting
paper that uh base model plus rack
usually works pretty well uh but
sometimes using rack on F2 model can
also like get better performance so I
think about 40% of the time rack on fun
model so it does pretty well uh in
almost no cases where phun is like
perform better on on ML um on this
Benchmark but but this not necessarily
the case right because they use some use
cases um where fune can work better and
usually they are related to to the use
case of um Behavior based failures so so
you can notice something like uh when
the appos are incorrect relevant or when
the appos don't follow the requested
format and there are a lot of ways to
like enforce this Behavior like
prompting vators uh use a better model
like when we talk about and F tuning so
so yeah so I was to say
like think rack replace F tuning is just
more of like what problems do you have
and there different solutions and
whether rack of funing works best for
for the problems is depend it's like
really
depends that's great yeah and I think it
was it was also fantastic that you were
able to pull a new set of slides uh to
do a deep dive I think uh uh the content
uh you will always have uh uh yeah a
reference U for for the for the audience
we also shared the link to your website
where they can find it um so then maybe
let's also dive a little bit deeper um
so especially now with uh this agentic
architectures um yeah how do we keep
this uh complex systems that now
interact with different like
databases uh uh uh different Services
how do we keep them safe and and
secure um so I think there a questions
of um putting God rail so so like at the
same time like when we talk about like
how to get some s here right so we need
to understand what the risk are that
where in the systems uh that things are
fall are filling and then we can put gut
rails around where the potential
failures can happen so so one Ty of
failures is um I say it's like um um
depending what you okay so like first
one kind failure is when using an
external API and you entally lick those
sensitive data to external API so some
company do that by just like not allow
just not using external API so you don't
want to S more like open AI anthropic or
Google on together uh some other
companies maybe try to say like put a um
input guard rail so you can detect like
sensitive information and then try to
block those requests from going out uh
or some model yeah so some companies
might try to host everything inhouse
instead uh another kind of like Risk is
uh let's say um I would say it's like um
also ra input is like prom injections or
like prom um so prom injections sometime
like sometime not own PR injection can
be uh harmful like it could be like
annoying like some people trying to get
AI to like say bad things right but can
be like more harmful when you give AI
models access to tools uh to so so you
can execute things um so in that case
you you have to like so do it from a
different approaches so of different
angle so first you have like makes a
system not allows a system to autoally
execute dangerous query so say if you
have give access to S database don't
give it right access right like you can
you can you can give it only read only
access like it can select but it cannot
like delete or remove update anything
right um so so so like the system
doesn't execute those uh similarly if
you give ACC to email then you can just
give it read access and not like allow
it to like send an email or like did
didn't email so so I think it's very
important to make the system um not
vulnerable to you in the first place uh
the second things is uh you can also
like uh detect queries so you can see
like um some people have um tried to
like do intent classifications of your
query to see like what is what is user
really trying to do and if it's like a
bad intention you don't serve it I think
one one really good practice I do
recommend a lot of companies do is like
the find schope questions so like uh so
somebody told me recently it's like they
went to a coffee shop when they have a
bot to ask about coffee and the boss
just refuse to answer any questions
unrelated not related to coffee and I
think it's good because I even think
okay if if you do create a chat bu for
Enterprise right to talk about like car
insurance there's no reason why it
should also like respond to question
about immigrations or like vaccines so
it's very clearly like detect like
whether the query is in scope for it or
not or just reject it uh so any other
queries could be like the model can
reveal sensitive informations so I think
one thing is like let's
say um let's say like you give model
access is a database and then a mod
retrieve the the information from the
data so maybe maybe Alexandro right you
created a request but then youra and my
data in the same database and the model
can retrieve not just your data but also
my data and so can reval it to users uh
so that could be like a race uh another
way could happen is because a model was
trained on sensitive data and then as a
inference time it was like uh reate the
information from the train data so the
solution to this like in the first one
um you can build a model predi so there
a lot of like pii information detector
so try should detect whether the pii was
like included in the in the response and
block it and the solution is like not
train model on sensitive data in the
first place but it's harder to enforce
because you a lot of us are not building
our own models we have to depend on
model providers so I do think this like
it's very important to have more
transparency are training data but
unfortunately uh companies just becoming
more crative on the train data because
training data is the competive advantage
so companies don't want to reveal too
much so there like uh they say we don't
want computer you know but at the same
time it can protect them from scrutiny
and like uh legislature so yeah so so
the whole Space is quite uh
scary yeah no that is really great I
mean I think I think the the topic of of
um yeah safety security guard rails um
yeah could be its own its own Deep dive
um it does feel like we're back in the
90s with SQL injections uh yeah and um I
do remember you in in one of your your
uh uh posts SL documents uh you had a
very nice um uh uh uh overarching
architectural blueprint of basically the
archetypal um design for for for um yeah
agentic systems which has basically the
input guard rails the output guard rail
so yeah I can put it it's like longer
talk I don't really want um it's like
it's it's pretty complex system so yeah
so like you have like so B Contex
construction here you have G rails so
you so basically you want to put guard
rails everywhere the system can fail
like if it's like fail as a input do it
if it fail as an output do it if it fail
as my retrieval do it uh so yeah so got
is just maybe like um I was like fancier
try
cash yeah yeah yeah that's that's super
interesting okay so so let's um yeah
maybe ask one or two more questions
before we wrap up but one question that
I think uh has come up or uh multiple
questions that we can uh I guess reduce
to to a single one um we have different
people with different backgrounds um I
think uh you know some coming in you
know for you know with uh in PhD or or
um yeah basically uh uh still University
others uh in software engineering uh
roles and others in machine learning
engineering roles um so yeah what career
advice would you give for someone um
yeah that wants to get uh into the field
or basically into a job of not just
machine learning engineering but
specifically AI engineering
yeah uh B career advice um I feel like
career is something very personal um so
I would probably try to
understand um yeah I think it's really
hard to give like blanketed um advice
that works for everyone because sometime
it works for me it works for my friends
but not working for other people and the
same time I try advice work on my
friends just doesn't work for me um so
so I would say this like
um yeah I don't think one most important
skill is like the learning she learned
and then problem solving um
like was probably spend less time like I
don't know uh so so this some heris uh
one I forgot the name of it but
basically idea is that um how long
something will last in the future is the
same amount of time that it has last in
the past right so so like uh so say like
if something has been around for 10
years and we can use your to estimate
that be last for another 10 years so so
relationship it's a same relationship
right you just oh yeah L effect thank
you John um so the first thing is like
um like if you in relationship with
someone like you just start dating last
week it's unlikely you know like you
don't want to plan something like a year
in the future but you've been with
someone for like four years maybe it's
reasonable to to like plan something a
year in the future so the same thing
it's like a lot of new things that comes
out a lot of them and probably some of
them will be very important but then if
they really important just like wait a
little bit you know like and see if
that's a case you don't immediately have
to jump on or like if you jump on do ask
questions because a lot of marketing
space so i' been like actually a
compiling but I call like marketing
speech uh translator because some
sometimes we come like oh here's this
fancy technique to do this and that
alternative is something very simple but
dress up in something like let's try
asking a lot of questions uh also like
when you read something like who is
publishing it because sometime um for
example like you see like some startup
coming out with like crazy exciting
things and it realizes like everyone who
is pushing for it as investor in the
company so like understanding who are
like propagating the information so
basically just ask a lot of questions
well that that is great that is great
and and and I really like that point of
um yeah also the fundamentals and the
things that will last I mean you
highlighted that that rag Concepts like
that you know have been you know uh
there for for a long time and that
proves it's uh resilience and you do
have also a very nice blog post uh on
measuring uh your career progression uh
I think you do delve into some some
similar areas so indeed uh I think uh uh
uh uh unfortunately
uh we are uh now out of time uh we have
had a very very great discussion uh uh
with a lot of really great questions um
so I do really want to give a massive
massive thanks to
chip great presentation yeah uh sorry uh
please uh go ahead yeah chip oh no thank
you so much everyone uh I really
appreciate your time and I'm sorry if I
spoke too fast and I'm sorry I couldn't
get your own the questions
uh I'm on social media um so yeah please
um say hi um I'm pretty bad social media
because it fight incredibly distracting
uh so so every has this fight of like I
know that like something like LinkedIn
and Twitter is could be good could have
me useful information but like it's like
trying to balance the cost and benefits
you know sometime yes it can give me
like 10% information but like 90% of
which is like distractions so anyway uh
and so yeah I I don't know how you
answer that question but yeah I hope I
see you all at some point in the future
uh and this um if you're rich child uh I
would try my best to respond but
sometime I cannot get to like one of
that yeah yeah no and this is this is
really great uh so Chip I think uh yeah
the whole presentation was great um the
questions also were great um and one
reminder that indeed the talk will be uh
recorded and is available uh in a few
days uh in learning. acm.org so you will
be able to replay it multiple times in
case you missed anything so nothing to
worry uh as all this information is
going to be available there as well as
the slide deck which was also shared um
please uh also a reminder to fill out
our quick survey uh you can suggest uh
future topics or speakers so please do
take a moment to to fill uh those up as
it helps us to make it ever better
um so on behalf of ACM uh of Chip and
myself thanks again for joining and I
hope you join us uh once again in the
future and with this uh we conclude our
talk and we wish you a great week ahead
thank you
all questions can I get the questions
that I couldn't get you so I try to
address them
Ask follow-up questions or revisit key timestamps.
The video features a presentation by Chip Huyen on the transition from Machine Learning (ML) engineering to AI engineering. Huyen explains that AI engineering involves building applications using foundation models. She covers essential topics such as evaluation challenges, context enhancement techniques like Retrieval-Augmented Generation (RAG), and the necessity of guardrails in agentic architectures. The discussion highlights that while the tools and hype evolve rapidly, many core concepts are foundational and resilient. She advises focusing on problem-solving over jumping onto the latest hyped technology and emphasizes the importance of evaluating applications based on specific failure modes.
Videos recently processed by our community