HomeVideos

From ML Engineering to AI Engineering

Now Playing

From ML Engineering to AI Engineering

Transcript

1340 segments

0:06

great

0:09

great

0:11

so we are ready to go so hello everyone

0:17

and uh welcome to today's uh ACM Tech

0:21

talk so this uh webcast is part of acm's

0:25

lifelong commitment um uh for Learning

0:28

and professional development um serving

0:30

This Global membership for computing

0:32

professionals and students um I will be

0:35

today your moderator my name is

0:36

Alejandra sedo uh I am currently member

0:40

at large uh for acm's uh governing board

0:43

as well as uh director of engineering

0:45

science and product at zando uh I am

0:48

also uh scientific adviser at The

0:50

Institute for ethical AI uh where I've

0:52

LED contributions to EU policy including

0:54

things like AI act data act etc etc so

0:58

for those of you who may be a little bit

1:01

unfamiliar with the ACM um or what it

1:04

has to offer here is uh some information

1:07

so ACM offers educational and

1:09

professional development resources that

1:11

bolster skills and enhance career

1:13

opportunities uh you can see some of the

1:15

highlights on your screen so these are

1:17

basically some of the uh key areas uh

1:21

that ACM is known for uh some of the

1:24

things to um highlight that are worth

1:26

noting is that ACM provides access to

1:29

what is referred to as the ACM Digital

1:31

Library this is the world's uh most

1:33

comprehensive database for computing

1:35

literature um ACM also provides leading

1:38

Publications and Global conferences that

1:39

draw top experts from a broad spectrum

1:42

of computing topics um they also provide

1:45

support for Education and Research uh

1:47

including curriculum development teacher

1:49

trading and uh perhaps also you may have

1:51

heard about the ACM uh touring and ACM

1:54

pricing Computing Awards um as well as

1:57

the ACM code of ethics uh which is a

1:59

collection of principles and guidelines

2:01

designed to help Computing professionals

2:03

make make ethically responsible

2:05

decisions so we have also published just

2:08

recently uh things like the um

2:11

generative AI principles which are of

2:13

course using this as the backbone as

2:15

well as the principles for algorithmic

2:17

responsibility uh which also served as

2:20

uh the backbone for several of other of

2:22

the ACM uh policy products uh before we

2:26

get started um let's mention a couple of

2:29

housekeeping rules rules um if you have

2:31

any questions at any time please uh type

2:34

them in the zoom's uh Q&A button I know

2:38

that you can also use the chat but the

2:39

Q&A button will allow us to collect it

2:41

more systematically uh we will then be

2:44

able to um yeah raise them as as the as

2:46

the um uh talk finishes the session is

2:50

being broadcasted uh uh and recorded uh

2:53

and it will be archived so you'll

2:55

receive an email notification when it's

2:56

made available and check learning.

2:59

acm.org for updates on upcoming

3:03

webcasts so today we actually have a

3:06

very very exciting topic uh with uh also

3:10

a very exciting uh presenter uh the

3:13

presentation is titled from ml

3:15

engineering to AI engineering and we

3:18

will be having uh chip uh who uh Works

3:22

to accelerate data analytics on gpus at

3:25

Voltron data as their VP of Open Source

3:29

uh and AI uh previously she was uh at

3:32

snorkel Ai and Nvidia and founded an AI

3:36

infrastructure startup uh that was

3:38

acquired uh she has also taught machine

3:40

Learning Systems design at Stanford and

3:43

she's the author of the book designing

3:45

machine Learning Systems which is an

3:47

Amazon uh bestseller in AI a lot of the

3:50

insights today are are going to be uh

3:53

providing a lot of depth on this so with

3:56

that uh and without further Ado um uh ch

4:00

please uh take it

4:02

away thank you so much aleandro for the

4:05

very intro and hi everyone I'm looking

4:08

at the chat right now and this is

4:10

incredible uh diversity in uh the

4:13

audience so feel free to like ask

4:16

questions when I'm speaking and I know

4:18

is that sometime going to get excited I

4:19

speak pretty fast so feel free to tell

4:22

me just slow down as well um so yes

4:26

so wait oh yeah my somehow my video was

4:29

not on I hope you can see me now hey hi

4:32

again um so yeah so today I'm talking

4:34

about from ma shielding engineer to AI

4:37

engineering as alandro mentions uh I

4:40

wrote um I'm the author of Designing

4:42

Machining systems and I'm working on a

4:44

new book that is similar like just built

4:47

on top of Designing Machining systems

4:50

but focused more on Foundation model so

4:52

we cover everything from like prom

4:54

engineering evaluations uh F tuning um

4:59

context r for actions agents and all the

5:02

things that you might have questions

5:03

about so in the process of working on

5:06

this book I do a lot of research so I

5:09

try to look at every

5:11

single AI tool out there and it's a very

5:15

time consuming things because there are

5:17

so many of them so to make my life

5:19

easier I only focus on repos at least

5:21

500 Stars so you can actually see them

5:24

on my blog uh on my website

5:27

h.com Lama poish so I think it's I find

5:31

pretty useful to like uh so like if I

5:34

wondering like who is doing inference

5:36

optimization so I can look at like I can

5:38

filter them out by the category of like

5:40

inference optimizations I want to see

5:42

who is doing Asian who is doing AI who

5:44

is doing cing um who is doing Serving so

5:48

it's been very helpful and you can see

5:50

like the number of changes like because

5:52

I notice a pattern just like I got a CO

5:54

the hype curve so you will see a lot

5:56

like sometime a repo come out and you

5:59

see a lot of excitement and a lot of

6:01

stars but then nobody ever uses this

6:04

again and it was just like flatten out

6:06

so I think it's if want to track and not

6:07

just a number of stares but like how

6:10

fast this is growing um okay so today we

6:13

talk about a engineering so engineering

6:16

is the process of building applications

6:19

with Foundation model so I use the term

6:21

Foundation models because it's a super

6:24

set of large language model um so I I do

6:27

think it's like we we seeing the

6:30

conversion of a lot of different

6:32

modalities when I started out in AI in

6:35

2014 we saw that like NLP natural

6:38

language processing was his own field uh

6:40

computer vision was his own own field

6:42

like reinforcement learning was its own

6:44

field but do see that like things are

6:46

like coming together so we see this like

6:48

a lot of models today like

6:51

gbd4 or Gemini they own incorporate not

6:55

just text but also images and we see

6:57

reinforcement learning being used train

6:59

model is like reinforcement learning

7:00

from like human preference so it's been

7:02

very interesting to see that um and

7:05

Foundation I think the term was coed by

7:07

Stanford because it also denote uh the

7:10

ideas that the concept that Foundation

7:12

model just ler foundation and it there a

7:15

need to be adapted to specific

7:18

needs um so there are a lot of changes

7:21

in um from like machine learning

7:23

engineer to engineering so one thing is

7:26

that models are getting really big big

7:29

and very few people can afford to train

7:32

them so we see they are like model

7:34

providers model developers who develop

7:36

models and make them available to use as

7:39

a service so this service you can access

7:41

through API or you can download them

7:42

like as an open weight model and you can

7:45

host them to S so this means that like

7:48

anyone even those with no AI uh

7:52

knowledge can now leverage a bu

7:54

applications so before right like always

7:57

is a company who could afford to build a

7:59

applications as a a big a company to

8:01

have data who can train their own model

8:04

and then these companies and then these

8:06

AI products are incorporated into their

8:08

main product but now like you can just

8:11

build it but then now you need to figure

8:12

out a way how to serve those

8:14

applications so one thing is very

8:16

exciting that we see a lot of interface

8:18

to serve the AI applications such at

8:20

like plugin so it can incorporate AI

8:22

into the browsers Google Docs or like um

8:25

vs code you can also see uh I'm very

8:28

excited about like getting like

8:30

embodiment of AI so for example we can

8:32

have like audio or like 3D character I

8:35

think like the ultimate goal is that we

8:37

have like AI embody in like physical

8:39

robots that can help us do things rather

8:43

well um so yeah so I think that AI has

8:45

become a common component in a software

8:47

engineering the same way the JavaScript

8:49

and databases are so uh I noticed like

8:52

engineering is coming like closer to F

8:54

speack than like traditional data

8:57

scientist so another changes uh is that

9:00

we are moving away from like uh we are

9:02

moving from close ended evaluation to

9:05

open-ended it's not moving away it's mon

9:08

need combinations of of of a lot of ways

9:11

to evaluate so Foundation models are can

9:14

gener open-ended responses and this

9:17

makes them really really powerful

9:19

because now with open-ended responses

9:21

you can adapt them to like different use

9:25

cases oh thank you yes I'm G to try to

9:28

speak slower um yeah

9:31

so yeah open-ended responses make um

9:35

being being open-ended makes partion

9:38

models useful for lot of use cases

9:40

because now can ask like say gbd4 should

9:43

do translation should write essay should

9:45

write code uh should do math however the

9:48

open-ended nature of models also make it

9:51

really really hard to evaluate so before

9:54

if you have say a classification task

9:56

right so you have the model who

9:59

maybe like one of those three categories

10:02

and say if you expect category like a

10:04

positive and the respon is not positive

10:07

then it knows that the respon is

10:08

incorrect but now it's very hard because

10:11

um there are many different ways to

10:14

express the same idea and even if some

10:16

model out put something that's like not

10:18

the same as expected sentence expected

10:20

text it's really hard to tell uh whether

10:23

is um is incorrect and also like as

10:26

models get more intelligent it also

10:28

becomes harder to evaluate them so

10:31

imagine like anyone can tell if the

10:34

solution to a first grade math problem

10:37

is wrong but it's very few people can

10:39

tell whether the solution to like a PhD

10:42

level math questions can um is wrong so

10:46

now that like less obvious mistake we

10:49

can detect from Ai and I so the

10:51

incorrectness or like inaccuracy of AI

10:54

now you have to do fact check in one

10:56

critical thinking and make it a lot

10:57

harder to evaluate

11:00

uh so we are seeing a lot of new

11:01

evaluation approaches for the

11:04

applications so one thing uh one

11:06

approach is comparative evaluations so

11:09

basically you you see um with this

11:12

approach you see two different responses

11:15

side by side and users pick which one is

11:18

better so this different from AB testing

11:20

because in AB testing you only see one

11:23

response uh at a time uh another

11:26

approach I that used to cause a lot of

11:28

controver icies but I do see that is

11:31

getting more adoptions which is like

11:33

using AI models as a judge so you can

11:36

ask um sometime you think it's like

11:39

using um evaluating whether an AI model

11:42

is good by like asking his friend so you

11:45

have a model to generate uh responses

11:48

and another another model or is the same

11:50

model to evaluate whether that respon is

11:52

good or not so there are many different

11:54

ways uh ai ai judges can be formulated

11:57

so one way is just like like a scoring

12:00

like a reward model uh so scoring is you

12:03

ask a model to Output like maybe a score

12:04

from one to five on the uh respond with

12:08

is good a reward model um or like some

12:11

others other models can also do

12:13

comparative evaluation for you so you

12:15

can ask a model to like pick which of

12:18

the two responses that humans would

12:20

prefer so I'm very excited about this

12:22

category of AI judges because if we

12:25

could uh find a reliable way to like

12:29

predict what humans prefer that could be

12:32

the goal right we wouldn't need a lot of

12:33

like humans annotation to create uh data

12:36

to train the model because like human

12:38

preference is is what like a lot of

12:40

people want for their models because for

12:43

this model to align with uh what humans

12:46

want um yeah and like other AI as a

12:49

judge is like you just ask a model like

12:50

hey does this s correct I given the

12:53

answer or like is this like relevant to

12:55

the questions but it's it's exciting but

12:58

it's like very hard to use because you

13:00

need a lot of rigor in making those

13:03

matrics standardized um that uh we are

13:07

doing okay I'm still speaking pretty

13:09

fast and I don't see

13:14

um I don't see the

13:17

audience it's VI out approach evaluation

13:21

we won't get back to this later maybe in

13:23

the Q&A I love that um so I wish had a

13:25

conversation yesterday with a few

13:27

friends and some the smart artist people

13:29

I know I was asking them like how do you

13:31

evaluate the model and they were like oh

13:33

we have about like five pet respond like

13:36

five queries five prompts that we use to

13:39

test and was like wow um things it's

13:42

like Vibe check is very important but

13:45

it's not sufficient so you do need to

13:47

like interact with the models every day

13:49

uh like to see the to menu see that to

13:52

like for for um because over time you

13:56

might have better idea like as you

13:57

direct the model you have a better idea

13:58

of whether what what consider a good

14:00

response one is the hardest challenge of

14:02

evaluations it's like creating the guid

14:04

like um so I so my friend he used to

14:09

like um like create guy life annotators

14:12

but then you you need to like it's

14:14

really hard to Define like uh some uh

14:17

you sometime need like to train people

14:18

to do so um so so yeah so like one thing

14:22

is you have to create very clear

14:23

guidelines to like what is good and bad

14:25

and only by interacting with the models

14:27

that you get have a better understanding

14:29

of like is response really good or not

14:32

because sometime a response can be good

14:34

but not not sympathetic so LinkedIn has

14:36

an interesting case study so they found

14:38

out that like um for they call is like

14:41

job assessment chatbot so when a

14:43

candidate see a job uh posting uh it can

14:46

ask AI like hey am I going to be a good

14:48

candidate for the job posting so the AI

14:51

can respond like no you're not right so

14:54

even though it's correct it's not very

14:56

sympathetic so they need to stay the

14:59

response the guy what could be a good

15:01

response so the respon is like maybe

15:03

like more encouraging like uh today

15:05

you're not but if you go to this skill

15:08

ABC then it would be a better fit and

15:10

here are the resources you can use to

15:12

like go to to make yourself um suitable

15:15

for this

15:17

role so I do think that evaluation is

15:19

the biggest challenge for JF AI

15:22

applications so one uh one col of

15:26

evaluation is like if we don't have a

15:29

way to evaluate whether model responses

15:33

are good or not just like too much like

15:35

it can make AI application too risky for

15:37

a lot of applications so today we see a

15:40

lot of Enterprises using um deploy AI

15:44

applications not because these

15:46

applications are potentially the most

15:48

impactful but they are the easiest one

15:50

to evaluate so for example a lot of

15:52

people want to do say um recommended

15:55

system because for recommended systems

15:57

you can evaluate the impact very clearly

15:59

because it will increase something like

16:01

click through rate or purchase through

16:03

rate or a lot people do coding because

16:05

with coding you can evaluate model using

16:08

a functional correctness so is a Model

16:10

gener A piece of code you can you can

16:12

execute the code and see if it runs and

16:15

whether it produce expected results so

16:17

because of that um if we can fight a

16:20

which to reliably evaluate responses it

16:23

believe that we can unblock a lot of

16:25

exciting applications

16:29

um okay so another things is like um I

16:31

also see more of the transitioning from

16:34

like future engineering to context

16:36

contractions so with classical machine

16:39

learning you need to like engineer

16:41

features like what could be uh so so

16:43

that for age of the query it would need

16:45

like creative features is important to

16:48

uh that's you think would be useful for

16:50

the model to make predictions so with J

16:53

AI model Foundation model was they use

16:55

prompt so for queries you need like uh

16:58

you need to inject useful information to

17:01

the prompts uh to the context so that

17:03

the model can respond uh to the query so

17:07

um Hallucination is an interesting uh is

17:10

an interesting uh term because uh

17:13

hallucination make it sound like a bad

17:15

thing but it's not always necessarily a

17:17

bad thing like you do one model like

17:19

sometime like be creative so it's

17:21

actually very useful for cre creative

17:23

applications um but sometime like for a

17:26

lot of use cases when you demand a

17:28

strict um uh like so some case when like

17:32

curacy is very very important like legal

17:35

or medical you you don't want models to

17:38

like make up things so we do see that

17:40

like many Studies have showed that like

17:43

models are more likely to hallucinate

17:45

when it when they don't have the right

17:48

information so uh so models have this I

17:52

think it's like a bias with actions so

17:54

even when they don't know something or

17:56

they unsure they just somehow like have

17:58

to respond and it had to like make up

18:00

something like if you if you don't know

18:02

the answer um I'm not quite sure like

18:04

how to correct that so I think part of

18:06

the training because like in the

18:08

training process right uh for on the

18:09

training us have like prompt response

18:12

every prompt isn't so there had to be a

18:14

response so I do think it's like to

18:16

correct that bias is going to be quite

18:18

challenging but anyway in the meantime

18:21

when you work with this model we have

18:22

seen that um models can perform much

18:25

better and hallucinate a lot less when

18:28

they apprach provided with the necessary

18:30

information to answer the question um so

18:33

like how do we enhance context like how

18:35

do we provide models with the correct

18:38

like with a relevant information uh

18:40

necessary for it to like respond so we

18:44

can do it to way like once it like with

18:46

retrieval so you provide access to your

18:49

databases quick documents images table

18:52

data so that they the model can retrieve

18:54

the relevant information from the

18:56

databases or you can also give model

18:59

access to tools so that it can gather

19:02

information sself so some tool like web

19:04

search is extremely useful so when um

19:08

the pattern of giving model access to

19:10

tools it also known as a gentic pattern

19:13

and we can go into more and chooses in

19:15

the Q&A if you're interested I think

19:17

it's very exciting but it can be a

19:19

little bit overhyped because you're

19:21

going to think of like um yeah we can go

19:23

more into this later so here's some like

19:26

visualizations of what enhancement uh

19:28

cont enhancement would retrieval look

19:30

like so it gives a model access to

19:32

external data databases so which be like

19:34

a collection of documents or images and

19:37

you have a retrieval system so like

19:38

given a query it will retrieve the

19:41

relevant information from the from the

19:43

databases and then the context is Joy

19:46

with a query so that the model can

19:48

generate a respon and some back to

19:51

users um so R for is old science so it's

19:55

not new I think I believe that's the

19:57

first digital

19:59

retriever systems would describe back in

20:01

the 1920s so it's like a century ago and

20:05

retrieval has been a backbone of a

20:08

search and recommended systems um so

20:11

it's it's not new so they have many

20:12

technology algorithms developed for

20:16

retrieval search so in general there are

20:18

two main approaches to a retrieval one a

20:21

turnbas so first of all like keyword

20:23

search like if you have a query like um

20:27

about say Transformer architecture it

20:30

can look into the database and search

20:32

for all the document that contain

20:33

Transformer architecture uh so there are

20:36

lot of term based Solutions um that

20:39

available like bm25 and elastic search

20:43

none of them is new but they have very

20:45

very powerful and like a lot easier to

20:47

stud it than the other approach which is

20:49

embedding based um retrial or like some

20:53

people call it a lot people call it like

20:54

uh Vector search so Vector search is

20:57

more complex um is like take longer um

21:01

is is more comp comp computationally

21:04

expensive but it can potentially provide

21:07

like much stronger performance because

21:08

you can have a lot of room for fun um

21:11

but at the same time I do things it's

21:12

like um if you're just studying out I do

21:14

things that like um I do think that a

21:17

termbase solution like bm25 is a very

21:20

robust um approach and is a very very

21:24

strong basine to beat so um when you

21:28

reach retri data right you don't just

21:29

retrieve structur data you also want to

21:32

retrieve like tabular data as well so

21:34

maybe you have like a c database with a

21:36

lot of SQ tables and I have an example

21:38

here so retrieving table data is very

21:42

different from retrieving like

21:44

unstructured data because now to aeve

21:46

retrieve data from from this a SQL table

21:49

you need to like write SQL query so the

21:51

first step would be like you have to

21:53

convert from like query like natural

21:55

language into the SQL query uh the

21:58

flavors that your database can

22:00

understand and then after that you

22:02

execute the SQL um the SQL query so this

22:06

like a bit trickier um and it can be

22:09

quite dangerous so say like what if like

22:11

people want to like execute a dangerous

22:14

SLE query like let's drop the table or

22:16

something like that so you would need to

22:18

like put in gut rail like like what kind

22:20

of queries that you can uh the model can

22:23

execute um so so yeah so in this case um

22:27

Al Texas pretty tricky because I choose

22:29

Jared a SLE query you need not just a

22:32

query but also need to understand the

22:33

table schemas so I say if you have a lot

22:36

of tables and the table schemas is too

22:37

long to fit into context length it could

22:40

be uh it could be hard uh so so and also

22:43

like like pick like what tables are

22:45

relevant so some people are trying to

22:47

build uh embeddings for tables based on

22:51

the table description and table schemas

22:53

so that for given query it can

22:55

automatically retrieve all these

22:57

relevant tables and then based on the

22:59

schemas of this relevant table jate a SQ

23:03

query for

23:04

those um yeah so so here example of like

23:09

H text queries can convert into s query

23:12

and it would need uh to know that you

23:15

are using the table

23:18

sales um so yeah you can also use retri

23:21

tool so here's what it looks like so

23:22

basically you need to give models uh you

23:25

need to define a set of that you that's

23:29

a model G take so action could be like

23:31

using a search API using a news API

23:34

using weather I think um the most

23:37

exciting tools that uh people wanted for

23:40

for AI like especially on CH stuff with

23:43

like web browsing because now with web

23:46

browsings you can allow model to stay up

23:48

to date so with our web browsing right

23:50

like um informations can be like um

23:53

models can get outdated so you ask a

23:55

model about something that happened

23:56

today and the information is not in the

23:58

database or it was not in the model

24:00

training data the model won't be able to

24:03

respond but with like web search the

24:05

model can stay a lot more relevant um I

24:08

think my website is like very

24:09

exciting uh okay so so I grouped on of

24:13

this um retrieval from databases ands

24:17

into what they call like people call

24:18

like context constructions um yeah so

24:21

like the go to enhance context so that

24:23

the US so that the models can answer

24:27

questions better and also reduce

24:30

hallucinations okay so so another change

24:33

um I know that we have been going

24:35

through like a lot of content I hope

24:37

that it's uh it's fine it's okay but now

24:39

feel free to ask questions we don't have

24:40

time to go over them um so the model's

24:42

going to get a Bigg like become bigger

24:45

so I remember in the early day uh we

24:48

talk like oh my God one billion um

24:51

parameters with like so much just crazy

24:53

like I remember when the first GPT model

24:55

came out so it was like uh 1 million

24:59

parameters and was like oh my God that

25:01

is so big and now we're like hundred of

25:03

billions and I think people are aiming

25:04

for like a trillion parameter models so

25:08

the the number of parameters uh is not

25:10

the only things that matter right it

25:11

also depends depends on how sparse the

25:14

model is because like the more SP model

25:17

is uh So the faster so the less

25:19

computational computation uh

25:21

computational resources are required to

25:24

run the model but anyway in general

25:26

bigger models require more uh more

25:29

computational resources to run so they

25:31

are more expensive to run they also have

25:33

higher higher latency and also it's just

25:35

like if you want to host those big

25:37

models it's require more expertise

25:40

because like it just don't simply put on

25:42

the machine you need to work on like

25:43

distributor system how to load dat this

25:45

you you know in um in um how how to sh

25:48

data so it's a lot harder to operate

25:51

larger model and then smaller model uh

25:53

so that also making that like um a lot

25:55

of Technologies to make this model run

25:59

faster and makes them smaller very

26:00

exciting so so one thing uh that I'm

26:03

very excited about is like inference

26:06

optimizations so inference optimization

26:08

is not new um so when I'm when I'm like

26:11

just uh so actually I used to take a

26:13

course of Stanford about um five years

26:15

ago not four years ago wow 10 fles um

26:19

but um I was like covering inference

26:21

optimizations and then I revisit the

26:23

topic

26:26

again sorry people still cannot hear me

26:30

um okay I'm going to try to speak

26:33

slower um yeah so so when I revisit the

26:36

topic of inference optimizations

26:38

recently and I realized that a lot of of

26:41

core techniques for in inference

26:43

optimizations still remain the same um

26:46

so so like the qu thing for example like

26:49

um quantizations quantization extremely

26:51

exciting it works pretty well like for

26:54

wide range of models and applications um

26:57

it's is not new right it use the

26:58

precisions from like um float uh 32 uh

27:03

to like float 16 and now8 bit and now

27:06

people even going for like four bit um

27:09

so another technique is um is low rank

27:12

low rank factorizations which is

27:14

basically like the core of like on the

27:16

Laura fing technique and also like they

27:18

like um sparcity with pruning and now

27:21

there new mixture of expert and also

27:24

with distillations where you trained a

27:26

smaller model to uh to imitate the

27:30

behavior of the larger model and of

27:32

course you are things like cash uh

27:34

parameter efficient F tuning so all of

27:36

that space is very

27:38

exciting okay so that is pretty much for

27:41

my presentations just for one last thing

27:43

I want you to cover uh so like see like

27:46

think of like for there are so many

27:50

adaptation techniques out there like to

27:53

your adap model but I would generally

27:55

put them into like three category like

27:57

one is like just prom engineering uh

27:59

like so you try to write instruct um

28:02

write clear instructions they compose

28:04

like a big task into like smaller

28:07

subtabs add think a lot of examples and

28:09

then the next one is like context

28:12

optimization so here I put rack but they

28:14

can be on so like action API as well so

28:16

the general retrieval plus actions and

28:19

the other category is like f tuning um

28:22

so fing I would say is like it's never

28:25

the first answer so I would say over

28:27

time as you progress through the

28:29

applications I think it's very important

28:31

to like get the most manage out of PR

28:33

engineering with like use very use good

28:37

examples with very clear guidelight for

28:40

like the model like good good example of

28:41

like how the model should act for

28:43

example if you want model to say give a

28:46

score from one to five Define very

28:48

clearly what is what does one mean what

28:50

what is an example of in one look like

28:53

and why it get it get a score of one and

28:55

the same thing one two three four five

28:58

um and after that you can try with like

29:00

simple like retrieval for example like

29:02

BM uh like keyword search or bm25 and

29:06

after that you can choose whether go to

29:08

more complex retrieval uh like vector

29:11

search because it would require you to

29:13

like use an embedding model and use a

29:16

vector database like sorry Vector search

29:18

which can be used as part as part of a

29:21

vector database so so yeah so like

29:23

because it's add more component to the

29:25

system it it increase the complexity or

29:28

it can also go as a f tuning rout uh and

29:30

then you can also like Joy together like

29:32

using r with a f tune model or using R

29:35

to like create um training examples like

29:39

uh to to fune the model so yeah so I

29:42

think that what what it looks like um um

29:45

so in general I would categorize so uh I

29:49

would decide what technique you use also

29:52

depends on what kind Behavior you want

29:55

the application to have or like what

29:57

kind of issues you want to correct for

29:59

the applications so I would in general

30:01

think of them as like uh failure modes

30:03

could either be information based so

30:05

like when the response are incorrect uh

30:08

or just outdated or another what I call

30:11

is a behavior based failure mode so when

30:14

the responses are like factually correct

30:16

but it's like either irrelevant uh or it

30:19

just doesn't follow the expected format

30:21

so for the information based um arrrows

30:24

I find the rack like context

30:26

optimizations helps really well but for

30:28

Behavior based uh like getting the model

30:31

should like follow the format then you

30:33

can try that with like um prompting

30:36

adding more example using vators um um

30:41

using some f technique like um

30:44

constraint samplings or using like

30:46

bigger models because as models become

30:49

more powerful they also getting better

30:51

at like following instructions and Jing

30:53

the correct format or you can fight tune

30:55

so F tun I think of like as a last L of

30:58

fans uh because fight tuning is exciting

31:01

but it also require both a higher a

31:03

front cost because you need to curate

31:06

data for f tuning and it also require

31:09

continual maintenance because you have a

31:11

model and now what right you have to

31:13

update it regularly and like what if the

31:15

newer model base model comes out do you

31:17

want to leverage a new base model which

31:19

mean you have to like maybe fight T

31:21

again so it's a lot of like cost

31:23

associated with f tuning that you really

31:25

should consider before committing into

31:28

it okay so um um so this is like a

31:32

little un ready to talk it's about like

31:34

what our work is doing so we build GPU

31:36

native CER engines um so the idea is

31:39

that gpus now are everywhere so I think

31:42

one one consequence of AI is that like

31:46

now a lot of companies have gpus and a

31:48

lot of people are using gpus for

31:50

inference and training but I do think a

31:53

big category is like data processing um

31:57

data processing um like ETL I do think

32:00

this like can be very suitable for gpus

32:03

because uh let's say if you want a model

32:06

if if you want to process a billion rows

32:08

of data you can just Farm z m to like

32:11

thousand of fours of GPU so you can

32:14

speed things up and we can also like

32:15

just make things like uh a lot cheaper

32:18

um we we also also contribute to open

32:20

source project like IIs and apach arrow

32:22

um yeah and if you want to find me you

32:25

can find me on my blog LinkedIn Discord

32:28

uh so yeah thank you so much

32:33

everyone great great let's please all

32:36

get our virtual hands together to give a

32:40

Applause to chip uh this was a fantastic

32:43

talk also big props uh chip for also

32:47

keeping an eye on the chat uh very you

32:50

know active uh uh reactions on on some

32:53

of the comments and even the questions

32:56

so we have uh lot of really great

32:59

discussion and much more to come uh we

33:03

have been able to group some of the

33:05

questions into different topics so we

33:08

will dive into what seems to be um areas

33:11

of uh hallucinations rag prompt

33:15

engineering AI judge evaluation uh gen

33:19

limitations and a couple of others but

33:21

maybe why don't we start with a more

33:23

General and overarching question so uh

33:27

sh you know how do you uh you know how

33:29

do people know what to learn about gen

33:32

given that the field is just changing so

33:38

fast so

33:41

um so gen AI a lot applications enabl by

33:46

AI gen AI are new but I a lot of

33:50

Technology surrounding it is not right

33:52

so so for example like vector like a

33:54

retrieval is not new at all um or um f

33:58

like on a lot of inference optimization

34:00

like

34:01

quantizations um distillations and also

34:03

not new the concept of language modeling

34:06

is not new like it was introduced um in

34:09

like

34:09

1951 um it's like a really really good

34:12

paper like CL shanon paper so a lot of

34:15

the concept are not new so so I would

34:16

say this like there are certain

34:18

fundamentals that I would recommend

34:20

people to get started so realiz this

34:23

learning learning approach of like going

34:25

in depth So like um I know actually try

34:28

not to read news because I find news

34:30

very distracting and usually there a

34:32

companies with the biggest marketing

34:34

budget usually like dominate the The

34:36

Narrative so so I I try to like Focus

34:39

like okay what problems I want to solve

34:41

and then I look into the solution to Sol

34:42

that problems and I fight out that like

34:44

one skills that would never go away is

34:46

problem problem solving like given a

34:48

problem like look possible solutions and

34:52

don't just like pick the like fanciest

34:54

uh or like most hyped solution right

34:56

pick the solution that is like works

34:59

best for you so yeah like pick a

35:01

solution a pick a problem you want to

35:02

solve and look at different solutions

35:05

and then in the process uh I think like

35:08

like on demand learning it's like when

35:10

you enter a new sub Challenge on that

35:13

like you fight out like learn more into

35:14

the concept like why is it the problem

35:17

Oh another thing is like when reading or

35:19

learning something um people tend to ask

35:22

like what and who so like okay what is

35:25

this doing or like who's doing this

35:27

right I think an important question is

35:29

like why uh so for example like if I

35:31

read about like Mi of expert I was like

35:34

um why do it using eight experts instead

35:37

of not like if we read about

35:38

quantization it was like why are we

35:41

doing eight bit why why can't we just

35:43

use one bit like what's the limitation

35:45

here like what what makes this hard so

35:46

like we try to understand like what uh

35:48

yeah why are people doing this um yeah

35:51

and why is this

35:53

hard great yeah know I really like that

35:56

um so ultimately on understanding that

35:58

uh it's built on foundations um you know

36:00

getting your hands dirty um you know and

36:03

also uh uh some really good tips U there

36:06

were quite a lot of questions

36:08

interesting questions on the topic of

36:11

hallucinations so maybe we before we

36:14

dive into some of the specific um

36:17

questions uh on this maybe can you also

36:20

explain a little bit more about like

36:22

what are hallucinations in the context

36:24

of gen and also what are your thoughts

36:28

about

36:29

hallucinations um so it's pretty funny

36:33

because like the term hallucinations in

36:36

AI has changed over time like originally

36:39

right everything that created by AI is

36:40

hallucinated because it's not real right

36:42

it's AI came about AI created those um

36:46

but I think it's like H hen announce

36:49

usually means that when um it's only

36:52

really a problem when AI creates

36:56

something gener something that effectual

36:58

incorrect so so hogen is not a bad thing

37:01

like say if I want to use AI to gener

37:03

artwork the hogen is actually good

37:06

because it's like you want something

37:07

like some creative some creativity

37:09

something that hasn't existed before um

37:12

so like um so what I focus on like

37:15

factual in inconsistency so like factual

37:18

inconsistency um so so you want to

37:20

measure so I could defy them like a two

37:23

types of factual inconsistency one I CL

37:26

as a local uh local factual

37:29

inconsistency so it's just like when you

37:33

have a fixed knowledge area and whether

37:38

it's um whether the respond is like

37:40

consistent with it or not so say let's

37:43

say it's like if you have an essay to

37:45

say the sky is purple and the moral J is

37:48

like oh the sky is purple then even

37:50

though like is not correct but it's

37:53

consistent with the essay so I would

37:55

still say especially like uh consistent

37:58

um but there like Global incon factual

38:01

inconsistency right is like you just

38:02

like basing only information out there

38:04

for example like um do vaccines cause

38:08

autism right so so so so you don't want

38:11

to like um so ask a question and AI has

38:14

to leverage on the knowledge it knows or

38:16

has access to to answer those questions

38:19

and sometimes the hardest part of it is

38:21

trying to determine not the

38:23

verifications but should determine what

38:25

the facts are because when you leave AI

38:28

to itself like going through the

38:29

internet and um with a lot of like

38:32

misinformation and people just like

38:33

putting for like like uh it's really

38:36

hard to differentiate um even like for

38:38

for humans right like uh it's really

38:40

hard so so so you need to make the task

38:43

easier for AI we focus on a local

38:46

factual inconsistency which means that

38:48

we provide AI with a ways or guidelines

38:52

or access to like information that's

38:54

relevant so the job AI now is like to

38:56

focus on like exract information from

38:58

that instead of trening to mean like

39:00

what is the correct what is the factual

39:03

things to do why is that why I think

39:06

like contact enhancement is like so

39:08

useful for for to to deal with uh

39:11

factual

39:13

inconsistency yeah I know that's that's

39:15

quite interesting so so I like that um

39:19

yeah kind of direction of uh in factual

39:22

contexts we need to make AI boring and

39:26

uh you know in others uh it has to be uh

39:28

creative uh so that's really interesting

39:31

and maybe also then diving into the

39:34

topic of uh rag um the interesting thing

39:37

is you know we have quite a large

39:39

audience uh and we're getting all the

39:41

spectrum of questions on very very

39:43

specialized and very Advanced also some

39:46

some more more

39:47

foundational um before diving into

39:50

questions related to rag I think it

39:52

would be uh great if you can maybe paint

39:56

a picture set some intuition of um you

39:59

know what what rag could look like in

40:02

practice in regards to prompt

40:04

engineering like what would a prompt

40:06

look like in order to suddenly make an

40:09

llm being able to talk to a database

40:12

like how can uh AI do that uh if you can

40:15

maybe just give us some of that that

40:17

intuition less architecturally and maybe

40:19

just like you know what does that look

40:20

like in a

40:22

prompt um so so

40:25

so this is uh let's go through this

40:28

right so let's say I have a query uh of

40:32

like

40:34

um tell me about the

40:37

movie Inside Out so the retri like first

40:41

query is sent to a retri systems So like

40:45

um it could just be like an API keyw

40:48

search and it's it's like going to some

40:50

documents so and it's just using example

40:54

keyword search so it identifies the

40:56

keyword here is inside that and it was

40:58

just search on the document just in

41:00

which one contained inside out and that

41:02

it just like paste a document into like

41:05

here owns the documents for inside out

41:07

and here's a query Tell me about the

41:10

movie Inside Out and the model takes

41:11

that as the context to generate a

41:14

response and say like inside out is a

41:16

movie that pix up blah blah blah and

41:17

send it back to users

41:21

yeahh yeah and I think I think then the

41:23

follow-up question to that um uh which

41:26

it answered uh it had asked basically

41:28

did rag actually replace fine

41:32

tuning so they addresses uh they address

41:36

different problems so maybe let me see

41:39

again actually have another slide so I

41:41

just give a talk on um on on on Tuesday

41:45

um that actually covering this

41:49

like yeah let me see um so I see like um

41:55

so so I do think if I to approach um to

42:00

picking the solutions like sometime it

42:02

can be solution Focus right oriented

42:04

it's like oh here's a solution and want

42:06

you use this how do I use this another

42:08

approach is like problem oriented it's

42:10

like okay what problems do I have and

42:13

then I see like what solution is best

42:14

suited for that problem so I would

42:16

categorize them as like

42:18

um usually if you want to enhance con um

42:23

if you want to fix information based

42:25

failures when the respond usually like

42:28

um oated so here's example of like an uh

42:32

an oated response right like if it's ask

42:35

like how

42:36

many booms studio and Booms has to Swift

42:39

released and if it respond like 10 11

42:42

it's likely because the cut of dat was

42:45

before the DAT the last un boom was

42:47

released or you can also see that's like

42:50

you can also check yourself like

42:51

sometime if find models are more likely

42:53

to hallucinate when um you ask them

42:56

about very very Niche information just

42:59

unlikely to appear in the training data

43:01

or um or just um like very rare like it

43:06

doesn't appear a lot so so you can do

43:08

that like by so that's when if R is very

43:11

useful but then if you see the problem

43:13

with um so there a lot of research

43:15

showing that like rack and Ral can work

43:19

quite well for like a lot of bench mark

43:21

so you can see a base model plus rack uh

43:24

on this paper and it's very interesting

43:26

paper that uh base model plus rack

43:28

usually works pretty well uh but

43:30

sometimes using rack on F2 model can

43:34

also like get better performance so I

43:36

think about 40% of the time rack on fun

43:39

model so it does pretty well uh in

43:41

almost no cases where phun is like

43:43

perform better on on ML um on this

43:47

Benchmark but but this not necessarily

43:49

the case right because they use some use

43:51

cases um where fune can work better and

43:55

usually they are related to to the use

43:58

case of um Behavior based failures so so

44:02

you can notice something like uh when

44:04

the appos are incorrect relevant or when

44:06

the appos don't follow the requested

44:07

format and there are a lot of ways to

44:10

like enforce this Behavior like

44:12

prompting vators uh use a better model

44:15

like when we talk about and F tuning so

44:17

so yeah so I was to say

44:19

like think rack replace F tuning is just

44:22

more of like what problems do you have

44:24

and there different solutions and

44:26

whether rack of funing works best for

44:29

for the problems is depend it's like

44:31

really

44:32

depends that's great yeah and I think it

44:35

was it was also fantastic that you were

44:36

able to pull a new set of slides uh to

44:39

do a deep dive I think uh uh the content

44:43

uh you will always have uh uh yeah a

44:45

reference U for for the for the audience

44:47

we also shared the link to your website

44:49

where they can find it um so then maybe

44:52

let's also dive a little bit deeper um

44:55

so especially now with uh this agentic

44:59

architectures um yeah how do we keep

45:02

this uh complex systems that now

45:06

interact with different like

45:09

databases uh uh uh different Services

45:12

how do we keep them safe and and

45:17

secure um so I think there a questions

45:21

of um putting God rail so so like at the

45:24

same time like when we talk about like

45:26

how to get some s here right so we need

45:27

to understand what the risk are that

45:30

where in the systems uh that things are

45:32

fall are filling and then we can put gut

45:34

rails around where the potential

45:36

failures can happen so so one Ty of

45:39

failures is um I say it's like um um

45:44

depending what you okay so like first

45:46

one kind failure is when using an

45:48

external API and you entally lick those

45:52

sensitive data to external API so some

45:55

company do that by just like not allow

45:58

just not using external API so you don't

46:01

want to S more like open AI anthropic or

46:03

Google on together uh some other

46:06

companies maybe try to say like put a um

46:08

input guard rail so you can detect like

46:11

sensitive information and then try to

46:13

block those requests from going out uh

46:16

or some model yeah so some companies

46:18

might try to host everything inhouse

46:20

instead uh another kind of like Risk is

46:24

uh let's say um I would say it's like um

46:27

also ra input is like prom injections or

46:30

like prom um so prom injections sometime

46:34

like sometime not own PR injection can

46:36

be uh harmful like it could be like

46:39

annoying like some people trying to get

46:40

AI to like say bad things right but can

46:43

be like more harmful when you give AI

46:46

models access to tools uh to so so you

46:49

can execute things um so in that case

46:52

you you have to like so do it from a

46:54

different approaches so of different

46:55

angle so first you have like makes a

46:58

system not allows a system to autoally

47:00

execute dangerous query so say if you

47:03

have give access to S database don't

47:05

give it right access right like you can

47:07

you can you can give it only read only

47:09

access like it can select but it cannot

47:11

like delete or remove update anything

47:14

right um so so so like the system

47:17

doesn't execute those uh similarly if

47:20

you give ACC to email then you can just

47:22

give it read access and not like allow

47:24

it to like send an email or like did

47:27

didn't email so so I think it's very

47:28

important to make the system um not

47:32

vulnerable to you in the first place uh

47:35

the second things is uh you can also

47:37

like uh detect queries so you can see

47:39

like um some people have um tried to

47:42

like do intent classifications of your

47:44

query to see like what is what is user

47:47

really trying to do and if it's like a

47:50

bad intention you don't serve it I think

47:52

one one really good practice I do

47:54

recommend a lot of companies do is like

47:56

the find schope questions so like uh so

48:00

somebody told me recently it's like they

48:02

went to a coffee shop when they have a

48:04

bot to ask about coffee and the boss

48:06

just refuse to answer any questions

48:08

unrelated not related to coffee and I

48:11

think it's good because I even think

48:12

okay if if you do create a chat bu for

48:15

Enterprise right to talk about like car

48:17

insurance there's no reason why it

48:19

should also like respond to question

48:20

about immigrations or like vaccines so

48:23

it's very clearly like detect like

48:25

whether the query is in scope for it or

48:27

not or just reject it uh so any other

48:31

queries could be like the model can

48:33

reveal sensitive informations so I think

48:36

one thing is like let's

48:38

say um let's say like you give model

48:40

access is a database and then a mod

48:43

retrieve the the information from the

48:45

data so maybe maybe Alexandro right you

48:48

created a request but then youra and my

48:51

data in the same database and the model

48:53

can retrieve not just your data but also

48:54

my data and so can reval it to users uh

48:58

so that could be like a race uh another

49:00

way could happen is because a model was

49:02

trained on sensitive data and then as a

49:05

inference time it was like uh reate the

49:08

information from the train data so the

49:11

solution to this like in the first one

49:13

um you can build a model predi so there

49:16

a lot of like pii information detector

49:19

so try should detect whether the pii was

49:21

like included in the in the response and

49:24

block it and the solution is like not

49:26

train model on sensitive data in the

49:29

first place but it's harder to enforce

49:31

because you a lot of us are not building

49:33

our own models we have to depend on

49:35

model providers so I do think this like

49:38

it's very important to have more

49:41

transparency are training data but

49:44

unfortunately uh companies just becoming

49:46

more crative on the train data because

49:48

training data is the competive advantage

49:51

so companies don't want to reveal too

49:53

much so there like uh they say we don't

49:56

want computer you know but at the same

49:58

time it can protect them from scrutiny

50:00

and like uh legislature so yeah so so

50:02

the whole Space is quite uh

50:05

scary yeah no that is really great I

50:08

mean I think I think the the topic of of

50:10

um yeah safety security guard rails um

50:15

yeah could be its own its own Deep dive

50:18

um it does feel like we're back in the

50:20

90s with SQL injections uh yeah and um I

50:24

do remember you in in one of your your

50:27

uh uh posts SL documents uh you had a

50:31

very nice um uh uh uh overarching

50:35

architectural blueprint of basically the

50:38

archetypal um design for for for um yeah

50:43

agentic systems which has basically the

50:44

input guard rails the output guard rail

50:46

so yeah I can put it it's like longer

50:50

talk I don't really want um it's like

50:53

it's it's pretty complex system so yeah

50:55

so like you have like so B Contex

50:57

construction here you have G rails so

50:59

you so basically you want to put guard

51:01

rails everywhere the system can fail

51:04

like if it's like fail as a input do it

51:06

if it fail as an output do it if it fail

51:08

as my retrieval do it uh so yeah so got

51:12

is just maybe like um I was like fancier

51:16

try

51:17

cash yeah yeah yeah that's that's super

51:20

interesting okay so so let's um yeah

51:24

maybe ask one or two more questions

51:26

before we wrap up but one question that

51:28

I think uh has come up or uh multiple

51:31

questions that we can uh I guess reduce

51:33

to to a single one um we have different

51:36

people with different backgrounds um I

51:39

think uh you know some coming in you

51:41

know for you know with uh in PhD or or

51:45

um yeah basically uh uh still University

51:47

others uh in software engineering uh

51:50

roles and others in machine learning

51:51

engineering roles um so yeah what career

51:55

advice would you give for someone um

51:58

yeah that wants to get uh into the field

52:01

or basically into a job of not just

52:04

machine learning engineering but

52:06

specifically AI engineering

52:09

yeah uh B career advice um I feel like

52:14

career is something very personal um so

52:18

I would probably try to

52:20

understand um yeah I think it's really

52:22

hard to give like blanketed um advice

52:25

that works for everyone because sometime

52:27

it works for me it works for my friends

52:30

but not working for other people and the

52:31

same time I try advice work on my

52:33

friends just doesn't work for me um so

52:37

so I would say this like

52:40

um yeah I don't think one most important

52:43

skill is like the learning she learned

52:45

and then problem solving um

52:48

like was probably spend less time like I

52:53

don't know uh so so this some heris uh

52:56

one I forgot the name of it but

52:58

basically idea is that um how long

53:01

something will last in the future is the

53:04

same amount of time that it has last in

53:06

the past right so so like uh so say like

53:09

if something has been around for 10

53:10

years and we can use your to estimate

53:13

that be last for another 10 years so so

53:16

relationship it's a same relationship

53:17

right you just oh yeah L effect thank

53:19

you John um so the first thing is like

53:22

um like if you in relationship with

53:23

someone like you just start dating last

53:25

week it's unlikely you know like you

53:26

don't want to plan something like a year

53:28

in the future but you've been with

53:29

someone for like four years maybe it's

53:31

reasonable to to like plan something a

53:33

year in the future so the same thing

53:35

it's like a lot of new things that comes

53:37

out a lot of them and probably some of

53:40

them will be very important but then if

53:42

they really important just like wait a

53:44

little bit you know like and see if

53:45

that's a case you don't immediately have

53:47

to jump on or like if you jump on do ask

53:50

questions because a lot of marketing

53:52

space so i' been like actually a

53:54

compiling but I call like marketing

53:56

speech uh translator because some

53:58

sometimes we come like oh here's this

54:00

fancy technique to do this and that

54:01

alternative is something very simple but

54:03

dress up in something like let's try

54:04

asking a lot of questions uh also like

54:07

when you read something like who is

54:08

publishing it because sometime um for

54:11

example like you see like some startup

54:13

coming out with like crazy exciting

54:15

things and it realizes like everyone who

54:17

is pushing for it as investor in the

54:19

company so like understanding who are

54:21

like propagating the information so

54:24

basically just ask a lot of questions

54:27

well that that is great that is great

54:29

and and and I really like that point of

54:31

um yeah also the fundamentals and the

54:33

things that will last I mean you

54:35

highlighted that that rag Concepts like

54:37

that you know have been you know uh

54:39

there for for a long time and that

54:41

proves it's uh resilience and you do

54:44

have also a very nice blog post uh on

54:46

measuring uh your career progression uh

54:49

I think you do delve into some some

54:50

similar areas so indeed uh I think uh uh

54:55

uh uh unfortunately

54:56

uh we are uh now out of time uh we have

55:01

had a very very great discussion uh uh

55:03

with a lot of really great questions um

55:06

so I do really want to give a massive

55:08

massive thanks to

55:10

chip great presentation yeah uh sorry uh

55:14

please uh go ahead yeah chip oh no thank

55:18

you so much everyone uh I really

55:20

appreciate your time and I'm sorry if I

55:22

spoke too fast and I'm sorry I couldn't

55:24

get your own the questions

55:26

uh I'm on social media um so yeah please

55:31

um say hi um I'm pretty bad social media

55:35

because it fight incredibly distracting

55:37

uh so so every has this fight of like I

55:40

know that like something like LinkedIn

55:41

and Twitter is could be good could have

55:43

me useful information but like it's like

55:45

trying to balance the cost and benefits

55:48

you know sometime yes it can give me

55:49

like 10% information but like 90% of

55:52

which is like distractions so anyway uh

55:55

and so yeah I I don't know how you

55:57

answer that question but yeah I hope I

55:59

see you all at some point in the future

56:02

uh and this um if you're rich child uh I

56:06

would try my best to respond but

56:08

sometime I cannot get to like one of

56:09

that yeah yeah no and this is this is

56:13

really great uh so Chip I think uh yeah

56:15

the whole presentation was great um the

56:18

questions also were great um and one

56:21

reminder that indeed the talk will be uh

56:24

recorded and is available uh in a few

56:26

days uh in learning. acm.org so you will

56:30

be able to replay it multiple times in

56:33

case you missed anything so nothing to

56:35

worry uh as all this information is

56:38

going to be available there as well as

56:40

the slide deck which was also shared um

56:42

please uh also a reminder to fill out

56:45

our quick survey uh you can suggest uh

56:48

future topics or speakers so please do

56:51

take a moment to to fill uh those up as

56:53

it helps us to make it ever better

56:56

um so on behalf of ACM uh of Chip and

57:00

myself thanks again for joining and I

57:02

hope you join us uh once again in the

57:04

future and with this uh we conclude our

57:07

talk and we wish you a great week ahead

57:11

thank you

57:12

all questions can I get the questions

57:15

that I couldn't get you so I try to

57:17

address them

Interactive Summary

The video features a presentation by Chip Huyen on the transition from Machine Learning (ML) engineering to AI engineering. Huyen explains that AI engineering involves building applications using foundation models. She covers essential topics such as evaluation challenges, context enhancement techniques like Retrieval-Augmented Generation (RAG), and the necessity of guardrails in agentic architectures. The discussion highlights that while the tools and hype evolve rapidly, many core concepts are foundational and resilient. She advises focusing on problem-solving over jumping onto the latest hyped technology and emphasizes the importance of evaluating applications based on specific failure modes.

Suggested questions

4 ready-made prompts