HomeVideos

ByteCast Ep88: Ricardo Baeza Yates

Now Playing

ByteCast Ep88: Ricardo Baeza Yates

Transcript

1921 segments

0:01

This is ACM Bytecast, a podcast series

0:04

from the Association for Computing

0:06

[music] Machinery, the world's largest

0:08

education and scientific computing

0:10

society.

0:11

>> [music]

0:11

>> We talk to researchers, practitioners,

0:14

and innovators who are at the [music]

0:15

intersection of computing research and

0:17

practice.

0:19

They share their experiences, the

0:20

lessons they've learned, and their own

0:22

visions for the future of computing.

0:24

I am your host, Juan Miguel de Hoyo.

0:27

>> [music]

0:30

>> At its core, search is about helping

0:32

people find what they need in a world

0:34

overflowing with information.

0:37

So, it sounds simple, but it's actually

0:39

a deeply difficult problem.

0:41

How do you understand what someone is

0:42

looking for, sift through enormous

0:44

amounts of information, and decide what

0:46

is most useful,

0:48

relevant, and trustworthy in that

0:50

moment?

0:51

And once you think about it that way,

0:53

you realize that search is pretty much

0:54

everywhere. It's in the questions we

0:56

type into a search engine, the videos

0:58

and articles recommended to us, the

1:00

products we discover, the maps that

1:02

guide us, increasingly, the AI tools we

1:04

turn to for answers.

1:06

Behind all of these everyday moments is

1:08

a deep field of research called

1:10

information retrieval,

1:12

the science of helping people find the

1:13

right information at the right time from

1:16

an overwhelming amount of data.

1:19

Few people have helped shape that field

1:20

as deeply as Ricardo Baeza-Yates.

1:23

Ricardo is a Chilean computer scientist

1:25

whose work in information retrieval

1:27

algorithms,

1:29

web web search, and data mining has

1:31

influenced how modern search and

1:33

recommendation systems are built and

1:34

understood. He is currently the search

1:37

chief scientist at you.com

1:40

and holds part-time professor

1:41

appointments at KTH Royal Institute of

1:44

Technology in Sweden,

1:46

Universitat Pompeu Fabra in Spain,

1:49

and the Universidad de Chile.

1:51

His career has moved across academia,

1:54

industry, and entrepreneurship. He

1:56

previously served as vice president of

1:57

research at Yahoo Labs,

1:59

holds 14 patents, and has co-founded

2:02

several startups in Chile and Spain,

2:03

including Theodora AI,

2:05

which focuses on mitigating

2:07

technological bias.

2:09

He has earned his engineering degree and

2:11

master's degree in computer science and

2:13

electrical engineering from the

2:14

Universidad de Chile,

2:16

and his PhD in computer science from the

2:19

University of Waterloo.

2:21

But, Ricardo's impact goes beyond

2:23

technical foundations. He has spent his

2:25

career mentoring researchers, building

2:27

institutions, and helping expand the

2:29

visibility of the Latin American

2:31

computing community around the world.

2:34

And that combination of pioneering

2:35

research and commitment to lifting

2:37

others up is what has led to his

2:38

recognition with ACM Luis Andre Barroso

2:41

Award, which honors fundamental

2:43

contributions to computing from

2:45

researchers in historically

2:46

underrepresented communities.

2:48

So, it's my honor and real pleasure,

2:50

actually, to spend time with you again,

2:52

Ricardo. It's been quite a while. So,

2:54

>> Thank you. So, it's so good to see you

2:55

again.

2:56

>> Yeah, when was the last time we saw each

2:58

other?

2:59

Was it ITU?

3:00

>> In Geneva in in the AI event that they

3:03

do every year.

3:05

>> Oh, yeah, that's true.

3:06

>> Yes.

3:07

>> Yeah. Will you be going this year?

3:09

>> No, no, I only went that year when I met

3:11

you.

3:12

>> Oh, okay.

3:14

Yeah, that's a lucky year.

3:15

>> Yes.

3:16

>> Yeah. So, before we continue, I kind of

3:18

wanted to ask you. I know I we talked a

3:20

bit about your background, but would you

3:22

mind introducing yourself and talking

3:24

about what you currently doing?

3:26

>> Yes, so no problem. So, currently I'm

3:29

sharing my time between industry and

3:31

academia, mainly industry in in in

3:33

you.com, as you mentioned, in San

3:35

Francisco, but also I I do have these

3:37

part-time commitments in several places,

3:39

and physically I have to be a few months

3:41

in a Stockholm because of my KTH

3:44

commitment that was even before I

3:45

started working with you.com, and so I'm

3:48

right now in Stockholm

3:49

doing research on responsibly AI, so

3:52

mainly with my PhD students, mainly in

3:54

Barcelona,

3:55

and my work here in KTH,

3:57

touching different problems on

3:59

responsible AI. Maybe the main one is is

4:01

evaluation, and looking at how to

4:03

evaluate harm and not success.

4:06

Because things should work, right, by

4:08

default. I mean, why we evaluate when

4:10

things work? We need to think what

4:12

happens when things doesn't work, and

4:14

how much harm that creates because

4:16

all errors are not equal. So, there are

4:18

errors that are more harmful than

4:21

others. And we are looking at that, so

4:23

different ways to minimize critical

4:25

errors, and also to try to measure. And

4:29

because of the regulation of the use of

4:30

AI needs that, for example, the

4:33

EU regulation is based on harm and risk,

4:37

but really we don't know how to measure

4:39

risk too well to get.

4:40

>> I totally agree, and I think that's a

4:42

very fair point. I think one of the

4:44

challenges we're living in now,

4:45

currently in the space, is that both

4:47

the technology for discoverability is so

4:50

pervasive. Like search is just We take

4:51

it for granted that search did not exist

4:53

like decades ago.

4:54

And so, the work that you were doing in

4:56

that space is really vital, but at the

4:57

same time, that space is also evolving

4:59

with in conjunction with a lot of

5:00

different technologies, like machine

5:02

learning, and how that impacts our

5:04

ability to disseminate and understand

5:07

information. Before we go in further, I

5:09

think one thing I really want to ask is

5:11

that sometimes when we have audiences

5:13

listening to this podcast, they may not

5:15

necessarily completely understand all

5:17

the technology.

5:18

And so, I was very curious if you could

5:20

explain what makes both either search or

5:22

responsible AI difficult as a topic to

5:25

kind of tackle.

5:27

>> In some sense, I'm combining the work of

5:29

those two fields because I did the com

5:31

mainly does

5:32

search APIs for AI AI agents.

5:35

And what do you want? If you If you want

5:38

to succeed with a AI agents, or what is

5:40

called the agentic AI,

5:42

is to have relevant and true

5:44

information.

5:46

And this is not easy if you use a

5:47

language model that will predict

5:50

everything.

5:51

And in these predictions, then some

5:53

mistakes can occur. I don't like this

5:55

word hallucination, because really

5:57

they're not hallucinating, they're more

5:59

like making mistakes.

6:00

Sometimes mistakes are creative, and

6:02

these could be interesting mistakes,

6:03

depending on the context.

6:05

But if you make a mistake in a legal

6:07

case, or in a medical case, or even in

6:10

other settings related to people,

6:12

this could be really harmful. And then

6:14

what you need to combine responsibly AI,

6:17

for example, to find the right

6:18

information. And that's why we're

6:19

building the best possible

6:22

search APIs for agents to make sure that

6:24

they have the correct information to

6:26

start with. Of course,

6:28

on the way down, mistakes can happen.

6:31

But at least starting with true

6:33

information. And today,

6:35

it's very hard to check anything that

6:38

the chatbot will give to you, because

6:39

you are not an expert. So if it's well

6:41

written,

6:42

and looks like correct, while it's not

6:43

correct. And very few people will find

6:45

the mistakes.

6:47

And this happening

6:48

all over again in in many contexts, in

6:50

many news. And many people, of course,

6:52

don't mention it.

6:53

But I know cases every day about this

6:57

problem.

6:58

>> I can totally see that happening. And

6:59

we're seeing some of that also appear in

7:01

the news, right? Like there's very

7:02

different cases, just going back even

7:05

years from now, since even when we first

7:06

met,

7:07

where you're seeing how we're beginning

7:10

to really grapple with

7:12

the impact of the technology. Is there

7:14

something that you think about search or

7:16

AI that we as engineers would know

7:18

about, but that the public usually

7:20

misunderstands?

7:21

>> I think even some computer scientists

7:23

don't understand very well the

7:24

limitations of data, machine learning,

7:26

and even us. So because I think the most

7:29

important problem is not the technology,

7:31

it's us, how we use it.

7:33

So for example,

7:34

if we don't understand that really the

7:36

system doesn't understand anything, and

7:38

it's only a very good prediction,

7:40

we run into trouble. So the first

7:42

problem I think that's true for everyone

7:45

is that we humanize this technology. We

7:47

talk about thinking, reasoning, seeing,

7:51

writing, reading. I know for all these

7:53

things to work, you need to have real

7:55

understanding.

7:57

So, first we should use the right

7:59

language. So, we should use, okay, this

8:01

system process data, generate data,

8:04

but that's it. Whenever you have a

8:06

system that looks at cat image and

8:08

transform the cat image to the word cat,

8:11

it's only changing the representation.

8:13

It's not understanding what a cat is

8:16

like we do.

8:17

So, I think this is even within some

8:19

computer scientists this is not well

8:22

seen because they use words that they

8:24

shouldn't use to explain how this system

8:26

works. So, we should be careful

8:29

on that. Weizenbaum, like 30 years ago,

8:32

the person that designed the first

8:33

chatbot, Eliza, said,

8:35

"We should never confuse computers with

8:38

humans."

8:39

But today we're doing every day.

8:41

>> Right. And it is kind of a blurred line

8:42

nowadays, right? Even with search, for

8:44

example, let's go search. Search as a

8:46

field has gone from keyword matching and

8:48

ranked links to semantic search, neural

8:50

systems, rags, or retrieval augmented

8:52

generation, and AI assistants that

8:55

synthesize answers. And do you consider

8:58

that

8:59

are we still doing some sort of

9:01

information retrieval when we're giving

9:02

synthesized answers by AI? Are those

9:04

valid information retrievals, you think?

9:07

>> For me, as an information retrieval

9:08

researcher,

9:10

I would say no. I know that many people

9:12

it's moving away of that standard search

9:14

engines that are using AI summaries

9:16

because the AI summaries many times are

9:18

wrong.

9:19

And some people look only at that. Well,

9:21

some people even use chatbots to instead

9:24

of search.

9:25

And there many consequences there. One

9:27

is that they're not really searching,

9:29

they're predicting information. That's

9:31

one problem. But also, for example, they

9:33

are using 10 times more energy than

9:35

searching.

9:36

So, they're wasting energy. So, that's

9:38

another I think that's another concern I

9:40

have, the use of resources

9:42

that are are important for us. And there

9:44

we need to be fair, it's not only

9:46

chatbots that we're using a lot of

9:47

energy with TikTok, Instagram, and many

9:50

other apps that are

9:51

helping us to waste time,

9:53

but also they're wasting energy. So, I

9:56

think many people is not conscious of

9:57

that, and somebody is paying that.

9:59

Probably we're paying part of it, but I

10:01

think we are not paying part of all our

10:04

waste of time that we do every day.

10:07

So, I think we need to be very careful

10:08

on this balance on how to use these

10:10

tools for the right thing. For example,

10:12

I believe this can be good

10:15

tools to translate information, to help

10:18

you in other languages, to summarize

10:20

information, to start information,

10:23

but if we use it for just writing,

10:26

basically we have all this problem of

10:28

cognitive delegation,

10:30

cognitive offloading. Basically, if we

10:33

do it, it's not too much trouble because

10:35

we have learned how to do it.

10:37

But for young people that they haven't

10:39

learned yet to how to do it, I wonder

10:41

what will happen in 10 more years when

10:43

people cannot write anything,

10:45

not because they don't know how to

10:46

write, it's because they don't know how

10:47

to think, which is this is the important

10:49

part.

10:50

They will be like I would call it

10:51

cognitive zombies.

10:53

Yeah, this is not evolution, it's

10:54

involution. And I wrote a paper about

10:56

that about the how we are not evolving

10:59

in search because we're using too much

11:01

AI for the wrong things. Now, if we

11:03

combine AI the right way with search, we

11:05

can get

11:07

very good interesting solutions. Like

11:08

for example, if you have relevant

11:10

information and you want to summarize

11:12

it, then language model can create the

11:14

best answer possible,

11:16

but you know that it's all true because

11:17

it started from true information.

11:20

But when you start predicting

11:21

information, then you run into troubles.

11:23

And this is something that it's very

11:24

hard to fix because it's part of the

11:26

architecture.

11:27

The architecture was designed to predict

11:30

the most probable text, was not designed

11:33

to look at knowledge database and give

11:36

you the real fact.

11:39

>> So, one of the things that's really

11:40

interesting now, obviously, I think is

11:42

retrieval augmented generation or rags,

11:43

right? So, for folks who do not know

11:46

what a rag is, it's to put it simply,

11:48

like high-level,

11:49

it means that an AI looks up the

11:50

relevant information first and then uses

11:53

that information to give a more accurate

11:55

answer instead of relying what it

11:56

already {quote} "knows" based off of the

11:58

model. And to me, personally, it does

12:00

just connect sort of connect modern AI

12:02

systems back to information retrieval.

12:05

Would that kind of fit what you're

12:06

envisioning or would you see it in a

12:08

more different direction?

12:09

>> Yes, but you know that even using rag,

12:12

this system sometimes in in

12:14

make up answers.

12:16

So, this is not enough. This improves

12:19

the so mistakes are less probable,

12:21

but because of this prediction approach,

12:24

this mistake can even happen. So, there

12:26

are some weird results like show that if

12:29

you put some random results on the rack,

12:32

the result may improve. So, this is like

12:34

counterintuitive. You if you put things

12:36

that are not interesting, then maybe the

12:38

the signal-to-noise ratio improves and

12:41

then you get a better answer. So,

12:43

we still don't understand how system

12:46

that have not 1 million, but billions of

12:50

parameters really

12:51

work. We can build them.

12:53

We know what is the architecture, but

12:55

really we cannot imagine in our minds

12:57

how these things interact.

12:59

When you're choosing 1 billion formulas

13:01

to compute something, right? No, it's

13:03

very very hard.

13:04

>> Absolutely. And I think as search

13:06

becomes more semantic and more

13:07

personalized and AI generated, those

13:09

stakes are going to just get higher. To

13:11

your point, like these systems don't

13:12

just find information, they can shape

13:13

what people see and trust and believe.

13:16

And that makes a lot of the work that

13:18

you're doing with biases and search and

13:20

recommender systems and responsible AI

13:22

especially important now.

13:23

>> Yeah, one big problem that really What

13:25

me is not only the biases, is the

13:27

cultural colonization behind the

13:28

systems.

13:29

So, my standard example, if I ask you

13:32

how many continents there are?

13:34

>> Yeah. Are you asking me?

13:35

>> How many continents there are for you?

13:36

>> Uh I'm pretty sure there's seven, but is

13:38

there a new one I didn't know about?

13:40

>> [laughter]

13:40

>> No, no, but but you study in the US,

13:42

right?

13:43

>> Right.

13:44

>> That's why for you is seven.

13:45

>> Yeah.

13:46

>> But in the European tradition is only

13:48

six because America is only one.

13:50

>> Oh, okay.

13:51

>> The American culture is already

13:52

colonizing the rest of the world because

13:54

that's why for example, the Olympic flag

13:56

has only five circles. And Antarctica is

13:58

not there because it doesn't compete in

14:00

the Olympiads.

14:01

But for the US and Canada and many other

14:04

countries now in Latin America for

14:05

example, it's not seven because North

14:07

and South America is two different

14:08

continents. So, that

14:10

that's a simple example of cultural

14:12

colonization, but there are more

14:13

complicated examples where you don't see

14:15

the all the possible points of view.

14:18

If we want to give true answers, we need

14:20

to give more than

14:22

all the possible points of view that

14:23

that are valid. Some some topics they're

14:25

not about

14:26

truth or not. They are about beliefs,

14:28

right? And for example, if people

14:30

believe we should control guns or not,

14:32

you should have both views. And so on.

14:34

And that's happened with many many

14:36

problems.

14:37

But you will not get all possible points

14:38

of views because you are trained with

14:41

biased data sets. And also you have

14:44

barriers against maybe some thoughts or

14:48

some beliefs that you think are wrong.

14:50

But maybe another culture don't think

14:52

the same. We are getting this

14:54

colonization that the next digital

14:56

issue.

14:57

Not only that, I think it's together

14:59

with this is that the best language

15:02

model supports around 200 languages.

15:04

But we have more than 7,000 languages

15:06

alive.

15:07

They did some analysis and 10% of the

15:09

people in the world don't speak

15:12

one of these supported languages.

15:14

Now, if you add that all the people that

15:15

don't have internet,

15:17

all the young people that cannot use a

15:19

computer or a device, all the old people

15:22

that will never use it

15:23

and if you do the right overlaps, you

15:26

get between 50 and 55% of the people

15:29

that cannot use these technologies. So,

15:31

when people say we are democratizing AI,

15:33

I really hate it because we are not

15:35

democratizing AI, we are increasing the

15:37

digital gap and much faster than before

15:40

between people that cannot use this

15:43

technology

15:44

and people that can use it.

15:46

Now,

15:47

maybe the winners, and I will try to be

15:49

like provocative here. Maybe the winners

15:51

will be those people because those

15:52

people will keep thinking.

15:55

They also have more privacy.

15:57

So, maybe they are the the poor people

15:59

today, but if these things don't go

16:01

well, they are the the future of the

16:03

world.

16:04

>> In a very good perspective actually. I

16:06

think people will see our faces for as I

16:08

imagine, but we're both from the global

16:09

south. And I think we have we have had

16:12

these conversations on kind of the

16:14

discrepancies between what technology

16:16

means for people in developed countries

16:18

versus developing countries. So, what

16:20

I'm going to try and do actually in this

16:21

conversation is let's take a step back

16:23

from the tech the technology aspects for

16:25

a little bit.

16:26

>> You are from the global south, but you

16:27

are from the north hemisphere.

16:29

>> I am from I was raised in the north

16:31

hemisphere. That's true. So,

16:33

>> But,

16:34

which country is your global south?

16:36

>> I was in the I was born in the

16:37

Philippines.

16:37

>> That's what I'm saying, but this is in

16:39

the north hemisphere.

16:39

>> Is it the north hemisphere?

16:40

>> Yes, it is.

16:41

>> Obviously, this is a good example of

16:42

where the technology bias comes in.

16:47

>> Yeah, but this is I say because this is

16:49

the just a convenient mention to talk

16:50

about the global south because most of

16:52

developing countries are in the south,

16:53

but it even includes Mongolia, which is

16:55

north of Philippines.

16:56

>> [laughter]

16:56

>> Yeah, that's absolutely

16:58

>> make much sense for me because I love

16:59

geography.

17:01

So, that's

17:02

all yeah. So,

17:03

but remember the equator goes through

17:05

Indonesia, which is south of

17:06

Philippines.

17:07

>> You know, they don't Okay, so that's

17:09

definitely a bias for sure cuz they told

17:11

me, "Oh, you know, we're so close to the

17:12

equator." But, you're absolutely right.

17:14

>> [laughter]

17:15

>> Let's amend that and you're you're

17:16

absolutely correct. I think the language

17:18

there is, to your point as well, like

17:19

just the language that we've used global

17:21

south is kind of a misnomer to some

17:22

degrees when we talk about developing

17:24

countries. But, I really want to step

17:25

back and talk about your origins. It

17:28

informs kind of a lot of your opinions.

17:30

And to understand why those opinions

17:32

were exist, like I don't understand

17:34

correctly, you started out in Chile and

17:36

then moved into computer science.

17:38

Chile's not necessarily at that point in

17:39

time, if I understand correctly, not one

17:41

of the traditional centers for

17:42

computing. So, I was just kind of

17:44

curious what incentivized you to get

17:46

into computer science, actually?

17:48

>> That's an interesting question. I think

17:50

I never thought about it.

17:52

>> Really?

17:52

>> Because I I never planned things. So, so

17:54

I take opportunities.

17:56

For example, for me, I think there is

17:58

another reason that maybe was the main

18:01

driver of my career is that I learned to

18:04

read and write very early by my

18:06

grandfather, my maternal grandfather.

18:08

So, that opened my mind to the world.

18:11

And then I had a great amazing

18:13

British teacher in primary school that

18:16

had a course called general knowledge.

18:18

Because she was not a teacher, but she

18:20

wanted to teach. So, she teach general

18:22

knowledge, and this opens even more the

18:24

world that general knowledge includes

18:27

many things. So, this was like a very

18:29

interesting class for me, like you will

18:30

not know what you will learn every week.

18:33

And this is nice.

18:34

And then

18:35

I guess when you have very good teachers

18:37

of I think the teachers help a lot on

18:40

the side what you will study. So, I

18:42

think a person can do many things, but

18:44

depending on the best teachers,

18:46

they will drive you to there to that

18:48

path. So, I had very good teachers in

18:50

math

18:52

and physics in high school. So, I

18:54

decided to be an astronomer or an

18:56

engineer.

18:57

So, I started engineering.

18:59

I had never seen a computer computer

19:02

when I entered university.

19:04

And then I had to do the first course on

19:06

on programming. I discovered that I love

19:09

the logic behind algorithms and really

19:11

the combination of math and logic

19:14

captured me. I was studying electrical

19:16

engineering. I did finish that.

19:18

But really I realized that what I really

19:20

liked was computer science, and then I

19:22

had an amazing professor of algorithms.

19:24

Maybe that's the reason I started

19:27

computer science. So I did a master with

19:28

him, Ricardo Poblete. I think he was one

19:30

of my best teachers I had.

19:32

Because he did his PhD at Waterloo, and

19:35

other people in Chile did it for

19:37

historical reasons.

19:38

I went there, got a scholarship. It was

19:41

not easy at that time to get a

19:42

scholarship, especially from Chile.

19:45

And I found also a great supervisor,

19:48

Gaston Gonnet,

19:49

that even

19:51

made me love more algorithms.

19:53

And they were working on the Oxford

19:54

English Dictionary project. So basically

19:56

at that time

19:58

that was the one of the largest files.

20:00

It was 570 megabytes in 1986.

20:04

So that was a very large file.

20:06

It didn't fit in one single CD-ROM. At

20:08

that time it fit in four CD-ROMs. Later

20:10

the CD-ROMs improved, and then you could

20:12

fit it in one. But the problem was how

20:14

to search that, how to search that. And

20:16

I started with a sequential search

20:18

algorithms in my PhD thesis. And then

20:21

naturally when the web arrived,

20:23

I finished my PhD in '89. So when the

20:26

web came in '91, '92, '93,

20:29

although the ideas from 1989,

20:32

naturally I said, "Okay, I can apply all

20:33

all these search skills to

20:36

to the web." And then I passed to do

20:38

information retrieval, which is

20:39

basically searching algorithm but using

20:41

an index, like a inverted index. That's

20:43

the classical approach. And then I

20:45

started to work on on that.

20:47

With a Brazilian friend, which was a

20:49

friend of Andre Luis Barroso, because he

20:51

worked at Google, too. And Thierry

20:53

Ribeyre decided one day to do something

20:55

crazy. Why we don't do a good book on

20:57

information retrieval, because it seems

20:58

that there's no modern book.

21:01

And he said, "Yes." And that became in

21:03

'99 the

21:04

this modern information retrieval, the

21:06

most cited book on the topic. So,

21:08

so that's a not a standard thing that

21:10

two American Americans

21:12

write the most cited textbook in a topic

21:14

like this. This is not true in I think

21:16

in any other topic. So,

21:17

so sometimes for crazy ideas, you get

21:20

interesting results. So, I think this

21:21

was the beginning.

21:23

And then that's maybe because of that

21:25

reason I became the VP of research of

21:27

Yahoo in 2006 when I started the lab in

21:30

Europe in Barcelona.

21:33

And after 10 years, Yahoo was sold to

21:36

Verizon, so I left that.

21:38

I became CTO of Intent, which was

21:40

another semantic search company.

21:43

Then I went to be director of research

21:45

of the Institute for Experiential AI of

21:48

Northeastern University.

21:50

Interestingly, my manager there was

21:51

Usama Fayyad, which was the creator of

21:53

Yahoo Labs.

21:55

And he was my skip manager at Yahoo, so

21:57

And by the way, it's thanks to Prabhakar

21:59

Raghavan that that works at Google. I

22:02

I started in Yahoo, so I I should say

22:04

thanks to him, too.

22:06

And then

22:07

when the

22:08

the things didn't went very well for

22:10

responsible AI because of current

22:12

governments, not mentioning any

22:14

political thing,

22:16

I decided to move to again back to

22:18

industry and also to do my responsible

22:20

AI research in Europe, where it's more

22:22

appreciated than in the US.

22:23

>> You've had a very interesting journey, I

22:26

would say. I didn't know all of it, of

22:28

course, right? So, I just it's me also

22:29

discovering it with everyone else. And I

22:31

find it incredibly fascinating that what

22:34

you seem to say that

22:36

things are just happening clandestinely,

22:39

but I'm very curious if like the

22:42

experience you have growing outside of

22:44

where like computing was happening like

22:45

a lot of that research kind of

22:46

influenced

22:48

how you approach kind of the application

22:50

of responsible AI and how approaches

22:52

kind of where your interests are right

22:54

now.

22:54

>> As I said before, I never had any plans.

22:57

So, so this is not a planned life.

22:59

That's better than a plan because I

23:00

think plans uh yourself.

23:03

Maybe if I had a plan, I would be in

23:04

Chile.

23:06

But because of that, I'm not in Chile,

23:07

and that's good and bad for the country.

23:10

But because I can represent Chile

23:12

abroad.

23:13

But I think when you just take

23:15

opportunities, and for me any change is

23:17

an opportunity, even when looks like a

23:19

bad change, I think it's an opportunity

23:21

to redesign yourself.

23:23

When I was at Yahoo, I was worried. I

23:25

think bias has been a topic that worries

23:27

me

23:28

always.

23:29

And I've been working on bias since

23:32

now almost 20 years.

23:34

So, before this was popular. For

23:36

example, my first paper on bias, which

23:38

is about the genealogical tree of the

23:41

web, was

23:42

what's the relation between search and

23:45

the content of the web. So, how search

23:47

changes

23:48

the content of the web.

23:50

How shape the content. Because if all

23:52

people take the top pages

23:54

in a search answer to write something

23:56

else in the web,

23:57

then basically you're reinforcing

23:59

whatever is the ranking of the search

24:01

engine. And then the search engine will

24:03

later say, "Okay, I was right. These are

24:05

the best results." But basically it's a

24:07

vicious loop. You are giving to the

24:09

people what to use to produce content.

24:13

So,

24:14

and it's very hard to mitigate that bias

24:16

because it's very hard to tell people,

24:17

"No, don't look at the top 10 results.

24:19

Look at the 100 results. Look at all of

24:21

them and choose the best information."

24:23

People don't do that. They do They click

24:25

on the two three first results and then

24:27

take they take something from there.

24:29

That first bias, which is called ranking

24:31

bias,

24:32

is mitigated by all search engines. So,

24:34

if people click more on the first

24:36

position, partly because it's in the

24:38

first position, you need to mitigate

24:39

that. Otherwise, you fool yourself.

24:41

And the rich get richer and the poor get

24:43

poorer.

24:44

But the second order bias that affects

24:46

the content is much harder. And we

24:48

showed that this was happening. That

24:50

basically there was

24:52

this thing, and that was in 2008. And

24:54

almost 20 years later, I think this is

24:56

this is even worse, but we cannot

24:58

measure because at that time

25:00

maybe there were like 100 billion pages

25:02

in the web. Now there are more than a

25:04

trillion and and they're mostly dynamic

25:06

and it's very complicated. So

25:09

and now there's a lot of AI-generated

25:11

content, which is the next problem.

25:14

And something that

25:15

is not true

25:17

maybe will be amplified.

25:19

And maybe in some topics we have more

25:21

false content than true content.

25:24

And then any model trained by this data

25:27

will think that the false content is the

25:29

right content because it's the

25:30

dominating content.

25:32

And we don't even know how much of this

25:34

happening and we don't even know in

25:36

which topics that is happening. It's so

25:38

complicated because the number of fake

25:40

news and fake content today is much

25:42

larger

25:43

than before and of course it's much

25:45

larger every day.

25:47

>> No, that's absolutely true. I guess when

25:49

we're talking about generative AI

25:51

systems nowadays when we're using the

25:52

application

25:54

like even kind of what we're seeing in

25:57

terms of like authentication research

25:59

and application still needs to catch up

26:02

with some of the advances as well in

26:04

going around those effectively. Do

26:06

generative AI systems make bias easier

26:08

to detect or the output is explicit or

26:10

is it do you think it's just hard

26:11

becoming harder and harder

26:13

nowadays?

26:14

>> Uh it's a very good question. I think it

26:16

becomes harder and harder because these

26:18

systems also have biases to themselves.

26:21

For example, that paper that showed that

26:23

if you write your CV

26:26

with ChatGPT, ChatGPT will find that

26:28

better than if you write it with Gemini

26:30

or Claude.

26:31

They look like they recognize their

26:34

generation. They prefer these systems.

26:37

And of course if people is using for

26:39

example ChatGPT to evaluate people,

26:41

which is I don't think it's a good idea

26:43

because data doesn't represent well

26:45

people.

26:46

In fact, any problems where people is

26:48

involved, data is not a a good

26:51

representation of the problem.

26:53

You got a like a vicious loop again. So,

26:55

okay.

26:56

The system chose mainly people that use

26:59

the same model to write their CV, for

27:02

example.

27:03

And this is the case of CVs, but in in

27:06

many other cases is the same.

27:08

For example, there are no good tools

27:10

to detect if something was written by a

27:12

model or not.

27:14

In fact, a few years ago we wrote a

27:15

paper saying that

27:17

one regulation that exists is that any

27:20

model that is published

27:22

should come with a tool

27:24

to detect if that model did this text or

27:27

not.

27:28

And you can do it with watermarks and

27:30

digital watermarks and other things. So,

27:32

you can do that.

27:33

And this will be very important because

27:35

then you can say, "Okay, this was not

27:36

written by a human or partially written

27:37

by a human,

27:39

but was mainly AI-written." And this is

27:41

important in many contexts like exams,

27:44

like CVs, and news, and so on. So, we

27:47

don't have that.

27:48

And we need that

27:50

more and more every day.

27:52

>> No, it's true. I do think it's funny

27:53

when you look at social media and you

27:56

can totally tell when something has been

27:58

written by ChatGPT. There's a certain

28:00

cadence in terms of how something is

28:02

written, right? Like there's always the

28:04

three points, but then they say

28:05

something like, "But honestly, blah blah

28:07

blah."

28:07

>> This is true for everyone or only for

28:09

us? And we can take that. This is for me

28:12

What is What is me? Some people like

28:13

that.

28:14

>> Yeah, some people do. Like

28:16

I see it on LinkedIn as well. So, I

28:17

mean, I'm like, "Oh, okay, well." For

28:19

some people, I guess it's the the ease

28:20

of efficiency to it. But at the same

28:22

time, again, for us, there's a cognitive

28:23

aspect of us where we're like, "Okay,

28:25

maybe we can tell that's from ChatGPT,

28:27

for example." But would you say there is

28:29

one particular type of bias or some

28:31

biases in search that or in these type

28:34

of things where us as users would never

28:37

really notice?

28:38

>> Well, I think the secondary bias that I

28:40

explained people will never notice. Like

28:42

basically that people is reinforcing the

28:45

beliefs, the ranking beliefs of search

28:47

engines because of this use of content

28:50

coming from search

28:52

to create new content in the web. This

28:54

is like a loop and we prove it

28:57

18 years ago, and today it might be

28:59

worse because now will be combined with

29:01

generated content that may be not even

29:03

true.

29:04

Well, always there's some true content

29:06

that's not true coming from people,

29:08

but now generative AI

29:10

amplify that by

29:12

a factor of I would say 1,000. I don't

29:14

know what is the factor,

29:15

but it's so easy to basically write a

29:19

a blog with your name, say whatever I

29:21

want to say,

29:23

and it's so easy and people is doing

29:24

that, especially in politics. Like in

29:26

2024,

29:28

more than 100 countries had elections

29:30

that played a role. Like fake news

29:32

played a role. Like suddenly in country

29:34

that very close to my heart, people that

29:37

lie during politics and invented lies

29:39

and that lies were enough to convince

29:41

people that if they were true, was

29:44

complicated. I think that creates fear.

29:47

People are easily manipulated.

29:49

>> No, it's true and we can see how bias

29:51

and personalization can bring us it to

29:54

basically the feedback loop problem,

29:55

right? Where a system can show us

29:57

information, we can click, we watch, we

29:59

ignore or share stuff and then the

30:01

system kind of learns from that behavior

30:03

and reinforces that behavior. In some

30:05

ways that could make some

30:05

recommendations useful. Ideally, that's

30:07

kind of where we wanted to go with the

30:08

technology, but it can also narrow what

30:10

we see or amplify existing patterns.

30:13

>> And one important problem related to

30:14

that I think I forgot because I think I

30:17

I mentioned all my worries, but one

30:18

worry we haven't talked is mental

30:20

health. Because if you have a cognitive

30:22

issue,

30:23

if you have a mental health issue, but

30:24

you cannot distinguish

30:26

the reality basically from the fiction,

30:29

and then you have all these teenagers

30:31

that using chatbots are being helped to

30:33

commit suicide,

30:34

or even we already have two two mass

30:36

murderers that were helped by chatbot on

30:39

the planning stage.

30:40

And this is the beginning of something

30:42

that is much larger. We have millions of

30:44

people

30:45

basically interacting with this system

30:47

and thinking that they have a digital

30:48

friend that will know what is better for

30:50

their future. Most of recent study where

30:53

they asked many people to basically

30:55

interact 30 minutes with these tools

30:57

and then uh decide if they follow their

31:00

advice or not. And 70% of the people

31:02

decided to follow that advice.

31:04

And of course, after some period, they

31:06

were not better because the advice was

31:08

without context, right? After half an

31:10

hour, you don't know a person. You don't

31:11

know what is the right context. And and

31:13

basically, many cases that these people

31:14

were worse than before.

31:16

But they trusted this system because we

31:18

have an impostor syndrome. We say, "AI."

31:21

We think that the AI is better than us

31:23

because they have processed the whole

31:24

web. They must know more, right?

31:27

People follow this this advice. And

31:30

I guess the best case is the Adam Rain,

31:33

the teenager that died in California

31:35

last year, 16 years old. The system

31:37

mentioned six more times the word

31:39

suicide.

31:40

The system sent 300 warnings to OpenAI,

31:44

and nothing was done because there was

31:45

no human interloop.

31:47

Al- although I think the human should be

31:48

on the loop, not in the loop. He even

31:50

showed the rope

31:52

and said to the system, "Is this rope

31:54

good enough?" And the system said,

31:55

"Yes." And he hanged himself, and he

31:57

died. This is an example, but this is

32:00

the ones that we know. But how many

32:03

maybe we don't know because not

32:05

everything goes to the news, and also

32:06

not everything is recorded. It's found

32:08

by someone. Maybe there are many cases

32:10

that no one has found that are related

32:12

to a bad use of chatbot. Of course,

32:15

chatbots are not responsible. This is a

32:17

person doing this. If the system sends

32:19

300 warnings, how many warnings do you

32:21

need to send to stop the interaction?

32:24

This is my problem. You should do

32:26

something. And after this case, OpenAI

32:28

did parental controls, which they should

32:30

have done earlier because this is

32:32

obvious. I mean, you cannot have a

32:33

teenager

32:35

using this without any parental control,

32:37

any without any very specific guidelines

32:40

to avoid bad usage.

32:43

>> ACM ByteCast is available on Apple

32:46

Podcasts, [music]

32:47

Google Podcasts, Overcast, Podbean, and

32:50

Spotify. If you're enjoying this

32:52

episode, [music] please subscribe and

32:54

leave us a review on your favorite

32:55

platform.

32:59

>> That's a really interesting point, and I

33:00

guess there's a question here for me at

33:02

least where we're looking at these

33:04

systems, let's say the search systems,

33:06

let's say the machine learning

33:07

information retrieval that we're

33:08

deriving from

33:09

like querying ChatGPT or certain

33:12

systems.

33:13

With the advent of what are the changes

33:15

that we should expect in terms of

33:16

responsibility for these systems?

33:19

>> That's a very good question because I

33:20

don't know why for the AI people think

33:22

the responsibility is different from

33:24

other fields.

33:25

And the example usually I give is that

33:27

if you have a car and you have a problem

33:29

in your car,

33:31

you will never go and check who did the

33:33

piece that is failing, right? You go and

33:35

talk to your dealer and say, "Okay, you

33:37

the car was has a problem. You can just

33:39

maybe the guarantee or maybe someone

33:42

will fix it, but you talk to the maker."

33:45

The same is true for software, for AI,

33:47

for chatbots. It's the same. So,

33:49

if a chatbot

33:52

produces a

33:53

kind of harmful incident,

33:55

the company that put that on the market

33:59

or in the public, this is the

34:01

responsible one. So, it's very easy. So,

34:03

in this case, and that's why in the case

34:05

I was mentioning, there's a case against

34:08

OpenAI for allowing, for example, 300

34:10

warnings,

34:11

which is weird.

34:13

And the same it will be true for any

34:15

other system. So,

34:17

even if in in the fine print of the

34:18

terms of usage or condition, you say,

34:20

"I'm not

34:21

liable of bad use." That's not

34:23

completely true. I mean,

34:25

if you

34:26

try to forbid any bad usage and you try

34:30

to control that, yes.

34:32

But if you don't do anything to control

34:34

the bad use, this is completely

34:36

different. This is like selling guns

34:37

without any permits. Anyone can buy a

34:39

gun and you don't know how to show that

34:41

you're if you're crazy or not.

34:42

>> Well, I mean it's an interesting point

34:43

and I think that's one of the stuff.

34:44

That's probably how we met honestly in

34:47

Geneva, right? We were again at the UN.

34:49

>> talking about that, yes.

34:51

>> We were at the AI for Good Summit for

34:52

sure. And those are some of the

34:54

questions that come up in that summit.

34:56

One thing I think that is interesting to

34:58

think about, we ideally could live in a

35:01

world where people are trying to enact

35:04

good when we're developing systems and

35:06

such as NLP systems.

35:08

But do you think a system can become

35:10

biased even if it's no one is

35:12

intentionally designing it to be biased?

35:14

>> Oh, that's for sure. I mean, most of

35:16

these problems we have

35:18

no one has had the intention that these

35:20

incidents appear. I think that they

35:23

never thought that this kind of usage

35:25

may happen.

35:26

And the first one was in 2023 with a

35:28

different system, not with OpenAI.

35:31

But if you think a bit more at the

35:33

beginning, when you are in the design

35:35

phase, for example, you do some kind of

35:37

red teaming with a different kind of

35:38

stakeholders, especially users,

35:41

maybe you will come up with this kind of

35:43

usage and say, "Okay, someone some

35:44

mother a mother will say, we need

35:46

parental controls." I mean, I don't want

35:47

my son to to use this without any

35:50

control. This is for me like a like if

35:52

you're a parent and you have a gun

35:54

without it's not locked in your home

35:56

and suddenly suddenly many people have

35:59

lost their children because they have a

36:01

gun without lock at their home.

36:03

But this is for me common sense.

36:05

Well, as in Spanish, we say that common

36:07

sense is the less common of the senses.

36:10

So, this is the problem.

36:12

So, this is the problem and and this is

36:13

the same and these tools don't have any

36:15

common sense.

36:16

So, you just so they're very easy to

36:18

break. Even if you have a guardrail, you

36:20

can say, "Imagine that you're in a

36:22

fiction story and we have this context

36:25

and these people and then

36:27

the system will predict that you are not

36:29

talking about you, but you are

36:30

interacting about something else that is

36:32

completely fictional.

36:34

You will make the system will produce

36:35

information that shouldn't be produced

36:38

and then help the person to whatever

36:41

he or she wants.

36:43

I see these These are the learning.

36:44

These are very easy to manipulate

36:45

because they don't understand the world.

36:48

Today with a friend

36:49

he was saying that the he found from

36:51

ChatGPT seven free

36:54

museums that he can visit in Stockholm.

36:57

I said, "Are you sure they're free?"

36:59

And then he checked. And one was free,

37:01

but for people younger than 19. Like he

37:03

was not younger than 19.

37:05

Another one said, "Oh, it's free after

37:06

5:00 p.m." So, yeah, it's free, but it

37:08

depends on the context. And I'm sure

37:10

that after I asked him, "Okay, how many

37:12

of the seven are free?" were free,

37:14

probably maybe one or two. But yes, it

37:17

looked like they are free, but they were

37:18

on some condition that the system was

37:20

not taking account because they don't

37:21

understand the word. They're just

37:22

predicting the word and they in the

37:24

training data saw the word free.

37:27

Sorry, didn't saw.

37:29

Process the word free.

37:31

>> Yeah, sure.

37:31

>> to feel you need to understand. It may

37:34

may infer that the the museum is free,

37:37

but it's not. So many details. So,

37:39

it's so hard to solve because by design,

37:42

by architecture, these systems are built

37:44

to predict things, not to know things.

37:47

It's There's no knowledge database.

37:49

If they had a knowledge database, that

37:51

would be very different. This is the

37:52

next step, I think. And in the future we

37:54

will have knowledge databases.

37:55

And the big company that have large

37:57

knowledge databases will profit from

37:58

that. Also, they will have real

38:01

logical inference like classical AI like

38:04

predicting Okay, can you go from here to

38:08

here? So, there's some causality and

38:10

then we can infer things. Yes.

38:12

And the last one I already mentioned it,

38:14

is common sense.

38:15

But that's so hard, you know. Common

38:17

sense implies that you do the right

38:19

thing even if you don't know it.

38:21

Like this is like a I guess it's

38:24

hardwired in our brain from our genetic

38:27

history.

38:28

And if you see something wrong,

38:31

you know what to do, right? Most people.

38:33

The people that don't have common sense,

38:34

they die.

38:36

But other people that have common sense,

38:37

they save themselves because they do the

38:39

right thing. Okay, we need to run this

38:40

way.

38:41

>> I totally that's such an interesting

38:42

point and I I agree with you. I was

38:44

listening to talk that Ken Perlin gave

38:47

and he talked about virtual reality.

38:49

But the way he he described like how

38:52

children learn is completely different.

38:54

Like the semantic language, the things

38:56

that they learn in terms of

38:57

communicating with others is very

38:59

different because of the advent of the

39:00

technology existing

39:02

during the period that they were young.

39:04

For us, we did not have that technology.

39:06

So,

39:07

even like what common sense is or how we

39:10

communicate in a common sense way or a

39:12

pragmatic way is completely different

39:14

based on how the the technology has

39:15

impacted us. I'm just curious like when

39:18

you were starting out like with the

39:20

research that you were doing in

39:21

information retrieval, did you ever

39:22

imagine that we would be in a world

39:24

where we're using this technology in

39:27

this way?

39:28

>> No. No, I don't think so. Some funny

39:31

anecdote, we invited the Don Knuth, who

39:34

famous Turing Award, my my role model in

39:36

algorithms,

39:37

to give a Q&A in the Latin American

39:39

conference 2 years ago. He of course he

39:41

did it online because he doesn't travel

39:43

much. He's more than 80 years old.

39:46

All the people want to talk to him about

39:48

AI. But he didn't want to talk too much

39:49

about that. Because I guess he didn't

39:51

like it or maybe he didn't know enough.

39:53

But something he said was very

39:54

interesting because he loves algorithms.

39:56

And he said, I tried to quote him

39:58

exactly, he said, I never thought

40:00

that we will use algorithms that we

40:02

don't understand completely, right? I

40:04

mean, we know how to build them, but we

40:06

don't really don't understand how they

40:07

they what they are doing.

40:09

And it was interesting thought, right?

40:10

Like we are losing control of

40:13

we can do so amazing things that

40:18

amazing text most of the time, right?

40:21

That we are losing control of the

40:22

output. So, before I get I guess we

40:24

could predict the output. Today, we

40:26

can't. In an algorithm

40:28

classical algorithm, you you can say,

40:30

"Okay, what will be the output?" And you

40:31

can

40:32

tell the output. And then you can check

40:34

it if it's right or wrong.

40:36

Okay, the the numbers will be sorted.

40:38

And then you check that or or whatever.

40:40

You will find the

40:41

the second largest of the set or

40:43

whatever. So, the classical algorithm

40:45

But today we can't do that. And maybe

40:47

that I never thought about that. Like

40:48

basically that verification we passed

40:51

from the problem of solving problems

40:54

to the problem of verifying solutions.

40:56

So, this is a change. For example, that

40:58

is happening from for example with

41:00

emitters. Like, okay, you have all these

41:02

possible security problems, and now you

41:04

need to verify how bad it is.

41:06

How you can fix it.

41:08

If the fix that the system provides is

41:09

correct and so on.

41:11

The problem before we we had too many

41:13

problems to solve. Now AI will do that

41:15

even in math. And there were some some

41:17

interesting results in math recently.

41:19

Even Cruz wrote an interesting report

41:21

about a math problem that he proposed

41:24

that AI found solution.

41:26

But now we don't have enough people to

41:27

verify these things.

41:29

Like, okay, here there are 10,000

41:32

security issues.

41:33

Here there are

41:34

100 theorems from Erdos problems.

41:38

Check if they're correct or not. So,

41:39

it's interesting. We still need humans,

41:41

but in a different phase and in a phase

41:43

that in some sense may be more

41:44

complicated

41:46

if we put the level of knowledge higher,

41:49

the verification will be more

41:51

complicated because everything every

41:52

time we're doing more complex things.

41:54

Maybe that will improve humans. You the

41:57

the best computer scientists will not be

41:59

the ones that can program well.

42:01

Will be the one that can do a great

42:03

architecture, a great orchestration, a

42:05

great integration of agents, and mainly

42:08

a great validation

42:10

and verification of results.

42:13

>> Now, that's an interesting point. Like,

42:15

I would say if I were to step back and

42:17

I'm going to ask you this question just

42:18

because we're on this topic, like

42:20

from your perspective,

42:22

what do you think like a healthier and

42:25

more accountable system for information,

42:28

like an information feedback loop would

42:29

be? Like, let's say in search or in AI

42:32

systems.

42:33

>> I think we need to keep the classical

42:35

search. So, but next you need we need to

42:37

know when something's relevant or not.

42:40

We need to do fact checking and for that

42:41

we need to have a knowledge basis. The

42:43

problem that knowledge is also bigger

42:45

and bigger and part of this knowledge

42:46

can only be verified by very few people.

42:49

Like, for example,

42:51

things that you have done in your life,

42:52

you are the only one that can verify

42:54

that fast.

42:55

And maybe if I try to do verify that

42:58

searching the web, there are like half

42:59

of them I cannot verify because they are

43:01

not in the web. So, this is the problem

43:03

that the knowledge is becoming much

43:04

larger and also the people that can

43:06

verify the knowledge is becoming a

43:08

scarce. There's one people, two people

43:11

that can verify that. For example, about

43:13

your life, maybe you are the only one.

43:15

There's not no one else in your family

43:16

that knows everything about you and

43:18

that's true. Because they were not with

43:20

you, they didn't have the same

43:21

experience and so on. So, this is I

43:23

think the problem. I mean, how we verify

43:25

information is like we need to build

43:28

different ways to verify information and

43:30

and maybe we need to work better on some

43:32

different kind of relevance.

43:34

This is a research problem, I think. We

43:36

don't know how to build

43:38

relevance for this new world of agents

43:40

that basically will

43:43

generate something that looks true. Now,

43:46

one very easy experiment is that if you

43:49

have something that you are not sure,

43:51

you can ask the same to many different

43:53

models.

43:54

And if they all disagree, probably no

43:56

one is giving you the right answer.

43:58

But if they all agree, probably they're

44:00

giving you the right answer. The problem

44:02

is what happens is you have partial

44:04

agreement. Maybe this will be the level

44:06

of confidence, okay? How many models

44:08

agree

44:09

will be the level of confidence. But

44:10

still could be a lot of biases in

44:12

training data, could be a lot of false

44:14

information. This is a very interesting

44:16

example of a person that did the

44:17

following.

44:19

This person was a female researcher that

44:21

basically published an archive a few

44:23

articles

44:24

about an illness that didn't exist.

44:27

And the paper, if you read it, clearly

44:29

is not a good paper. Like it says that

44:32

it's fake. The paper says it's fake.

44:34

But this was

44:36

basically used by all models.

44:38

And if you ask about that illness, you

44:41

will get an answer. And the answer is

44:44

wrong because the illness doesn't exist.

44:46

But they process these papers, something

44:48

like bloxonomia,

44:50

and it's something that says that you

44:51

get the red eyes because of watching too

44:54

much a computer screen,

44:56

like from blue lights.

44:58

This doesn't exist, but

45:00

people will believe that exist. If

45:03

someone ask, "How do you call Maybe we

45:05

can do the experiment or do you later.

45:07

How do you call the illness of getting

45:09

red eyes because of blue light?" Maybe

45:11

they fixed it.

45:13

They fixed it, but how many things like

45:15

this

45:16

are already learned by these systems?

45:18

It's not the one, it's many, but you

45:20

don't know them.

45:22

Right.

45:23

>> Yeah, that's true. And that's an

45:24

interesting That's a really interesting

45:25

anecdote. I was just thinking, "Wow, I

45:26

should put some eye drops on." But

45:28

>> [laughter]

45:29

>> yeah, no, these are very interesting

45:30

points, and I think the question I have

45:32

just for anyone who's listening

45:34

actually, who might be just interested

45:36

in the field, be it AI, be it just

45:39

classical information retrieval, what

45:41

would you tell someone who wants to

45:43

build these powerful systems, but also

45:44

just wants to avoid causing harm?

45:46

>> I would say first that you need to

45:48

understand very well how these systems

45:50

work. So, you need to understand the

45:52

limitations of data, like for example,

45:54

that data doesn't represent well people,

45:56

the limitation of machine learning that

45:58

the system don't really understand, but

45:59

they predict things.

46:01

And I gave a talk at keynote at SIGMOD

46:03

and later maybe it's in YouTube about

46:07

the limitations of data machine learning

46:08

and us. And also what about uses of this

46:12

because if you learn an example that can

46:14

make harm to people, maybe you will stop

46:17

using them.

46:18

But then

46:19

if you really want to use this system

46:21

well, I will say let's use the ACM

46:23

principles for responsible AI systems.

46:26

I was one of the main co-authors of

46:27

these principles together with Gina

46:29

Matthews. For me, the first one is the

46:31

main one.

46:32

Show that your system is legitimate and

46:34

also you have all the competence

46:37

to do it. What means that the system is

46:39

legitimate? Well, it's ethically sound,

46:41

it's legal

46:43

and also it's based on science. There

46:45

are a lot of pseudo-scientific

46:47

applications. Basically, predicting

46:49

things that really are not related to

46:51

the data. And then competency means you

46:54

know well AI, you know computer science,

46:57

you know the domain expertise of the

46:58

problem. For example, if it's a health

47:01

system, you have doctors working with

47:02

you and so on. And very important, you

47:05

have the permission to do it because

47:06

many problems have been because of

47:08

people doing things that they shouldn't

47:10

supposed to do.

47:11

Because sometimes people don't think,

47:13

"Oh, this is a great idea. Let's do it."

47:15

But someone else has to authorize them

47:17

to do it.

47:18

And I think this is the case of the

47:21

child care

47:22

fraud model that was used in the

47:24

Netherlands that at the end implied that

47:27

the whole government resigned in 2021.

47:30

Probably the engineer that did that and

47:31

the Ministry of Social Affairs never

47:33

asked if the system was

47:35

okay or not. They did it and basically

47:38

didn't work.

47:39

>> No, I mean that's very good advice and

47:41

to be honest, I think useful for all of

47:43

us when we're thinking about being

47:44

responsible professionals actually in

47:46

computing.

47:47

>> Yeah, and these principles are also

47:48

translated to Spanish and soon to other

47:51

languages so more people can use them

47:53

and they can find it in the ACM website.

47:56

>> Nice. Uh it's always good to know. You

47:58

know, it's really funny cuz I I latched

47:59

onto something you said a while ago

48:01

where you said you're representing Chile

48:02

and the Latin American community outside

48:04

of the Latin American community, but to

48:07

be honest, like you have had an impact

48:09

on not just the representation side, but

48:11

also building research capacity and like

48:13

technology capacity in Latin America.

48:16

So, that includes, you know, the Center

48:17

for Web Research at the University of

48:18

Chile and mentoring, I think, about 34

48:21

PhD students, many from Latin America.

48:24

So, I wouldn't actually discount your

48:26

impact on people and like the community.

48:29

And so, I was just curious, like, just

48:31

from your perspective, like, what does

48:32

Latin America bring to the computer

48:34

science that the global field doesn't

48:36

necessarily hasn't taken a look at yet?

48:38

>> Another thing I'm proud of my 34 PhD

48:40

students is that half of them are women.

48:42

So, that's part of the

48:43

>> Oh, yeah.

48:44

>> and the gender priority affirmative

48:45

action regarding bias.

48:48

I think Latin America can help on giving

48:51

a different view of things.

48:54

Because sometimes you have this first

48:55

world view of things that basically you

48:58

think about the problem that is not the

48:59

problem of the whole earth, it's a

49:01

problem of

49:02

a given rich country.

49:05

You lose sight of the real problems.

49:07

Whenever or what I have said earlier

49:10

that you they say, "Okay, we're

49:11

democratizing." Yeah, but you are

49:12

democratizing in your country where all

49:14

people can use these tools because they

49:15

have the same language and they have

49:17

money and time to use it. But, in other

49:19

countries completely different. So, if

49:21

you go to one country where the language

49:23

is not supported and they don't have

49:25

very good internet, that means that only

49:27

maybe 20% of the people that speak

49:29

English, if you are lucky, can use these

49:31

tools.

49:32

This is something that you shouldn't

49:33

forget, that you should look at the

49:35

problem from a global point of view and

49:38

not from a local point of view. So, I

49:40

think that's important and from Latin

49:42

America, I think we we know that because

49:43

we suffered that. We basically we are

49:47

overlooked. I was the president of the

49:50

Latin American Center for studies that

49:52

basically is kind of the association of

49:54

all the departments of computer science

49:56

in almost all Latin America and that's

49:58

more than 100 institutions from Mexico

50:00

to Chile.

50:01

So, that helps you to see the

50:03

differences, to see for example, just to

50:06

say one thing. So, I think there are

50:07

more differences between Latin America

50:10

than between the US and Chile. So,

50:12

sometimes things are relative. So, the

50:14

levels of let's say number of PhDs in

50:17

computer science, the quality of the

50:19

research,

50:20

if you put that, you will find more

50:23

inequalities

50:24

in the same region rather than with the

50:26

rest of the world. And it's something

50:28

that I think in the US of Latin America

50:31

like the same thing. Everywhere is the

50:32

same. Which is not true. I mean, we are

50:35

One of the most important thing we have

50:37

on earth is the diversity of people,

50:39

diversity of cultures,

50:40

diversity of opinions, and talent is

50:44

everywhere. So, we can bring some

50:46

specific talent. We are not too many.

50:48

Like Chile is only 20 million people, so

50:50

nothing.

50:52

And even there countries that were

50:53

smaller that produce great people, let's

50:55

say Uruguay and Costa Rica.

50:58

So, we need to make use of the

50:59

diversity. I think part of the power of

51:02

humankind is the diversity. And

51:04

suddenly, some people doesn't like

51:06

diversity. And this is I think a problem

51:07

for all of us because our future depends

51:09

on diversity. Resilience depends on

51:11

diversity. Real innovation,

51:14

like good innovation. Remember,

51:15

innovation is not always positive. There

51:17

are a lot of bad innovation, but when we

51:18

use the word innovation, we have a bias

51:20

towards innovation is always good. No,

51:22

no.

51:23

Regulation is doesn't stop innovation.

51:26

Regulation stops bad innovation and we

51:28

need that. We need to have that.

51:30

So, we need to build the world

51:33

where everything is better because we

51:34

know what are the the the issues. And

51:36

for example, the financial world has a

51:38

lot of regulations and it still can

51:40

innovate. Technology I technology can be

51:42

the same. And this it will be good

51:44

because today I think the approach to

51:45

innovate is brute force. Let's use more

51:47

compute, more data, more energy.

51:51

But this is not the best way to

51:52

innovate. I love when deep seek appear

51:54

because I said, "Okay, the Chinese are

51:55

sinking

51:57

because they have some restrictions

51:59

and they have to improve too much the

52:01

same quality.

52:02

So, they use less energy,

52:04

less parameters, less compute, and the

52:07

results are not far. So, this is what we

52:09

need. We need to go back to thinking. We

52:11

need to go back quoting news in a

52:13

different way. We need to go back to

52:15

understand what is happening in the

52:16

system to do it better.

52:18

>> No, that's I think that's a great point.

52:19

There's always going to be a question of

52:21

responsibility.

52:22

>> We need to be responsible. That's why I

52:23

prefer to use a responsible AI and not

52:26

trustworthy AI

52:27

or AI safety that you don't take care of

52:30

what happens.

52:32

And of course, I don't like ethically AI

52:33

because AI is not human, so cannot be

52:35

ethical.

52:36

>> Yeah, those are interesting nuances to

52:37

bring up with you.

52:39

Yeah. Absolutely. One of the questions I

52:41

think I had like when it comes to how

52:43

you mentored students. And to be honest,

52:44

when you give any advice as well, which

52:46

I always appreciate, is kind of your

52:48

ability to elevate people to kind of be

52:50

inspired. And I was just curious if you

52:52

have any advice on how to develop like

52:55

kind of community to let students or any

52:58

young professional believe that they can

52:59

contribute at the highest international

53:00

level.

53:01

>> Well, that's a complicated question. I

53:03

think at the end means that you need to

53:05

volunteer for a lot of things. And you

53:07

know, I volunteer for ACM. I'm a

53:09

volunteer for ACM and a lot of people

53:10

give a lot of time to volunteer, either

53:12

working in committees, either doing this

53:15

podcast, many different ways, like

53:18

organizing conferences, mentoring

53:20

people. I think you need to volunteer to

53:22

build these communities and these

53:24

communities are built by networking, by

53:27

meeting people, by helping people.

53:29

I still answer all my email. I know some

53:31

people don't do that, but many times

53:33

it's a student that is asking, "Can you

53:34

share this paper? Can you help me with

53:37

this?" And so on. I try to answer all of

53:38

them because that's the only way to

53:40

improve the world, especially if this

53:42

request comes from developing countries.

53:44

But sadly, I don't see the same in many

53:46

people. So, most of the researchers are

53:49

only interested in their careers and

53:50

building only their success. I think

53:53

part of your success depends on what you

53:55

do to others. At the end, you are not

53:58

expecting any return, but life will

54:01

return something to you. I believe that.

54:02

And in my experience, that happens. Like

54:04

20 years later, because of something you

54:06

did to someone, you get something back

54:09

that is even better. So,

54:11

So, I think you need to be kind to

54:13

volunteer to for these things and to

54:14

help people, not only in computer

54:16

science, in any aspect. So, try to help

54:20

people. Look at the world. I mean, you

54:21

are not alone. Because today I feel that

54:24

most people think that they live alone

54:26

and they don't care about other people.

54:27

>> No, I think sometimes we can get that

54:29

feeling as well. I think this is

54:30

partially Personally, I think this is

54:32

why I think your receipt of the Luis

54:34

Federico Award is kind of a big deal. It

54:37

really brings together two important

54:38

parts of your career, which is a lot of

54:40

your technical contributions, but also

54:42

your broader impact in general to

54:43

community. And so,

54:46

I think one of the things I was thinking

54:48

about when I just heard about the news,

54:50

I was wondering how you felt about it,

54:51

like when you received it. What were you

54:53

feeling at that time?

54:55

>> It was a really great feeling. I never

54:57

met him, but I knew about him and I

54:59

found that he died young because of an

55:01

illness.

55:02

As I said, Bertier, my co-author of the

55:04

book, was his friend, so I knew about

55:06

him. And And in some sense, being South

55:08

American, I feel proud of receiving an

55:10

award that was given to because of a

55:13

South American researcher. So far, this

55:15

is amazing. And as soon when I will

55:17

receive it in the ACM award ceremony, I

55:20

guess I will feel more things because I

55:22

will be surrounded with with many people

55:23

that I know, that I care, and also

55:26

people that I admire, like Turing Award.

55:28

So, it will be like a nice moment in my

55:30

career.

55:31

>> I think so as well. It's an interesting

55:33

award

55:35

to receive, I'd say. And I think one of

55:37

the things I'm curious about for sure is

55:39

that because it does carry a deeper

55:41

meaning about visibility and

55:42

representation, and that's something

55:43

we've been just talking about, of

55:44

course,

55:45

is how do you hold those two meanings

55:47

together like for yourself?

55:49

>> Mhm.

55:51

Philosophical question.

55:52

>> [gasps]

55:53

>> So I think I feel lucky that in this

55:56

case I was born in South America because

55:58

otherwise

55:59

I would not be

56:01

able to win this award. And this have

56:04

other connotations. So for example, you

56:06

know, everyone is born in a random place

56:10

and could have been anywhere. And that's

56:12

why

56:13

because I'm in some sense an immigrant I

56:15

have lived in like in seven different

56:17

countries during my life. So I feel like

56:19

like I don't belong to any country

56:21

really.

56:22

I already have three passports. So so

56:25

I feel that I belong to many

56:26

communities.

56:28

One is a computer scientist. But I feel

56:30

that's part of the problem that is in

56:32

the world that we have these boundaries

56:34

that are completely made up. That

56:37

we have all the trivial issues. We have

56:39

a lot of problems that we created

56:41

because all the civilization is

56:43

basically a fiction and some

56:45

people like Harari and other people

56:47

point that well. Like like we have

56:49

invented most of the restrictions we

56:51

have. Money, property, basically

56:54

belonging to a country. And so all these

56:55

things are we are not here before.

56:58

And some of them are causing problems. I

57:00

mean, we live in a time of

57:03

growing inequality and I think that's

57:05

not sustainable.

57:07

Also, I will not talk about climate

57:09

change and other issues we have. So the

57:11

question is

57:12

is weird that we are the only animals

57:15

that have taken that path. Like we can

57:17

think

57:18

something that very few other animals

57:20

maybe are doing in the way that we are

57:21

doing.

57:23

But at the same time we are kind of

57:24

stupid because we are creating a world

57:26

that is not better for us. I will place

57:29

myself there. Like I cannot do much but

57:32

working responsibly is my small

57:34

contribution to improve the world.

57:36

>> Maybe you may seem it to think it's

57:37

small, but I think everything that we do

57:39

has meaning and value and the ability to

57:42

put that value to something that is

57:43

meaningful for others doubles that quite

57:45

a bit. We're coming towards the end by

57:47

the way. I don't know unless you want to

57:49

we can still keep talking after the

57:50

podcast of course, but I really want to

57:52

kind of just reflect on kind of your

57:55

career and your experience as a

57:57

professional. You've been doing this for

57:58

quite a while. You have seen a lot of

58:01

different things not the technology side

58:03

as we mentioned, but also the impact

58:04

side.

58:05

Do you have

58:07

any thoughts on how your mentorship has

58:09

shaped the way you think about impact?

58:11

>> As you said, I mean it's not only

58:14

about computer science. For example, I

58:16

love geography so I travel a lot and I

58:18

have seen many realities. I think

58:19

understanding the world also changes to

58:22

or seeing the contrast of India,

58:25

seeing how people live in Africa, seeing

58:28

very happy people that don't have

58:29

anything. I mean all these things show

58:31

that maybe the culture we are creating

58:34

probably is not the best for us

58:37

because we are we were not made that way

58:39

at the beginning. So evolution is

58:40

important. So I feel that maybe some

58:42

tribe in Amazonas is much more happy

58:44

than us. Probably that's the case and

58:47

they don't have many of the problems

58:49

that we have because they don't have

58:50

created those problems. Because at the

58:52

end of this problem we are created by

58:54

us. But I will go back to what I said

58:56

before. I think

58:57

how we can shape the world in just ways

59:00

that improve things.

59:02

In a small ways. Maybe at some point

59:05

maybe new generations will say stop this

59:07

and you can see it in some movements

59:09

like not

59:11

change it from how you eat

59:12

or on what things you do

59:15

or on how much exercise you do.

59:17

Maybe if more people will do this and at

59:21

some point that the political elites go

59:24

away we will have better

59:27

people governing

59:29

and we would have a better world, but

59:30

maybe I will not see that. But I think

59:32

you and I can contribute a little bit to

59:34

improve that, but I guess you have to be

59:37

a bit idealistic in spite that it's not

59:39

being very realistic because it's a very

59:41

hard work

59:42

to be against the dominant system. I

59:45

don't know about politics. I think many

59:47

people mix politics with having a better

59:50

world. I think all people wants to have

59:53

a better world. Maybe there are

59:54

different ways to do it, but we should

59:56

forget about politics when you want a

59:57

better world.

59:58

>> I think that's an interesting point. I

60:00

guess from a technological side, what

60:01

question about maybe let's say search AI

60:04

or even any technology do you think or

60:06

do you hope that the next generation

60:08

will solve or try to solve?

60:09

>> Well, I think the main question today is

60:11

where we should use AI or not. This is

60:13

the question that we should

60:15

carefully think because if we lose

60:18

cognitive skills, we are lost in some

60:20

sense in history. We need to decide

60:22

basic things about where and when to use

60:25

AI. For example,

60:27

when a kid should start using AI? For

60:30

what?

60:31

Until when?

60:33

These are basic questions that don't

60:34

have answers because we are going too

60:36

fast, but if we don't answer these

60:38

things soon, the inequalities will

60:40

appear in different ways. Maybe in some

60:42

sense in good ways because

60:44

the more developed countries are using

60:46

more AI than the developing countries

60:47

and then this could be

60:49

in their disadvantage, not in their

60:51

advantage.

60:52

>> Yeah. Now, that's a very good point.

60:54

>> Ricardo,

60:55

>> thank you so much for the conversation

60:56

by the way. Like, first off, it's really

60:58

nice to just see you and talk to you

61:00

again. It's been a while.

61:01

>> Yes.

61:01

>> I do really appreciate cuz I think

61:04

what I

61:04

sometimes forget is that there is a way

61:07

that we should see technology and you're

61:09

helping us see how search is not just a

61:11

technology, AI is not just a technology,

61:13

but it's also one of the ways that

61:15

people can make sense of the world. And

61:17

how can we improve the way we make sense

61:18

of that world, right? We've talked about

61:20

some of the foundation of information

61:21

retrieval. We talked about how search

61:23

has evolved and kind of the

61:24

responsibilities that come with building

61:25

these systems as well. And finally, I

61:28

think we got a chance to talk about

61:29

about like some of your life, some of

61:30

the stuff about mentorship, the

61:32

different things that we're seeing, and

61:33

the opportunities there are that should

61:34

be to elevating the Latin American

61:36

community as well. So, yeah. Your work

61:39

reminds us that computing is not only

61:40

about the systems we build, but also

61:43

about the people and the institutions we

61:44

help bring forward. So, yeah.

61:46

Congratulations again on the receipt of

61:48

the award. Thank you for joining the

61:50

podcast.

61:51

>> Thank you. Thank you for the opportunity

61:52

of having this podcast and maybe to

61:55

remember as you said, AI is a tool

61:57

and may help to solve our problems, but

62:00

the only people that can solve our

62:01

problems are ourselves. These are social

62:03

problems. These are not technological

62:04

problems.

62:05

And technology can help, but also can

62:07

create more problems and we need to try

62:09

to avoid that.

62:10

>> Absolutely.

62:11

>> Thank you.

62:11

>> Thank you.

62:14

>> ACM Bytecast [music] is a production of

62:15

the Association for Computing

62:17

Machinery's Practitioner Board. To learn

62:19

more about ACM [music] and its

62:20

activities, visit acm.org.

62:24

For more information about this and

62:26

other episodes, please visit our website

62:29

at learning.acm.org/bytecast. [music]

62:33

[music]

62:36

That's learning.acm.org/bytecast.

62:40

[music]

Interactive Summary

This episode of ACM Bytecast features Ricardo Baeza-Yates, a distinguished computer scientist and pioneer in information retrieval, web search, and data mining. The conversation explores the evolution of search, the complexities and risks of AI, and the critical importance of responsible AI development. Baeza-Yates shares his perspective on the cultural biases inherent in AI systems, the dangers of cognitive offloading, and the need for human oversight and ethical considerations in technological innovation. The discussion also touches on his personal journey from Chile to global academia and industry, his commitment to mentorship and diversity, and his recent recognition with the ACM Luis Andre Barroso Award.

Suggested questions

4 ready-made prompts