HomeVideos

Uber Minion: Agent-Ready Dev Platform Behind 11% of Merged PRs

Now Playing

Uber Minion: Agent-Ready Dev Platform Behind 11% of Merged PRs

Transcript

425 segments

0:06

So, behind Uber the app, which is used

0:08

by hundreds of millions of people

0:09

globally, there's an engineering

0:10

organization that has been building one

0:12

of the most mature internal AI platforms

0:14

in the industry by making deliberate and

0:16

compounding infrastructure bets over a

0:18

long period of time.

0:20

Today, Uber's background agent platform

0:22

called Minion is responsible for

0:23

generating over 11% of all merge PRs

0:26

across of Uber.

0:27

They also have a large-scale migration

0:29

management called Shepherd, which runs

0:31

migrations across a hundreds of millions

0:33

of lines of code. They have custom AI

0:35

tools for testing test generation, code

0:38

review, and more, all built internally

0:40

and compounding on top of each other.

0:42

So, my guest today is Nikhil, a software

0:44

engineer on Uber's developer platform

0:46

team. Nikhil has been at Uber since

0:48

2016. He came in as an intern, loved it,

0:50

and then just never left. So, he's been

0:52

there for a long time. And over that

0:54

time, he's been one of the leading

0:55

engineering voices behind Uber's agentic

0:57

shift. An early adopter of cloud code

0:59

inside of Uber, and he's been a key

1:01

force in building the primitives

1:02

underpin everything that you're about to

1:05

hear within Uber's development platform.

1:07

So, Nikhil, welcome and thank you so

1:08

much for being here.

1:09

Very flattering. Thanks for having me,

1:11

Lou. When it comes to building AI and

1:14

agent-first, you know, sort of AI-first

1:16

developer platform,

1:18

you already had an existing developer

1:19

platform, CI infrastructure, cloud

1:22

compute, development environments, you

1:23

know, sort of compounding investments

1:25

that you've made over those years. What

1:27

are those bets that you've placed that

1:29

have kind of taken you to where you are

1:30

today?

1:31

>> already getting to a point at Uber where

1:34

a engineer was not necessarily a

1:36

specialist. That is, they were not only

1:37

building Android code. They actually

1:39

happened to build Android, iOS, they

1:41

happened to build the backend, they

1:42

happened to build sometimes the

1:42

frontend. They actually happened to be a

1:44

fully product-focused engineer. So, if

1:47

you want to wear multiple hats, you need

1:48

to be able to switch between

1:49

environments, and environment isolation

1:51

wasn't a soft problem.

1:53

So, I actually built a

1:55

rudimentary container

1:58

for local development for Android

2:00

engineers.

2:01

And that rudimentary container, within

2:03

like a year, evolved into a prototype

2:06

called Devbox. Instead of

2:08

actually having this environment where

2:10

you had to set up and it wasn't

2:12

reproducible, now you'd have this box

2:14

and and you'd be able to just put

2:17

everything in the remote machine and

2:19

plan for a future where your engineers

2:22

are not really interfacing with the

2:25

local tooling

2:26

if they don't want to. In other words,

2:28

if if I don't want to have Android

2:30

tooling on my machine, I can use a

2:31

Devbox. If I want to build Go code, I

2:33

can use a Devbox. Yeah, exactly. So,

2:35

we're kind of, I guess, going from this

2:36

world before, right, onboarding humans.

2:40

And now this also translates over to to

2:42

agents and how do we onboard those? And

2:44

I think one question that's on a lot of

2:47

people's mind is what happens if, you

2:49

know, if we give agents their own

2:51

computers and allow them to do stuff.

2:52

And can we run them for a longer period

2:54

of time? Can we get them to do more

2:56

significant work? But that in itself is

2:58

a a new working model that we haven't

3:01

seen before, so it comes with it

3:03

challenges around security, challenges

3:05

around infrastructure, and also

3:06

wrangling large language models to kind

3:08

of do this longer horizon work. I know

3:11

that's something that Uber has also

3:13

started to experiment with. And you do

3:15

have an internal platform called Minion.

3:17

It'd be great if you could take us

3:18

through and sort of talk us, you know,

3:20

the inception of Minions, like where

3:22

that's at today and The key distinction,

3:24

I think, from Minion that that came

3:26

about was the idea of that

3:29

certain tasks are toil tasks. You are

3:32

working on a product feature or

3:33

something like that. Somewhere in the

3:35

middle you get a ticket. The ticket says

3:38

fix this

3:39

massive vulnerability. You quickly

3:41

reshuffle your work.

3:43

You say, "Okay, I need to make this fix,

3:44

but to make this fix, I have to make the

3:46

fix, I have to test the fix, I have to

3:47

make sure it's working in production.

3:49

And I'm the only person who can look at

3:51

it because I own the fix.

3:53

But, the actual fix is minor.

3:56

It's actually the validation loop that

3:58

takes

3:59

a long time, and that's what I'm

4:01

required for.

4:02

So, what do I do?

4:04

I have to reset my schedule and go and

4:05

make the fix myself. Now, actually at

4:07

Uber, as the stack grows the stack grew,

4:11

this kind of work became common work,

4:13

common engineering work. So, what is

4:16

Minions about?

4:17

We made a bunch of trade-offs. There are

4:19

toil tasks as a result of a lot of these

4:21

trade-offs that people have to deal

4:23

with. Can we give them a way

4:25

to

4:26

win back time

4:28

by

4:30

uh stack ranking their work

4:32

in a way where they can background some

4:34

of the toil,

4:35

right? So, that when they get to the

4:38

validation, 50% to to 70% of the work

4:42

required has been completed by a

4:44

background agent. So, first is the bet

4:47

on a a a specific tool that can do a

4:50

specific task.

4:51

Secondly, is that can do that task well.

4:55

So, a lot of the time people want to

4:57

say, "Write me a big feature from an

4:59

agent."

5:00

We cannot measure I cannot score on

5:02

that. I can score you on whether you can

5:04

do the basics. So, I tell you to change

5:06

a line of code, if you can figure out

5:08

whether you need to run tests also

5:10

as an agent, right? You need to

5:12

configure out, "Oh, I have a tree of

5:13

dependencies that may be affected by the

5:15

change I'm about to make." That I don't

5:17

have to explicitly tell you. This is the

5:18

bet. The bet is that toil work

5:21

that seems simple, but actually has

5:23

wide-ranging scale for a company like

5:25

Uber, can we automate the toil? And so,

5:27

I wrote like a little doc, and because

5:29

of that we founded like a platform where

5:31

we put this agent and a number of these

5:33

other specialist agents into our CI

5:36

infrastructure.

5:37

And we accept prompts in order to do

5:40

work. It goes through our CI checks, and

5:43

all our tests run,

5:45

and quality checks,

5:46

um you know, diff description, as much

5:49

metadata as we need. But the the end

5:51

goal is to actually remove the toil from

5:53

the engineer. And one key point I want

5:55

to make

5:56

the engineer is a is a sensitive

5:58

creature. That is

6:00

don't promise an engineer the world when

6:03

you give them a platform, give them a

6:06

valuable use case because you have to go

6:08

to where the pain lives. If I can prove

6:11

to an engineer that the the core value

6:14

proposition of taking certain painful

6:17

toil tasks away from you is beneficial,

6:20

they will naturally understand that AI

6:22

is not here necessarily to take away all

6:25

of the work, just the parts of the work

6:28

that are tedious, they require rigor and

6:30

tedium to to execute. When it comes to

6:33

those toil and those use cases, how are

6:35

you, you know, thinking about the sort

6:37

of use case identification and roll out

6:39

across Uber? Yeah, this is a big

6:42

challenge, I will admit.

6:45

Mainly because

6:46

you have the industry and the industry's

6:50

documentation of the various product use

6:53

cases for

6:55

each entity AI, right? And then you have

6:58

the internal conflict, which is not

7:00

everything will apply, but a lot of it

7:03

can in different ways. The common

7:05

feedback we get actually is that I don't

7:08

know what to choose. I don't even want

7:11

to choose.

7:12

And

7:13

I think for us that's very big signal,

7:15

right? If it comes down to us to make a

7:17

choice, then we go quite strictly with

7:19

like our developer satisfaction and

7:21

data. And the satisfaction data is like

7:22

onboarding and stuff.

7:24

Toil work is taking away a lot of my

7:26

time. I barely get any time to do deep

7:28

work, right? You have very clear signal

7:30

that you kind of need to invest Oh,

7:32

okay, code reviews are taking a lot of

7:34

time. We can't be like Apple. We can't

7:35

just put out a tool and say, "You use it

7:37

this way." Actually, they're going to

7:39

tell us, "No, No, we use it this way.

7:41

You need to listen to what I'm saying

7:43

and I I there's a little bit of a

7:44

negotiation there, too. We can now take

7:46

what bigger swings, I'd say.

7:48

Oh, reviews are a bottleneck. Let's try

7:50

to make high-quality agentic reviews.

7:52

Is is toil work like cleaning up feature

7:54

flags or cleaning up things or doing

7:56

small changes a bottleneck?

7:58

Let's give you that ability to

8:00

get you out of that zone and then

8:02

measure directly. I Is this actually

8:04

helping you? I think that's what most

8:05

organizations working in the space are

8:07

really looking for, right? They're

8:08

looking for people with

8:10

ingenuity to act on boring feedback.

8:13

You're going to get every quarter you're

8:15

going to find out documentation

8:17

is [snorts] is a is is a pain point,

8:18

that toil is a pain point. The feeling

8:20

of productivity is as important as

8:23

actual productivity. These two go hand

8:25

in hand when you actually want to have a

8:26

productive software engineers or a

8:28

productive workforce, right? Uber has or

8:30

has been investing in also in a platform

8:32

called Shepherd. What does it look like

8:33

if we can also apply agents at scale

8:35

across our organization uh for something

8:37

like migrations or you know, sort of

8:40

upgrades and things like that. You know,

8:42

anyone who's worked with, you know, Java

8:44

or some of these other tools know that

8:45

they need upgrading all the time. So,

8:47

then you're going to have to do this

8:48

across a number of different services,

8:49

some of which may not necessarily be

8:52

actively developed upon. And it seems

8:54

like Shepherd is

8:56

then exploring and experimenting with

8:58

that type of like large-scale change.

9:01

So, the idea behind Shepherd

9:03

is to take the real work that needs to

9:06

get done,

9:07

but put it in the context of an

9:09

evergreen engineering uh

9:12

yeah, it's evergreen engineering, right?

9:13

Like that's that's how I would put it,

9:14

right? You don't want to

9:16

be chasing your tail sort of when you're

9:19

doing this work. The work is a fact of

9:22

engineering.

9:23

You have to be on the cutting edge

9:25

regardless. But, what we found, I think

9:27

even with things like Shepherd, is if

9:29

you can give people the platform to be

9:31

able to disseminate changes

9:33

and give them the tools to be able to

9:36

have a mental model of a large change.

9:38

They can operate on that primitive.

9:40

And that that's a primitive. At the end

9:42

of the day, a large-scale change is a

9:43

primitive that we support now, which

9:45

happened. Minion is sort of one form of

9:47

primitive that can push toil off into

9:49

the background. Then you have Shepherd,

9:50

which is more for these large-scale,

9:54

you know, mass changes. And then then

9:56

Shepherd is effectively going to

9:58

manage the sort of merge conflicts and

10:00

keeping those PRs up to date and give

10:02

you a sort of centralized view then for

10:04

those mass changes across across the

10:06

organization.

10:08

Right, exactly. And and I think

10:10

the the the the angle with Shepherd

10:13

ultimately is

10:15

we need to keep evergreen code. We want

10:18

to keep our code healthy, and we want to

10:20

give people the primitive to be able to

10:22

do that. And less like Minion, which is

10:24

targeting some level of toil in one way,

10:27

Shepherd is targeting some of this work

10:28

in a different way. Organizations where

10:31

the engineers, I guess, identify a

10:33

little bit more with the idea of being a

10:34

builder or they're a little bit more

10:36

outcome-oriented,

10:37

uh seems to be where the organizations

10:39

are moving faster towards agentic or

10:42

AI-driven development.

10:44

Is the fabric of Uber,

10:47

you know, kind of that builder-oriented

10:49

culture? And would you say is is that a

10:51

factor in kind of what drives your speed

10:55

of adoption of AI? Is that Mandating the

10:58

use of AI didn't really get an option

11:00

where it needed to be.

11:02

People had to realize for themselves. It

11:04

had to reach a point of capability

11:05

before people adopted it.

11:07

What does that tell you about the

11:08

culture of engineers in general? They're

11:10

all builders. In fact,

11:12

the reason why we even have 90% of the

11:14

primitives we have before AI that I'm

11:16

talking about

11:17

is because the builders need to keep

11:18

building. They need to move fast.

11:21

For example, Dev Pods is just a a

11:23

function of being able to parallelize

11:25

manually, essentially, right? Uh like

11:27

the Dev Containers. Uh

11:28

these remote dev environments is just

11:30

parallelism without agents, right? We we

11:33

we already reached that point where

11:35

our builders were demanding more easy

11:39

ways to build. So, what does that mean

11:41

for platform? It means for our platform

11:43

side,

11:44

we have to also think about

11:46

whether we can wear some of those hats

11:48

for them

11:49

and

11:50

consider

11:52

what a builder would scrutinize as a

11:55

waste of time sometimes, right? It's How

11:58

do I make them stick? is is kind of And

12:00

the culture

12:02

of making it stick comes from being able

12:04

to give people the ability to fail in a

12:06

safe way.

12:08

If I If I give you the tool

12:10

and I tell you use this

12:12

or else, that's not going to work. If I

12:14

tell you I give you the tool and that

12:16

tool says I can do X, Y, and Z and then

12:18

you actually put it in production and it

12:20

can do a quarter of X, a tenth of Y, and

12:23

a fiftieth of Z, I will fail there as

12:25

well. What I need is

12:28

the engineer to be able to see that I

12:32

know the pain they see.

12:34

That's what gets a builder really

12:36

hungry.

12:37

When they know that they can take the

12:39

tool and they can mold it to their

12:40

outcome. In other words,

12:43

maybe for us agents at our company, for

12:46

us agents are

12:47

the way to go because our engineers are

12:49

builders who need to be able to finally

12:52

configure and control what they do. And

12:54

not that we um we are necessarily

12:57

forcing them to work with an expectation

13:00

that these things do things that maybe

13:01

they're not capable of. Any sort of

13:03

parting wisdom or final thoughts that

13:05

you have that you'd like to share? It's

13:07

not actually that late. And in fact,

13:10

most people still don't have it figured

13:12

out. And there's room here. There's

13:14

runway here to figure this out. Um

13:17

because the primitives have always been

13:18

the same.

13:19

Can I move my engineers faster?

13:21

Can I take away the pain that they see

13:23

on a daily basis uh so that they can

13:25

focus on building high quality features.

13:28

That mission as a platform is the same

13:31

10 years ago and it's the same today.

13:33

The only difference is the means with

13:34

which you go about solving your problem.

13:36

All investments

13:38

may or may not pay off. And of course

13:40

it's a sense of prioritization. That

13:42

doesn't mean the primitives of high

13:43

quality software delivery and

13:46

um

13:47

rapid iteration. The These things are

13:49

important to actually most I would say,

13:51

that's my personal opinion, most tech

13:52

companies are going to want rapid

13:53

iteration. So you want to have some of

13:55

these fundamental and make sure your

13:57

engineers have a space to fail.

13:59

Because if they fail and it's expensive

14:01

for you, that ends up looking kind of

14:04

tough for most people because the it it

14:06

erodes trust, right? So you want to

14:08

build a platform that lets people have

14:10

the space to fail and protect people

14:12

from toil and focus on the same things

14:15

that they're always trying to do, which

14:16

is build high quality products and

14:18

features for the business. Amazing.

14:19

Nikhil, thank you so much. Thank you.

14:21

Thank you, Lou.

14:22

>> [music]

14:28

[music]

Interactive Summary

Uber's engineering organization has built a mature internal AI platform over a long period, making deliberate infrastructure investments. Their key platforms include Minion, which automates 'toil tasks' like minor fixes and validations, contributing to over 11% of all merge PRs, and Shepherd, designed for large-scale migrations and upgrades to ensure 'evergreen' code health across hundreds of millions of lines of code. The company also developed Devbox to provide reproducible remote development environments, addressing environment isolation for engineers who work across multiple platforms. Uber's success in AI adoption is attributed to its builder-oriented culture; instead of mandating AI, they identify engineer pain points through developer satisfaction data, providing valuable use cases that allow engineers to reclaim time from tedious tasks and focus on high-quality feature development, while also offering a safe space for experimentation and failure.

Suggested questions

5 ready-made prompts