HomeVideos

Background Agents Summit Day 2: From Foundations to Production

Now Playing

Background Agents Summit Day 2: From Foundations to Production

Transcript

332 segments

0:00

on time, right?

0:13

>> All right, we are live.

0:16

>> Great. Hello everyone. Welcome back to

0:18

day two of the Background Agent Summit.

0:21

I'm Philip and I'm here again with Will.

0:23

>> How's it going?

0:25

>> And for anyone who's new, so this is the

0:28

first-ever virtual Background Agent

0:30

Summit there is. We have in total 14

0:32

talks lined up. Yesterday was day one.

0:34

Um it was mostly about foundations and

0:36

we had seven sessions mostly focused

0:38

around infrastructure, security, context

0:41

engineering. And if you've missed it

0:43

yesterday or if you know, there were

0:45

some technical blips where the quality

0:47

of the streams wasn't as great, don't

0:49

worry, we've got it all lined up in HD.

0:52

We'll send the recordings to you via

0:53

email, so make sure to go to

0:54

backgroundagents.com/summit,

0:56

put in your email address and we'll send

0:58

all of the recordings to you. Don't

0:59

worry.

1:00

But

1:02

before we start with today's section

1:03

sessions, we wanted to do a quick

1:05

wrap-up. So a few themes kept surfacing

1:08

yesterday with that we want to pull

1:10

together. So the first one is really the

1:12

foundation matters more than the agent.

1:14

Unsurprisingly, yesterday was about

1:15

foundations, so we'll talk a bit about

1:17

that. So

1:19

um Alister from Stripe said it quite

1:21

plainly actually. He said the cloud dev

1:23

boxes they had were a strategic credit.

1:26

Um the agent worked because they had

1:28

reproducible um cloud dev environments

1:31

already. And today we have a session

1:32

from from Uber.

1:34

Um so I'll take it a bit away, but he's

1:35

also talking about DevPod, who's which

1:37

is their containerized dev environment

1:39

that they have been using for years. And

1:41

since they had these solutions in place,

1:43

it also allowed them to like start very

1:45

quickly with Background Agents, right?

1:47

Um but for most companies, the

1:49

infrastructure that they have today,

1:51

everything is running local, you have

1:52

like, you know, manual setup scripts and

1:54

all of that. And that's exactly the

1:56

reason why you um

1:58

basically land on the false summit that

2:00

we talked about in yesterday's keynote.

2:02

Um

2:03

I'm wondering um um Will, what your

2:05

thoughts were also on the foundations.

2:07

>> Well, yeah. I mean, I think that it's

2:09

really interesting to see how some

2:10

companies that already had like good

2:12

primitives there from the, you know,

2:15

sandbox, CD, dev box kind of

2:17

perspective, kind of had that leg up. Um

2:20

whereas the companies that did not had

2:22

to kind of wind up on the same approach,

2:25

uh but from the other direction, right?

2:26

So, like Cole who built OpenInspect, he

2:30

mentioned that that was the single

2:31

biggest friction point was that

2:33

environments themselves weren't actually

2:34

reproducible. And when the environment

2:36

wasn't reproducible, the agent wasn't

2:38

reproducible. So, like the the

2:40

primitives, the things that uh build up

2:43

to and wrap around the agent make more

2:46

of a difference than the tools that the

2:48

agent is calling, whether that's like

2:50

riff grep versus grep or like how it's

2:53

building out knowledge graphs and things

2:54

like that. So long as the agent can run

2:57

in an environment that's reproducible,

2:58

the agent itself is much more likely to

3:00

be reproducible. I thought that was a

3:01

really interesting take. Yeah.

3:04

And then that kind of takes us to the

3:06

next piece, right? Reproducibility is

3:09

a lot of that is coming from the world

3:11

of context. And Patrick at Tessell did a

3:14

really good talk at the end. If you

3:15

didn't have a chance to watch that, I

3:17

would suggest taking a look when we send

3:19

out the follow-up emails. Um you defined

3:21

it as like the context development life

3:24

cycle. Um and really tried to strike a

3:27

good balance of like uh generating,

3:30

testing, version controlling,

3:32

distributing, and even observing the

3:34

context aspects of the agent. What's

3:38

going into that environment, right? What

3:40

tools does it have access to? Um what is

3:43

the like most accurate, up-to-date

3:45

version of things? And then uh

3:48

from Cloudflare, Rajesh did a really

3:50

good uh talk about what they called

3:52

engineering codex, um which were

3:56

really rigid definitions of how the

3:59

agent is supposed to work and how the

4:02

applications that the agent works within

4:04

work themselves. And those codex uh

4:08

really improved the quality of the agent

4:10

themselves, right? This was all done in

4:13

backstage, which I thought was so cool.

4:15

Like there's there's this

4:18

kind of ephemeral thing that we were

4:20

doing for a really long for I guess at

4:22

this point it feels like a really long

4:23

time, but it hasn't been that long of

4:25

like IDP engineering and uh

4:29

giving developers a way to self-serve,

4:31

but all of that work has kind of

4:34

it it was nice to have as a developer,

4:37

but it's a must-have for an agent.

4:41

>> Yeah. And I think like adding to that,

4:44

so Patrick also coined the term um

4:46

context um development cycle, right? So

4:48

you really like have to be very

4:50

deliberate about how you inject context

4:52

and um and and when to refresh etc. So

4:55

it's not that simple and I think Kwak

4:57

from Genentech also made a really really

4:58

good point here that I want to um

5:00

um emphasize. So they let their agent

5:02

self-serve, right? They put it on a

5:04

nightly dream, like reflect on what it

5:06

has done and like create new skills and

5:08

like update the existing skills, and

5:10

that has generally improved the outcome.

5:12

But also to um what they observed is

5:14

that it really depends on on on what

5:16

kinds of skills, right? If it's a more

5:18

like factual context as to like always

5:20

go to these Slack channels and like

5:22

here's the ID and everything. Obviously

5:24

more context helps um because it gets

5:26

there faster, uses less tool calls. But

5:28

if it's more a diagnostic context or

5:30

like open context where you have

5:32

something around debugging, for example,

5:34

if you're a very rigid with like super

5:36

long checklist, it actually you might

5:38

end up with the same outcome quality,

5:39

but it just requires more tool calls, it

5:42

takes longer, is more expensive. So the

5:44

lesson really from him was um kind of

5:47

teach the agent what to look at but not

5:49

to what to look for. So,

5:51

I found that super super interesting as

5:53

well.

5:54

All right, let's go into the next

5:56

Sorry, do you want to add something?

5:59

>> I was saying yeah, it's it's important

6:00

not to over-engineer there. You know,

6:01

like I feel like there's still the

6:03

concept of over-engineering like an API

6:05

endpoint, like that still kind of goes

6:07

back to over-engineering the context

6:09

that you're giving an agent.

6:10

>> 100% Yeah, yeah. And it's so easy to

6:12

over-engineer with agents too cuz

6:14

producing is just so cheap.

6:16

Cool, so let's go into the next theme

6:18

that we that we

6:19

that stood out. Obviously, security was

6:21

a big one. We had two sessions around

6:23

that. So, that can't be an afterthought,

6:25

especially if you really go big on

6:27

autonomy and like don't and babysit the

6:30

agent anymore, right? So, we had Leo

6:32

Lorenzo and Chris from our own team who

6:33

talked about my prompt level guardrails,

6:36

which most of our local coding agents

6:38

rely on, keep failing. The like the

6:41

agent will just reason around them and

6:44

that's why the enforcement really has to

6:45

happen like a level below in depth. And

6:48

that's what B2 and Data Wallet doing.

6:50

They have allow and deny lists and scope

6:52

credentials and it really is like a

6:54

kernel level policy enforcement. And

6:57

we also had Stephen from Nono

7:00

from the open source side basically

7:01

sharing the same principles and that

7:04

just drove home the point that

7:06

security can't be an afterthought but if

7:08

you got it got it figured out then you

7:10

can really, you know, with good faith

7:12

let the agent run autonomously.

7:14

>> Yeah, I mean speaking of that autonomous

7:16

running,

7:17

I think Quake from from Genentech did a

7:20

really good job of talking about the

7:22

engineering that has to go into the

7:24

environment itself in order to make it

7:27

reliable enough to merge to main or to

7:30

do things where where it has a bit more

7:32

production capability. So, like

7:35

integration tests, canary deployments,

7:38

feature flags, the types of things that

7:39

you bring into

7:42

uh

7:42

production data and production systems.

7:45

If you

7:46

treat agents the same way by making the

7:49

environment the governance layer, then

7:52

you can do a lot more with your autonomy

7:55

within your you know, within your own

7:56

cloud within your background agents

7:58

ecosystem.

8:00

>> Yeah, exactly. So if if if the

8:01

environment is the governance layer, you

8:03

have the security figured out, then what

8:06

you can do is basically you can move

8:07

from software manufac- uh software

8:09

engineering to software manufacturing.

8:11

That's something that Cole Murray from

8:12

Open Inspect yesterday um kind of

8:14

coined. He was talking about, you know,

8:16

using fleets of agents like overseeing

8:18

them as an engineer. And that really

8:20

already goes into the the factory

8:22

analogy that we talked in the keynote

8:23

and we'll also hear more about later

8:25

today. And the interesting thing here is

8:28

again, if you have these two foundations

8:30

figured out, then you don't have to be a

8:32

super technical person to to work with

8:34

uh background agents because there's

8:35

like not so much that you can actually

8:38

uh do wrong. You don't even need to know

8:39

so much about security, right? So if you

8:41

have those two foundations in place,

8:43

even like PMs or designers can really um

8:46

self-serve more with agents and ship

8:47

actually code or something else. So we

8:49

see this at our at our own company where

8:51

we for example, like the slides that

8:53

you're looking at, we we did them with

8:54

Ona, right? Like and uh a PM did those.

8:57

So

8:57

um nobody that's like super technical.

9:00

So I found that super super exciting cuz

9:01

it's like very empowering if you got the

9:03

foundations figured out. Like it really

9:05

extends towards everyone else in the

9:07

company.

9:07

>> Yeah. And that that stretches so far.

9:10

Like it it's more than just even the

9:11

engineering work close to. Like the

9:13

um

9:14

with with Greg showing that genomics

9:17

could be done in a secure way accessing

9:20

like really secure classified data um

9:25

and being able to do background agent

9:27

tasks on it in a reliable secure way is

9:29

so cool. So I love seeing the potential

9:34

here, you you

9:36

Um so yesterday was primarily about the

9:40

building blocks. Um, you know, how do we

9:42

actually build these systems? How do we

9:44

design the architecture? What sorts of

9:46

trade-offs do we have to make? Um, and

9:49

what are the you know, kind of things

9:51

that we should be experimenting with and

9:54

how are those like leading to other

9:55

decisions down the road? But today is

9:58

more about how do we actually ship this

10:01

within a real organization and what

10:03

sorts of changes do this do background

10:06

agents unlock? So,

10:08

uh

10:10

Yeah, I mean the questions that we're

10:12

asking are are more along the lines of

10:13

like ROI. How do we improve the number

10:16

of uh PRs that agents generate that you

10:18

can actually like open and merge that

10:20

aren't just slop or drafts?

10:23

Um, how does that impact your review

10:25

process and things like that?

10:28

>> Yeah, and we're also going deeper into

10:30

uh what the end state could look like if

10:33

if if that's at least uh conceivable.

10:35

So, the idea of the software factory, we

10:37

talked about that multiple times. Today,

10:38

we have two talks around them. Uh and I

10:41

already saw a comment about dark

10:42

factories sounding shady and fun. So,

10:44

yeah, make sure to tune into that.

10:46

Shardul's a great guy. Um, so let's get

10:49

um quickly let's pull up the agenda and

10:51

go through so you know what to expect

10:53

and when. So, we start off with Joey

10:54

Wang from Harvey who built um who was

10:56

the engineering lead behind Specter,

10:58

their collaborative cloud agent

11:00

platform. Um, where really everything

11:02

revolves around the concept of sessions.

11:04

Then we have Lawrence at IO, Zach who

11:07

built software factory.dev, a project

11:09

where he was like running this factory

11:10

publicly for two weeks.

11:12

Um, then we have Nikhil from Uber. We

11:14

have Shardul from AWS on the dark

11:16

factories. And in the end, we have

11:18

Suhayl from um Monzo who's talking about

11:20

adoption of agents in a regulated bank

11:22

that has I think more than 3,000

11:24

microservices.

11:25

So, um really really interesting talks.

11:30

>> Well, yeah. I guess now's the time. go

11:32

ahead and kick things off for the day.

11:34

Everybody

11:36

tune in for Joey Wang and the rest of

11:39

these calls the rest of these talks.

11:40

It's going to be a really uh

11:42

star-power filled day.

11:44

>> Make sure to ask questions in the chat.

11:46

We will forward them to the to the

11:48

speakers in in case they're not live

11:50

here and we will follow up with the

11:52

answers also via email. Have fun.

Interactive Summary

The video is a recap of the first day and an introduction to the second day of the inaugural virtual Background Agent Summit. The hosts discuss key themes from the previous day, emphasizing that solid infrastructure and reproducible environments are essential for agent success. They cover topics like context management, security guardrails as a governance layer, and the potential for 'software manufacturing' to empower non-technical users to build with agents. Finally, they preview the agenda for day two, which focuses on shipping agent-based solutions within organizations, ROI, and real-world adoption in companies like Uber, AWS, and Monzo.

Suggested questions

3 ready-made prompts