HomeVideos

From Signal to PR: Anatomy of a Self-Improving Agent — Jason Lopatecki, Arize

Now Playing

From Signal to PR: Anatomy of a Self-Improving Agent — Jason Lopatecki, Arize

Transcript

509 segments

0:01

[music]

0:12

Well, thank thank you all. Um, let me

0:15

get set up here. So, not just the the

0:18

founder of Arise, but but I tend to

0:22

build an incredible amount of stuff. Um,

0:27

let's see if we get this going here.

0:30

Oh, sorry. One more second. Um, so not

0:35

just a founder here, but but also a

0:37

builder and I do my best to um uh to to

0:42

to build agents assistance. Um, we have

0:46

an agent in product.

0:48

We have an agent in product called Alex

0:51

and uh and a lot of I think a lot of my

0:53

experience has come from actually um

0:56

trying to make the stuff work and work

0:58

well. Our first version of our our own

1:01

agent frankly sucked. Uh it was many

1:04

years ago uh probably two years ago when

1:06

the first in the space to do it. Um and

1:09

a lot of what we have built uh has come

1:13

out of that our own experience in

1:14

building building this agent and and

1:17

signal is kind of our our next

1:18

generation of this which is trying to

1:20

automate a bunch of things which we do

1:23

every day uh and build it into products

1:25

that that people can use. Um so I'm

1:27

going to try to I'm going to go through

1:28

this this materials here. I'll try to go

1:30

fast and try to show you a lot of

1:31

product too. I'm a product person. Um,

1:34

so

1:36

if you've built a startup before, um,

1:38

you you've experienced this, your

1:40

platform's down, it's it's late at night

1:42

and and you and you want to go fix it.

1:44

Um, and and really the the it takes a

1:48

lot of energy to go do that. And we're

1:49

going to talk about like the automation

1:51

we've built a little bit and and what

1:53

what the future looks like and and I

1:54

truly believe um that the future of the

1:58

the observability space is is actually

2:00

changing massively right now. Why is

2:03

that? Well, observability used to be for

2:06

humans. Used to be a UI you click, a

2:08

graph you click, something you look at.

2:10

Um and and today it's I would argue it's

2:13

a lot of 2.0 which is like this

2:15

combination of coding agent. Those of

2:17

you who built skills skills for um

2:20

Pyroscope, Google Cloud or or whatnot,

2:23

that these these skills help you with

2:24

your your human debugging these systems.

2:27

Um and and really telemetry is like this

2:31

smoke uh thrown off of your system that

2:35

can allow these agents to go make fixes.

2:38

It tells you what path in the code it

2:40

took. Without that, you're guessing and

2:43

there's a million paths it could have

2:44

taken. the the data thrown off by your

2:47

system allows um allows you to to go go

2:51

use agents to go debug your software.

2:55

Evals add another layer to this. Um but

2:58

really what we're at here is is how do I

3:01

build systems that autonomously fix

3:04

themselves? Really that that is what

3:06

we're after. Both both AI agents, I put

3:08

AI into my my my system. How do I have

3:10

this thing just improve itself? and and

3:13

today we're kind of in the 2.0 which is

3:15

a human making fixes and reviewing

3:17

things. Um but there's a future we're

3:19

all driving towards and throwing off

3:22

traces, throwing off logs, throwing off

3:24

way more than you normally would and

3:26

having agents run at this for a

3:28

continuous loop is where we're going.

3:30

You can build at agent speed, but today

3:33

you can't improve your systems really at

3:35

this agent speed. So those of us feel

3:37

this this kind of governor happening

3:39

within our our our our products. um and

3:43

and the bottleneck is actually not the

3:45

fix anymore. So those of us who've used

3:47

these systems and and use used um coding

3:50

agents with with skills, the the the

3:53

bottlenecks, a lot of the the confidence

3:55

in in do I have it right? You know, a

3:58

lot of this is is about is this fix the

4:00

right one to push? Um and and so these

4:05

are kind of the challenges here. And

4:06

then how do you how do you build this

4:08

loop in a way that just moves faster? Um

4:11

and and a little bit of the way we we've

4:14

kind of come to do it and we do it in

4:16

our system is we've kind of inverted

4:18

this this loop which is like a human you

4:21

know looks at things and an agent uh

4:23

fixes it to a person now can wake up

4:27

with with an idea of the issues based

4:29

upon the errors occurred in their

4:30

system. So so the agent is actually you

4:33

know maybe it's not a a fix itself but

4:35

it's putting up an issue. It's looking

4:37

at the data before a human even looks at

4:39

it. Um and and what you move from there

4:42

is is kind of humans grabbing tickets to

4:45

to having some amount of evidence um

4:48

some deep evidence relative to whatever

4:49

you're looking at already sitting in

4:51

front of you by the time you actually

4:53

even look at it.

4:55

And and human review is kind of one

4:58

thing but but a lot of times maybe

4:59

you're driving this little investigation

5:01

a bit from where it started. So that's

5:03

the reality of where we are today is

5:04

there's still maybe it's not human

5:06

reviewing but human driving the the step

5:08

two and three. Um but but this is kind

5:11

of what what we view the loop as. And

5:13

really what it is is there's you know

5:15

there's an event that occurs that you're

5:17

kind of kicking things off on or you're

5:18

looking at periodically. Um and then

5:20

there's some context around that which

5:22

is really driven by skills. Um, I guess

5:25

a question for all of you. Who's created

5:27

skills in this room? Who's created a

5:30

skill that that that interfaces to an

5:32

observability platform?

5:35

Okay, handful. Okay, cool. Awesome. Um,

5:38

so the magic of of of skills that that

5:41

that connect to observability platforms

5:44

um is it can gather the context. The

5:45

agent can decide what it needs, what it

5:47

needs to look at um to to start to

5:49

troubleshoot what you have there. Um,

5:52

and then there's idea of triggers which

5:54

are like periodic and and um and uh and

5:57

event based. And so the future

5:58

observability actually looks a lot more

6:00

like this than it does clicking around

6:03

graphana UI.

6:05

So first off evidence well normally

6:08

these like or are or what do you start

6:11

with what do you look at? Uh traces are

6:13

are pretty nice logs as well. uh but but

6:16

you know most of the AI systems these

6:17

days have like traces at the core of of

6:20

the agent framework. So, so you kind of

6:21

start with with looking at traces and

6:23

this is this could be periodic you know

6:25

every five minutes this could be based

6:27

upon an event an error and normally

6:30

there's some combinations of these which

6:32

is um you know some some like uh context

6:36

and log you know context and skills used

6:38

to put together logs maybe there's the

6:40

repo uh you want kind of a combination

6:42

of all this together um to understand

6:44

what to go fix the repo tells you the

6:47

code path that you know the you know all

6:50

tells you everything that's there, the

6:52

the production logs or traces that the

6:54

agent pulls down. Um, normally our

6:56

skills actually pull pull little temp

6:58

files down into the the repo. Um, so

7:00

that you kind of have this this this

7:03

idea of what actually happened, what the

7:05

code is there enough and and all that

7:07

together to put up a fix. Um, so it's

7:10

this combination of the right data and

7:12

file format in the repo along with your

7:14

code in the repo. That's kind of the

7:15

magic of this skills which are

7:17

composable for the agent to go actually

7:19

put a fix. Um, and a lot of this, some

7:23

of you, a lot of you probably do this

7:24

locally today. You you run this locally.

7:27

You have an agent that you you kick up,

7:28

maybe you're spinning up, but it's on

7:30

your laptop. And I think we all feel

7:33

this this this move from from this

7:35

laptop um to to maybe to to basically

7:38

sandboxes. Um

7:41

and and really the sandbox is this this

7:43

running environment where um based upon

7:46

an event or a periodic you know a

7:48

periodic event you can kick this thing

7:50

off and it does the same thing you were

7:52

doing locally get it working locally

7:54

first locally on your laptop and then

7:56

event based based upon the observability

7:58

platforms like ourselves. Um you can

8:00

trigger these on a schedule or or kind

8:03

of you know every every error that comes

8:05

up. Um, and generally, you know, you

8:09

generally it's kind of put putting the

8:10

loop together to do this and and um, and

8:13

I want to kind of give you one example.

8:15

So, this is Alex, our agent. This is a a

8:17

real example. It's a very simple one.

8:19

And then I'm going to show you what it

8:20

looks like in product. Um, this what we

8:23

use every day. Um, but this is just an

8:26

example where um, we had a stream

8:28

canceled event. So um so Alex is is

8:32

basically um Alex is is basically our

8:35

our inproduct assistant. Um to-do update

8:38

is is a uh is a is a way of of managing

8:41

kind of its its task list. Um and it was

8:44

trying to you know I'll walk you through

8:46

the the error in a second but basically

8:48

um it's calling a bunch of these two

8:50

to-do updates and kind of um errors out

8:52

and and so for us it's it's you know how

8:56

do I put the data together um to debug

8:58

this? how do we do it automatically and

8:59

and signal is just something that's

9:01

running in the background for us that's

9:02

putting up like issues relative to these

9:04

things. Um this was uh a kind of one or

9:08

two line fix that it comes up with.

9:10

These are these are ideal but a lot of

9:12

times the fixes are bigger. Um and and

9:15

the bigger it is the more likely a

9:16

human's involved in kind of like

9:18

spearheading it over the line. But again

9:20

it's about that that cold start. Can I

9:22

start with like all this information on

9:24

the issue and guide it the rest of the

9:26

way is kind of where we are right now.

9:27

Um, and for us, your job kind of moves

9:30

from responder to reviewer. Um, and and

9:33

and the view is like traces and evals

9:35

don't go away in any way, shape, or

9:37

form. They're just they're they're a key

9:39

part of the loop. Now, you're going to

9:41

trace 10 times more. You're going to log

9:43

10 times more because that helps you

9:46

know what path your software took.

9:48

Before, you wouldn't do that because

9:50

because humans can't dig through all the

9:53

logs. It's just noise. But by logging

9:56

and tracing more of your like is it

9:59

every inch of your software? Maybe in

10:00

some places. Um by logging and tracing

10:04

orders and orders of magnitude more than

10:06

we do today, we can actually create

10:08

these continuous loops that know what

10:09

path was taking your software and and

10:11

and actually have it fix itself. So this

10:14

is kind of my vision for where I think

10:16

things are going. um in a way and and

10:18

for us I'll show you signal in a second

10:20

and you'll see all these I mean I feel

10:23

like there's there's this think of this

10:24

as an an SR you know something that

10:26

helps you debug maybe SR for for AI um

10:29

but I feel like there's a lot of black

10:30

boxes out there like oh there's a SR

10:33

agent that does this or S agent that

10:35

does this all we're really trying to do

10:37

ourselves is take your local debugging

10:39

experience with cloud code cursor and

10:43

run it periodically so pick your sandbox

10:46

pick your harness, pick your skills,

10:48

we'll pre-bank a bunch of things with

10:50

you. So, we're just trying to again take

10:52

the things we were doing locally and

10:54

actually run them um uh you know, run

10:56

them in a system. So, we believe in you

10:58

know an open approach um to this and um

11:02

I I'll give you a demo of what this

11:03

looks like um from a product

11:06

perspective. So,

11:10

so this is um this is a a financial

11:13

trading agent. Um given what you saw in

11:16

the previous

11:17

uh presentation, I would not recommend

11:20

doing a financial trading agent. Um they

11:22

they they're unlikely to make you money.

11:24

Uh at least not not yet. Um maybe

11:26

there's some people uh doing it good.

11:28

But long story short is this one's, you

11:30

know, uh people asking questions about

11:32

stock trading right now and it's giving

11:33

giving answers. Um there's a lot of ways

11:36

this this can fail. And so this this

11:39

gives you this is a rise. It's a

11:41

platform. So first off from let me

11:43

describe the products we have. Uh this

11:45

is AX which is our our SAS platform. Um

11:48

we also have Phoenix which is open

11:50

source if you just want to start

11:51

tomorrow. Um Signal right now is is just

11:54

available in in our AX SAS platform. Um

11:56

which also could be deployed VPC but but

11:59

this so so give you an idea of our

12:02

product lines. Uh if you want to try out

12:03

Signal, it's it's an AX. Um

12:07

and what it looks like is something like

12:08

this, which is it's just periodically

12:10

running and and kind of coming up with

12:13

like issues and you can hook it up to

12:16

your GitHub repo. It can create an issue

12:18

in your repo. You can create an

12:20

evaluator from this. Maybe u maybe

12:22

there's a a specific problem by which

12:25

you want to catch again. You can add

12:27

these to a data set. So if you want to

12:29

add these and and it has evidence

12:30

associated with this like traces um in

12:33

this case um in this one it has skills

12:36

like for Google cloud and some other

12:37

logging systems. So we can frontend a

12:40

bunch of places the data we're we're

12:42

pretty good at building I think these

12:43

skills to debug issues again uh but you

12:46

can add your own skills. So these

12:48

examples here you know traces running

12:50

out without a guardrail. Um there's

12:53

there's um uh you know safety safety

12:57

issues and intent issues and a lot of

13:00

this too is like you know how does this

13:02

work? How do I you know it feels a

13:04

little too black box to me. Well all

13:05

this is open and open box um in the

13:09

sense that um I can set up you know I

13:13

can set up the harness that I want it to

13:16

run on. This one's cloud code. I can

13:18

pick my sandboxes and sandbox systems.

13:20

Um, I can use cloud managed agents if I

13:22

want. I can use Arise sandbox. Uh, why

13:25

would I want to use Arise sandboxes

13:27

versus cloud managed agents? Well, a lot

13:30

of our customers um don't want to

13:32

connect their production systems to

13:33

Tanthropic. You know, you you you want

13:36

your s you want these these sandboxes to

13:38

debug your database or connect to it. So

13:41

we install in the VPC of a lot of you

13:44

know bigname companies out there um from

13:46

from Uber to um to bookings to you name

13:50

it and and these people don't want to

13:52

send their connections out but they'll

13:54

they'll use a s you know many many

13:56

companies um are very comfortable

13:59

installing a V into a VPC and actually

14:01

connecting it up. So you can use Arise

14:04

sandboxes or you can use Daytona or any

14:06

of any of these that you're comfortable

14:07

with um that you built relationships

14:09

with. Um and and then from a a platform

14:13

perspective, you know, we support

14:19

um running we support, you know,

14:22

tracking the different agents that

14:23

you're running. So you have this swarm

14:25

of agents. Maybe you've kicked off um

14:27

maybe you're kicking off a signal which

14:29

which is our agent. It's running

14:30

periodically. Maybe you're kicking off

14:32

your you know you've named another agent

14:34

um in the system. And these all support,

14:38

you know, viewing the session that that

14:40

ran, downloading the transcript, and you

14:43

can resume a clouded session locally,

14:45

too. So, the idea is that this thing's

14:47

constantly running. You're picking the

14:48

harness, the sandbox. You're deciding

14:50

the prompt if you want, hey, don't be

14:52

aggressive or look for, you know, look

14:55

for security issues. So, you're deciding

14:56

the prompts that drive this and you're

14:59

also deciding the skills that go along

15:01

with this. So um in

15:06

a preset here um I can add you know add

15:09

different skills. I can add my own

15:11

skills. I can link repos. I can I have

15:13

pre-baked skills too. Um so the ideas

15:16

observability platforms are really

15:18

starting to get are becoming tied to the

15:20

continuous loop to the the fix not just

15:24

the the signal. Um, and and you you want

15:28

to take your local experience you have

15:30

debugging the stuff locally. You want to

15:32

take the evals that are running and and

15:34

actually have these all work in

15:36

something that puts up a fix or at least

15:38

gets you a cold start and then I can

15:40

take it over locally if I want to

15:42

continue debugging uh from from here.

15:44

Um, so this gives you um a rough idea of

15:47

of kind of of of signal to PR um what

15:51

we're doing. Um, I did want to offer,

15:54

you know, questions if people people

15:56

have any questions on what we're doing

15:57

or how we see uh the industry evolving.

15:59

Happy to happy to answer. Thank you.

16:04

[applause]

16:13

>> Yeah. Go ahead.

16:20

question of like why can't we just

16:22

connect cloud code to your data and have

16:26

cloud code do all these things. I think

16:28

there's like a version of that question

16:29

that can probably be asked for these

16:31

autofixes, right? Like why not have

16:32

cloud code read the traces and push the

16:34

PR itself.

16:36

>> I'm curious how you would respond to

16:37

that question.

16:38

>> Yeah. Uh so so why wouldn't have cloud

16:41

code kind of hook to your your data and

16:43

just just do it? Um the answer is like

16:46

you should um like like the vision and

16:49

what we do actually at at Arise is we

16:51

have uh a lot of skills. I think first

16:53

off to to make that really work well you

16:55

have to do a bit of well-designed skills

16:57

in the data space. The skill like the

17:00

the important things of designing these

17:02

skills are are around really around

17:05

getting data you know finding the right

17:07

data first. So I want to find a group of

17:09

traces relative to a session or

17:10

something. Getting that data into the

17:13

repo in a file format. These harnesses

17:15

are magical with files. So you get the

17:17

file what happened. In some cases we

17:19

have 10meg files like sitting in the

17:21

repo. Um so it's designing the skill to

17:23

be really really well done with the um

17:26

with this data and and then giving

17:28

claude enough skills to be composable to

17:30

find issues. So the answer is absolutely

17:33

yes. like we have Pyroscope skills that

17:35

will find memory issues. We have facets

17:38

in Pyroscope that the skill knows how to

17:40

use. I can cohort by customer to see if

17:42

a customer is causing an issue. Um but

17:44

but you've got to kind of design the

17:45

skill surface area in a way that Claude

17:49

can really really work well and and and

17:52

it's not just like point Claude at the

17:54

data.

17:54

>> I see. Thank you.

17:57

>> Any other questions?

17:59

Anyone else?

18:01

>> Oh yeah. Okay. One more.

18:04

Thanks for the talk. Um, there was quite

18:06

a few mention of eval but you know I'm

18:08

looking at the traces so you know I

18:10

understand the concept of traces but

18:12

where where did the evals come in when

18:14

you have that signal that says hey

18:15

something broke in production.

18:17

>> Yeah. So so the so the eval typically

18:19

will the um the eval essentially are

18:22

running and being layered on typically

18:23

to the production traces something we

18:25

call online evals. Um let me see if this

18:28

one has an example here of it. Um so so

18:33

eval actually are data on the trace

18:35

itself and so the agent knows how to uh

18:39

grab the data from traces knows how to

18:42

um visualize and you know the skills to

18:44

basically pull data for for the

18:46

aggregate values of evals across the

18:47

traces so that so the the skills that

18:49

you give uh the harness allow it to get

18:53

the data on the eval from from the

18:54

traces. Um so eval are kind of like I

18:57

view them as le at least a first

19:00

generation eval evals which are elements

19:02

a judge um as a as a AI layer that

19:06

allows you to run periodically and and

19:08

assess your system but it's like it but

19:11

it's adding a little bit more you know

19:13

pre-processed information on the data

19:15

that that and then as signal is running

19:18

it's using data from the evals that were

19:20

layered on um in addition to all the raw

19:23

data that it has there Um it but it

19:25

tends to be like you build an eval for a

19:28

failure you've seen before a lot of

19:29

times. So I have these prompt injection

19:32

things that I'm trying to catch or

19:33

something or um or or a failure in the

19:36

way it's responded maybe to to something

19:38

before. So they they tend to be this

19:40

like you know u at least the LM as a

19:42

judge is tends to be like this this

19:44

thing you um preset up and then you can

19:47

actually create evaluators for failure.

19:49

Say you find this failure that's pretty

19:50

common and happening all the time. I can

19:52

create an eval so I can catch it next

19:54

time. I just you think of it as like

19:55

almost a an AI um assessment that's

19:59

always running. Uh the other note is the

20:01

element as a judge can run really at

20:03

scale. Well, every you know I have

20:05

customers who who lay you know layer

20:07

element as a judge across um their full

20:10

data set uh where where this tends to be

20:12

like you know uh more periodic on a lot

20:15

of data. So cool. Thank you.

Interactive Summary

The presentation explores the evolution of observability in software engineering, moving from human-centric UI monitoring to autonomous, agent-driven systems. The speaker, a founder at Arise, discusses how AI agents and 'skills' can automate debugging and maintenance by leveraging telemetry data (traces, logs) and continuous feedback loops. The goal is to move from being a responder to a reviewer, where agents identify and potentially fix production issues based on observability data, reducing the manual burden on developers.

Suggested questions

4 ready-made prompts