5 Rules for Building AI Agents That Work in Production | Nan Yu & Jacob Shumway
1273 segments
The models are really smart, but we're
just not using them enough. If you
overprompt these things, you're more
likely than not going to make it worse.
>> Give it as little instruction as
possible. Give it the tools to load
context. Don't give it context.
>> Here's all the context. Just figure out
what the right thing is. Write the issue
and work on it. Linear made the issue.
Took 6 minutes and then it gave us a PR
that we click run. You have [music] to
really break down what is the actual
workflow that your users want to do.
Computers can do a lot of work for us.
So let's get rid of all the work we
don't want to do and give it to
computers.
All right. Hey everyone. Today I'm
really excited to welcome Naan and Jacob
from Lineer do a special episode. We're
going to do a deep dive on how to build
a production agent end to end and how it
actually works and we're going to use
Linear's own agent as an example to keep
the discussion concrete and real. So
welcome guys. Yeah, good to be here. All
right. So why don't we start at the
super high level. Can you kind of
demystify this whole agent thing? What
does the explain like on five version of
what an agent is?
>> Yeah, sure. Um, I can answer that one.
So, at a high level, an agent is really
just calling an LLM in a loop. Um,
normally when you call an LLM, you're
going to give it one question, you're
get one answer back. But we want agents
to be really autonomous and be able to
accomplish more complex tasks. So,
typically what you'll do is you'll
define a goal, some milestone for the
agent to hit, and you'll give it tools
that allow it to build its own context,
and then just run it in a loop.
Question, answer, question, answer. On
each turn, it's going to call tools,
pull context in, and eventually it's
going to hit a point where it has enough
information to consider the goal
accomplished. It'll it'll like
synthesize this final response and send
it back.
>> Got it. Okay, that that makes sense.
Yeah, it's it's basically a model using
tools running on a loop. That's kind of
that's kind of like the one.
>> Yeah. Okay, cool.
>> Um Okay, then let's talk about the
linear agent. Maybe now you can show us
like the initial idea behind this thing
and then like now what it's kind of
evolving to.
>> Yeah. Yeah. I I think you know I I I
think it's actually pretty interesting
to expand on Jacob's answer a little
bit, right? Because you know you we you
we asked the question and I I think that
that is like the the correct sort of
technical definition. You know, you're
asking engineers what they're going to
say. Um, but I also think that when we
talk about agents like colloquially, you
know, we think of them as like products,
right? It's like it's like a packaging
of some sort of uh AI loop plus some
other stuff. And I think ultimately
these things are like a bunch of
different subsystems that are all, you
know, interacting and then there's like
one facade, right? It could be a chatbot
or something like that that that kind of
fronts all of it. So, you know, you if
even if you think of like the the
desktop agents people use, there's they
have all sorts of stuff built into them
like schedulers and uh and and these
these other sort of like side uh you
know, tertiary kind of features or
they're all kind of part of the agent as
a product too, right? So, I I think we
kind also have to think about it from
that perspective.
>> True. True. Yeah. Yeah. Sure. Do you
think about all this when you had the
initial idea or like what is your
initial spec or
>> Yeah. Yeah. Sure. So, I I'll show you
right now. I was working with Jacob the
other day. I'm like I dug this up. This
is a memo I wrote in like if you look
it's like the second half of uh 2025,
right? So it's like not that long ago in
human time, but in like AI time it's
it's you know ancient history. Um and
and I I think on on here, you know,
you'll you'll see like there's a lot of
focus that we had on like like computers
can do a lot of work for us. So let's
get rid of all the work we don't want to
do and give it to computers, right? Like
that that's like the central sort of
central thesis of this thing. And uh
it's you know and we kind of structured
this idea about um you know before we
really had this sort of technical idea
in our head. We had this idea about like
there's some kind of triggering event
and there's some context that happens
and you have instructions and it kind of
loops on on actions and then and ends
with some kind of result. Right? So like
we we we already had this sort of
concept but I I think what we uh we were
just we we weren't like ambitious enough
like we did we didn't think it could do
like really interesting things. were
like, "Hey, let's give, you know, let's
create this agent thing." We call it,
you know, we gave it a really like
robotic name. It's like robotic program
manager with the idea, right? And and
just give it the boring stuff, right?
And and I I think that what what we
what's changed, right, about how we
think about it today, right, which, you
know, we're in the second half of 2026,
so it's literally just a year later. Uh
is that we we don't think of just giving
it the boring stuff. Sure, you're going
to give the boring stuff, but also
there's it opens up a lot of like the
creative possibilities, right? Where
where you can do it interactively and
and uh and and it could be a big force
augment for like interesting work,
right? Not just not just the boring
stuff.
>> I don't want to offend anyone, but I
feel like um technical program manager
is like one of the most boring jobs do.
You're literally just [laughter]
like managing spreadsheets, trying to
track tickets and stuff, right? So So
this is what the agent started with,
right? But but now it can do now it kind
of serve multiple hats kind of do end to
end product development, right?
>> Yeah. Yeah. Exactly. And and and
depending on you know what you use it
for and the context that you kind of put
into it. It can do all sorts of very you
know creative things that we at that
point we just like didn't even think was
uh a reasonable thing to expect, right?
And things have moved very quickly.
>> So I'm really curious what your process
is. So you wrote this memo back in uh
third quarter and then uh Jacob just
build it in a couple days or what was
the [laughter] process? Yeah. That was
the process.
>> Yeah. What was the next step to to
actually build this thing to prototype
it or like play with it?
>> Yeah.
>> Jake, what was it the first version of
this that we actually built from here?
>> The first version was was really
prototypy. It was like we were calling
the LLM from the front end directly
>> and it was just flagged internal and we
like we gave it access. We have our
command menu with all our actions. We
gave it access to those as tools and
we're just like let's see what this can
do. It was It was really hacky.
>> Mhm.
>> This episode is brought to you by
Oceans. I hired someone through Oceans
for podcast post-production a few months
back and can't imagine running the
podcast without [music] his help. He's
proactive, picks up new tools fast, and
uses AI to compound everything that he
ships. Oceans doesn't just [music] place
assistants, they place operators. The
talent is AI fluent and delivers the
same output as a senior US hire at 3 to
[music] 5x less cost. They reject 99% of
applicants. So the person who lands on
your team is already operating from day
one. If you're scaling and need
marketing, [music] ops, finance, or EA
help, I highly recommend giving Oceans a
try. Check it out at oceansalent.com/
[music]
Peter. Now back to our episode.
>> Right now you can like, you know, push
code, you can manage tickets, ingest
stuff, but like what were some of the
initial use cases that kind of popped
out that that like, you know, you wanted
to prioritize first?
>> I think it it was simple things. I think
creating issues like it's such a simple
thing but I think that was one that
surfaced really early is useful actually
[snorts]
writing documents that sort of thing.
>> Yeah. I I I think one of the first like
real um use cases that we knew about
that we knew people wanted to do was
they would have uh you know this is
again this is like maybe even before
everyone had uh you know automatic note
takers for everything right like people
were still like handwriting notes and
things like that. they were like, "Hey,
I I I hand I I wrote some notes on the
sales call and they they said a bunch of
stuff. Can I just dump this in there and
just extract out, you know, like the
issues that we need to build, right, for
this or the bugs that were reported or
or those kinds of things, right?" So,
like it it was that, you know, that that
was the very basic operation that we're
like, "Okay, if we let's get that
because like I know that's valuable.
People have directly asked for that.
That's something that we felt, you know,
in our own workflows. So, if we can get
something working that can achieve that,
like that's a that's a reasonable
starting point, right?
For us, it's it's we didn't even know if
it was going to be uh we're going to
have a chatter interface, right? We're
like like there there's some place to
just dump a bunch of text and maybe, you
know, maybe it's just like a a text
field or something like that and you you
just hit submit and it goes, right? So
maybe maybe it looks more like that. So
I I think it was it was very open-ended,
right, when we when we first started,
which is also why we don't have like a
super robust spec for it, right? Like we
didn't go into this thinking that we
knew exactly what we needed to build. We
just we're just like here's some
experiments that some directions we
could try and let's just do it.
>> Yeah. So you launched it on Slack or
something for people to use internally.
>> Yeah. So the the the first production
version of this we we launched uh sort
of secretly right without like really
telling anybody that uh because we we
already had a Slack integration and
Slack you could always mention bots
right we everyone knows you know like
things like donut and things like that
where you know it's it's all very
procedural and like they're like these
different Slack bots. Uh so you could
mention linear before and you know it
give you like a form or something like
that and uh we we just like silently
hooked it up to this right we're just
like okay we're just going to do it and
then if anyone discovers it by app
mentioning in linear they can start
talking with it and then you know they
can they can do whatever they want and
and a lot of uh like a lot of usages for
it emerged that we didn't even expect
and like that that was and like we we
sort of had the inkling that that's what
would happen right is that people would
do things and like the obvious thing
would would be something like hey linear
make an issue to and then you would
describe it in a natural language and it
would do it. But people started
realizing that because they could read
the context, they could just say
something like linear do the right thing
or like look at what we just did, right?
Or just like something extreme, you
know, you can be super lazy,
>> right? Like people will be like, oh, I
just say at linear and then upwards
pointing finger emoji, right? Like it's
like they they do those kinds of things
and then linear will just
>> reason through like what happened? Like
I I know how to create issues. it looks
like there's some issues name so I'm
just going to make some issues and then
and then tell the user that I did that
right so like it it this kind of like
behavior was actually very emergent and
and we we didn't expect it to be able to
do this
>> interesting okay so just so understand
you built this thing and and like kind
of get the model you made the model
aware of all the internal APIs that has
or something people can use the UI for
the model can also do
>> and they just kind of let it loose let
loose for people to try
>> yeah effectively I mean I think Jake we
could probably talk about like you know
we tried certain techniques at first and
then we sort of landed on on a version
of it that have now
>> we tried to give it essentially just
like everything you can do in linear
which is a lot of different actions um
across all the different surface and we
ran into just like context issues and
hallucination uh we tried things like
let's give it our graphql schema and see
if it can write queries um and that that
didn't work really well and we kind of
ended up on a skills type of setup where
>> we we give it the act the ability a tool
essentially to load skills and then it
based on the request it'll just load up
the different skills that it needs and
that comes with a set of tools and
instructions around that.
>> The skills are defined by you guys like
create ticket skill or like you know
>> because we we have opinions on how
different things work. You know, if
you're writing an issue, how do you
think about setting a priority and how
do you think about writing the
description? So, we encode all of that
in these skills.
>> Got it. Okay, that makes sense. Okay.
>> Yeah, I I I think that that's you know
when you when you have a native agent
like this, right? Like people talk a lot
about like, hey, they build a CLI or
they build MCP or something like that
and it comes with a bunch of skills for
how to use it. When you build a native
agent, you can go bug wild with this,
right? You you could have hundreds of
these things and because you have
dynamic loading and you have control
over how everything works. Like you can
have a very smooth and opinionated way
of how it uses your app. And I I think
that this is one of the big advantages
of having a native agent, right? It's
like it's like you can you can just
bring you can just treat it like a like
a power user of the app. there's no
there's no uh variance right in in in
doing that. So I think that that's
that's what we that's where we ended up
and then because it can just you know
run a loop and decide what tool calls to
make and stuff like that. It it all of
this emergent behavior about just being
super lazy when you app mention it.
It'll just figure out the right thing to
do just kind of came out.
>> Can you guys show us well the product is
pretty polished now but can you guys
show us some examples of like tagging
linear in Slack in different channels
and see see what it does?
>> This is a conversation we just had.
there's some behavior that you know I'm
like look this looks a little weird and
I I so like here here's the pattern
right it's like it's very natural we
have a conversation in Slack and I'm
like tagging our designer Yan and Jacob
right and and about maybe some
suggestions about what we can do here
and we're we're trading
>> ideas right it's not like we have some
exact sense of what we want to do right
now right like we're like hey like this
feels bad here maybe I'll try these
things and then you know designer like
kind of gives his opinion here and then
I I try to clarify right? You were like
kind of finding where you know where the
actual problem is. Um and you know Jacob
uh you know raises a uh an objection and
I'm like look we can just here here's my
how we want to address that objection.
So like we're like finding the truth so
to speak and you know at the end of the
day like the message is just like just
do it right like at linear create
[laughter]
issue for me I'm going to just I want to
hold on to it and then just and then now
that linear can write code just just do
a pass like I I'll take a look at what
you did right like previously it was
just make an issue for me but here's all
the context just figure out what the
right thing is like we we argued about a
bunch of stuff we came we ended up
somewhere
>> so like figure out what somewhere is and
then write the issue Now it's write the
issue and work on it. So then linear uh
you know made the issue there's a
decision you know and then it's um it's
assigne is Jacob and it delegated to
itself right and then uh that was 18
minutes ago it took six minutes and then
it can gave us a PR that we can click
through right so like this was the whole
sort of loop of of here's an idea and
then we talked about it and we figured
out you know hopefully where we got to
and then linear sus it all out and then
and then made a PR
>> interesting and then and then now you
can just go in here and like play with
it and see if it's actually a good idea
or not, right?
>> Yeah, you can play with it and see if
it's a good idea or not. And and
importantly, right, like because it's
part of like the linear system, like
this is in Jacob's backlog, like it's in
his, you know, status to-do, right? So,
it's in his personal backlog, like he
doesn't lose track of it. You know, if
if it was just stuck in Slack, it would
just be like, you know, you have more
chats and then all of a sudden you just
lose track of it and then, you know, you
would,
>> you know, kind of cross your fingers
with search or hope that someone
remembers it or something like that,
right? So like the actual tracking like
aspect of this is still matters, right?
You you know because you have it in an
organized backlog that's in the right
project and everything. So that's that's
that's that's the ultimate result,
right? When you think about like what is
linear's goal, its goal is to like put a
ticket in the right place and also, you
know, accomplish the task and and uh
maybe this is actually kind of
interesting. It is like an interesting
product principle. So even though linear
actually did a work, it's assigned to
Jacob. Is is it like a principle like
every agent has to be tied to a human?
the the the vast majority of them are
right like there's going to be
situations where um you know it's it's
really you know the the system
effectively invokes itself right if like
you know if you're if you instrumented
like a data dog or something like that
there was an alarm that tripped and it
threw a you know put a bug into the
system and then linear try to solve it
like there's not no one really touches
it until the very end.
>> So at that point you're relying on the
agent to figure out who should review
the code right all the way there. this
is you know for something like this like
someone made a decision to like do this
right in this case it was Jacob who's
like okay cool I think we have enough
information now let's let's work on this
right so like that that way he he he he
has a handle to it and it's attributed
to him
>> got it okay what kind of information and
context can this agent access uh like
obviously all the tickets in line can
read slack and stuff or like you hook up
to you know gone and every everything
else yeah
>> uh yeah so this this was this context
was just from slack from this thread,
right? So, it can read Slack and and I I
think a lot of the
>> a lot of the benefit comes from
stitching all this together,
>> right? Like Slack by itself isn't
enough. But if you combine Slack with
the ability to read your codebase with
the ability to read like you know your
project description and your other
tickets and stuff like that, then all of
a sudden you can do something, right?
Because like what could have happened
here was uh you know we we we you know
we made the issue and said, "Oh, looks
like I found someone else who actually
did this already, right? There's there's
actually an open PR that exists in the
system." like it would have told us that
that's what happened, right? So because
it has access to all that stuff. It
doesn't just like you know bulldoze its
way through this. It like it knows it's
aware of everything else in the system.
>> Got it. Okay. So let's just go back to
the early like when you stealth launched
this agent in Slack, right? Like so
people probably started using it and
started getting some feedback. Was the
product was just like a simple basic
prompt and some tools that was the
product very complicated back then or is
it pretty simple? How do you improve it?
Um, at that point it it was it
technically had the ability to do a lot
of different things, but it wasn't
really good at it yet. I don't think
we'd arrived on our skill architecture
yet, which unlocked a lot of things. Um,
the primary thing people were using it
for was just creating issues. Um, so we
actually optimized pretty heavily around
that. We created a small little router
and for 80% of these use case, we routed
to a specific subprompt that was just
for creating issues like highly
optimized for that. Um, so that's what
most of our usage was in the really
early days.
>> So the structure of the product was like
there's like a main prompt that maybe
tells the agent what it can do and stuff
and then it it kind of routes to like
subprompts, right? Is that it?
>> Yeah. Yeah. Yeah. Or there's there's a
really small model that runs a router
and that will send it to either this
like big model main prompt that can do
anything and has access to all these
tools or these specific use cases like
creating an issue.
>> Oh, interesting. Okay. So there's a best
practice saying that like when you're
prototyping an agent, you should use the
best model available just to see what
it's capable of, but it sounds like you
guys actually use a mix of different
models for different tasks.
>> We do, but I would say we follow that
best practice for the most part. Uh we
tend to throw the biggest model on it
until we know that it's working well. We
build out some eval. We have a good idea
of like what success criteria looks like
and then you can start to optimize the
model down because you have a really
good framework in place to like make
sure it's still meet. Ideally, you want
to use the smallest model for the job,
right?
>> Yeah. Because Yeah. You want to save
save money, right? You [laughter]
>> Yeah. Okay. So, so then um it's like
very iterative. Like in the beginning,
you probably don't have a ton of eval
set up like automated evals and stuff.
>> Yeah.
>> Yeah.
>> Got it. Just to make this pretty
concrete, like uh let's take the create
ticket thing, right? You probably have
some emails for did they actually create
a ticket or not or like is the ticket
useful or how do you evaluate how good
it is?
>> Yeah. Um, yeah, I mean, yeah, pretty
much. We have eval. A lot of it comes
from just iterating and using it. So, a
user will use it in a way that isn't as
expected. We'll add that to our data set
for our eval.
Um, but we try to have a mix of like
objective and then more subjective
measures.
>> Objective and sub. Okay. So, objective
is like uh deterministically do this
thing or not.
>> Yeah, exactly. If the user says in
progress, make sure it always adds the
status in progress. that's really
deterministic. And then there's more
subjective things like did you structure
the description in a in a good way? Did
you extract the the right information
that should be the title into the title
field?
>> How do you evaluate like is that like a
yes no thing or is it like a scoring?
>> Yeah, it's it's LLM as a judge. So we'll
we'll yeah again we build out this data
set over time. Then we just have a a
scoring type of LLM that's like did this
extract the right information? this is
what it should be pretty much.
>> And then and then I don't know I don't
have a ton of experience doing this
stuff but like I I feel like you have to
you have this element as a judge which
sounds really fancy but you have to look
at the judge and be like hey is this
actually judging it correctly or not?
It's actually pretty manual, right?
>> Yeah. And we we actually try to use
those less often for that reason. I
think I think eval are most successful
when you're ensuring consistency
somewhere that consistency is important.
But consistency is not always important
for agents. They can have a lot of
variance in how they respond to things.
Um, and you really don't want to have
too many emails around that because then
it just false signals.
>> Okay. Interesting. Okay. So, you started
with uh creating a manage tickets and
then uh what are some other use cases
that you you decided to support before I
I don't remember when the thing was
first launched but like before the first
launch. So I I think the the the way
that you should think about this is that
like there's a because it's especially
because it's like um very purpose
specific. There's like this power law of
like use cases, right? Like if you think
about like what the purpose of linear
is, it's like you're you're trying to
structure uh the intention of the of the
company,
>> right? Like you you have ideas, you have
meetings, you have Slack conversations,
you have whatever, right? Discussions.
And at some point you make a decision.
And we're, you know, we were originally
like a system of record to codify those
decisions. So like if you if you just
like look at the distribution like the
vast majority of um how people use
linear like the agent, right, is like
codify decisions. Like I had this this
you know free form conversation with a
customer. They they complained about a
few things. Let's extract what those
things are and figure out what what to
do with them. or um you know the the
sort of thing that we demoed uh or or
even like hey we changed our opinion
about something right we we we had a we
had a big meeting about something just
like pull the notes and like we changed
so many opinions about this uh the spec
of this project or something like that
just go and update all this so that all
the the marketers and stuff like that
don't get false information. So [snorts]
all those are the those are like the
motions right that are are fairly common
but like ultimately this is powered by
you know Frontier LLMs right you can do
anything you can do anything that you
want right and like like what you know I
I I've used things I've used it for
things like being an interview grader we
have like a document in in one of our uh
teams that's like here's the criteria
for you know how we want to evaluate
someone's like interview process or
whatever it is and because I have the
granola MCP connected to my linear agent
right I can just be like, hey, that last
interview I just had, could you just
quickly give me a score on how we go
against this rubric, right? So, I don't
have to like read through all the notes
and remember it's like a starting point.
So, there's like a lot of things that
you can uh you can utilize it for as
long as you have uh the context
somewhere.
>> Got it. Okay. When you guys launched Lar
agent, it wasn't like uh here's the use
cases that actually pass all the eval
focus on those use cases in the
marketing, but like I guess the user can
do other things too if they want to
because it's just like a LM, right? Is
that Yeah.
>> I I I I think the the the way the way
that the way to think about it is like
>> it's it's almost like what's the biggest
problem in applied AI right now, right?
The biggest problem is not the agents
aren't smart enough, right? The the
problem is not the models are not
advanced enough. The problem is uh
there's like you know people talk about
like a capability overhang or like a
capacity overhang or something like that
which is like the models are really
smart but we're just not using them
enough. And and so like where where
we're all the evalu that we have are
focused on they're focused on like are
you uh almost like you know when when
the user says something they actually
want to accomplish a task. Did you did
you figure that out? Right? Because you
could here's a here's an opportunity you
can really help them accomplish
something like you know did you did you
understand that that's what they wanted
you to do or like were you a little bit
too eager and you went on and and and
did something like that was way too
expensive and annoying when the user
didn't want that. Right? Like so so we
we have things where um you know for
example in the Slack integration when
like you're having a conversation with
the agent and you ask a follow-up
question, right? And the agent thinks
they can answer it. There's a whole
internal process that goes through like
I think I can answer this question.
Should I interject and and and that's
that's the so the events are are really
focused around those kinds of like
ergonomic type of uh you know type of
moments
>> and after the product is live or even
during dog fooding uh is there some sort
of feedback loop after the agent
response I can do a thumbs up or thumbs
down or something like some feedback so
you can get con constant feedback.
>> Yeah. Yeah. Totally. And that that's
that's been very useful, right? A lot of
our evals are effectively derived from
those moments where someone goes like
this behaved in a weird way or a stupid
way like let me tell you why. And then
and then eventually that itself becomes
an eval.
>> Um and Jacob, you probably know like a
couple of these like the the funnier
earlier ones, right?
>> Yeah. Yeah. We've had we've had a ton of
these. Like we had we had a user like
call the agent dude once and then the
agent was like, "Oh, okay. I'm not going
to respond to you because that was like
you're not being formal with me. Um, so
we we've had a lot of interesting use
cases where we've had to really just
like dial these things into a really
really narrow zone. I feel like um like
you don't have to show the prompt, but
like I feel like when you write the
prompts and skills, you got to be a
little bit maybe it's more line around
principles and how they should think
versus like hey you should make sure
this is 140 characters longer like very
specific kind of right like you don't
want to restrict it too much right in
terms of what it can do.
>> Yeah, totally. And you also just want to
like give it as little instruction as
possible to be honest. Give it the tools
to load context. Don't give it context I
think is like a a principle we found
important.
>> Interesting. Because if you just give it
too many instructions, it'll just like
overfit on software.
>> Yeah. And it may not need that. And then
it may overemphasize on certain things
that actually aren't important for that
task. Um they're they're just smart
enough to get what they need if you give
it a really good defined goal.
>> Oh. because they can just like do
searches and stuff and figure out
themselves.
>> Yeah. Give it tools to load skills, give
it tools to load guidance, you know,
different things like that and it can
build its own context.
>> Got it. Interesting. It's kind of funny
like it's kind of because you probably
get a bunch of feedback about the agent
and then you probably have a something
the agent ingest the feedback and
synthesize it. So it's almost like the
agent's improving itself, right? It's
like a loop. [laughter]
Yeah, we have a mechanism where the
agent can report um essentially
functionality it can't do. So if the
user asks it to do something it doesn't
have a tool for it, it's going to call a
tool and report that back to us and then
we autoingjust that into you know go
check if there's already an issue for
this. If so add it there. If not create
a new issue. So we have like a
constantly streaming
>> nice
>> set of issues coming in around what we
can do.
>> Dude, do you remember like offhand like
what's the craziest thing the user like
a user asks agent to do? crazy things.
Um I I think in general like we're
actually okay with the user asking, you
know, as long as there's not like safety
concerns or like that sort of thing.
We're okay with if you want the agent to
write you a poem, that's fine. Let it
write you a poem. Um so we we give it
quite a bit of freedom in that regard.
Um I think trying to lock it down too
much can can end up in a frustrating
situation.
>> Got it. Okay. I I think when it when it
comes to like you know the the range of
things that people ask linear agent to
do it's it's interesting right because
it because it's like explicitly like for
work right it's associated with your
with your workspace and your your
development team and stuff like that
like people don't tend to go on wild
adventures but there are things which
surprise us right like uh a lot of
people use it for translation so they'll
they'll get feedback from customers or
something like that in in a language
they don't speak and they'll just
they'll just straight up just ask it for
translation or even set up an automation
to be like if something ever you know we
we have a lot of customers in France. So
if we ever get a a intercom ticket that
comes in in French, just translate it
for me before you when you file an issue
against it or something like that,
right? So like there's a lot of those
sorts of uh creative sort of use cases.
They're very on topic, you know, but
like we never thought that that would be
a thing that people would do
necessarily, right? Like that that was
they didn't even cross their minds that
that was a a possible uh a possible
thing.
>> Got it. Okay. Yeah. Yeah. Yeah. I I
think just like put the product in
people's hands and then they'll figure
out like new use cases will appear and
they can figure out which part to
improve.
>> Yeah. All right. Well, let me ask you
like some product questions about the
agent then. So like I I think before
this before linear agent came out,
linear was a platform for like you you
can tag like cursor agent and some other
agents on here, right? What was the
philosophy behind actually kind of going
off and building your own agent?
>> Yeah, I I think the um you know
assigning issues to agents was like kind
of what you're referring to, right? And
then you can and you can add mention
them in in in like comments and stuff
like that. Uh and I I think that we
wanted to be able to support the entire
software development life cycle intent
like ultimately that's what you know
that that that's where we we saw our own
sort of expertise right in in terms of
how we can have very good opinions that
you can adopt for your for your team. Um
and uh and you know having intelligence
and having an AI system uh there to take
action was a way to actually uh execute
on those things right because we we you
know before like if you think about like
the the you know the the olden days like
we would like write you know guides
right we'd write like something like the
linear method about how you how you
ought to think about your software
development process and uh so a lot of
in you know what we saw uh with uh you
know with with models capabilities right
is is basically the ability for us to be
like actually you don't just have to
read it and then execute on this
playbook that we're giving you. You can
just have linear itself execute the
playbook. So like I think ultimately
that's that's where that's what we saw.
So so you know specific agents that we
add to the system they'll they'll do
parts of that. They'll write the code or
maybe they'll they'll do the root cause
analysis or something like that for for
bug reports. Uh but you know none of
them really took the whole process end
to end like we wanted to.
>> Yeah, that makes sense. I I I I always
think like the the AI handles the middle
80% and then the humans in the first 10%
and last 10% and maybe maybe we're
almost at or maybe we're already at a
point where you know we we can just tell
we can tell linear agent like here's
some user problems here's like a high
level idea of a solution like go figure
out and then go put a PR up and then
I'll review it even for like bigger
features right
>> yeah yeah I mean what you're what you're
talking about like I I've heard people
say something like AI is a it's not an
endto-end solution it's a middle's
middle solution right it's kind of what
you're what you're talking about and and
and and you know right now like maybe
humans handle the first 10% last 10% but
like at some point maybe it's like the
first 0.1% and the last touch right it's
like you know the middle just get it's
going to get bigger until it basically
reaches the limits of the edges right so
like that's kind of what we're also uh
seeing right like we we're assuming that
that's going to happen so we want to
build like a system right that
facilitates that
>> I don't know dude so so so Jacob have
you like have you been following
discords around loops and goals and
stuff like that are you like doing all
that stuff you let it run.
>> Yeah, we're definitely experimenting
with that internally. I mean, there's
definitely cost trade-offs you have to
consider once you you get into these
long runninging agents. Um, and I but I
think models are getting really close to
where that's a really interesting thing.
Yeah, on one hand like not exactly like
I said like if you upload a linear
method in the markdown file and then the
agent will actually read through it and
like actually try to follow it, right?
But on the other hand, I feel like it's
so good at producing markdown files and
all this stuff by itself. And then
usually it's pretty long. Sometimes I I
don't even read it anymore. I I like,
[laughter] okay, you make a plan, go for
it. Like just just do it. And then and
then I look at the final output.
>> And I'm I'm like really paranoid that
like it'll just end up as like, you
know, as more and more agent produced
markdown files end up in the re repo and
like I don't read any of it,
>> it just turns to slop, dude.
That that's what happened.
>> Yeah. Yeah. And I I I I think like like
you know as people use AI systems they
they they start feeling these things as
well. And then like you know a lot of
times people oh you know skill issue or
something like that they were like well
you should just read your things you
should like tell it not to produce slap.
like yes but like like we it comes back
to the main problem of like what are the
defaults and and I think one of the one
of the opportunities we saw for you know
for introducing AI into linear as a
system rather because like there's an
there's an alternate way of the universe
where we just say like we have an MCP uh
you know MCP server and they just use
whatever to to connect to linear and
then that's it right like we don't have
like a native agent doing anything
>> um but I think having a native agent
lets us uh you know imbue our you know
opinions about what good product
management looks like, right? And what
it looks like is not producing endless
markdown documents that you that are
unreadable and full of extraneous
detail, right? Like like where you know
we can we can we can make some
determinations about uh how the process
ought to run in the in the best case.
>> I see. I see that that's a really good
Yeah, because if you just build an MCP
just a bunch of tools that people can
use then they might go off the rails but
like the agent actually has a bunch of
skills and instructions to kind of imbue
linear's values and product process
right
>> exactly
>> that makes sense yeah cool when you guys
decided this thing was ready to ship
like kari has high bar right so when you
guys and this thing like people are
using it for all kinds of stuff so how
do you decide if this thing is ready to
go
>> what was it like from the engineering
side Jacob I I could talk about the
product side but what was the what did
it look like from from the back end.
>> It felt like we released it early, not
too early, but I think we kind of leaned
on the side of like it's pretty good.
Let's put it out there and get more data
for how to improve it more. I think we
could have worked on it internally
forever honestly to make it perfect. Um
it with something subjective like this,
it's it's harder to get to that like
perfect spot. And so I think we hit a
point where like we just need to ship
it.
>> Yeah, got it. How about you not from the
pro side? Yeah, I I think from the
product side, you know,
if you if you think about shipping
products that have like that have like a
UI, right, the UI effectively limits
what you can do with it, right? It's
like a very UIdriven feature or
whatever. And uh so you can have some
definition of quality that's based on
eliminating everything that you didn't
intend in the first place. I think for
something like this, we had like core
hero use cases that were like we these
are the things we're going to demo.
These are things we think are going to
add a lot of value and they're a good
like first way for people to use the
agent,
>> right? And if we make those rock solid,
then the other stuff it's like look,
there's going to be a variance in in
reliability and things like that because
it's just the nature of this kind of uh
you know, this kind of tool and this
kind of technology and we're okay with
that, right? But as long as like the the
the hot paths uh that that we're
advocating for are are you know are
solid and we believe in them, then I
think that that's the bar for quality
that we're looking at.
>> Okay. the stuff that you know like
create a ticket, manage tickets, create
a PR like that kind of stuff like that
that you highlight in the market
marketing as long as those are good.
>> Yeah. Yeah. Exactly. Because those are
things are like like those are you know
those have those are in the warranty so
to speak, right?
>> Yeah. Got it. Okay, that makes sense.
Okay. So, so I guess just to like kind
of wrap up a little bit, I'm I'm sure a
lot of companies are thinking through
this right now. Should they just build
like an MCP or should they build a
native agent or like you know if you
build all this stuff then people don't
even use your website anymore. Then like
what do you do? Like do you have any
advice for uh builders or companies or
think about whether just to even build
their own agent or not?
>> Yeah, I mean my advice is like like what
what you know really you have to really
break down what is the actual workflow
that that your users want to do, right?
Like and and a lot of times it's it's
like super multi-step, right? Like no
one no one like sits down at their desk
and oneshots their whole job. Like that
that's not that doesn't happen, right?
So it it's it's it's an entire process
that goes into end and then like where
are the natural places where uh where
you want to hook into that where where
it makes sense to uh to hook into that
and and I think that that's really where
um you know like we said like you know
linear agent as a as a collection of
as a collection of subsystems right like
a lot of them are about figuring out
where the right entry points are because
the the most obvious thing is like look
there's a there's an inapp chatbot you
can have a chatbot right and then People
will point at that be like, "Oh, that's
the agent." It's like, "Well, that's
that's that's one way to interact with
the agent." It's necessary because if
you want to do any kind of like
follow-ups or you want to do any sort of
like multi-turn processes, you have to
have some surface to to do that. But
that's not where the entry points are.
The entry points are in the discussion
that you're having in Slack or in your
meeting debrief or when you're writing a
project update and you're trying to do
the research, right? Those are the those
are the the on-ramps
>> to like utilize the intelligence. So if
you if you give people good on-ramps,
right, the the uh the sort of
interactive chat, that's that's the
follow-up, right? That and that covers
like the long tale of things that people
want to do. So I I think that's what you
kind of have to do for for um for domain
specific agents, right? Like because
otherwise you you're deal with this
problem, which is like, you know, why
wouldn't I use quad or chatbt or
something like that for this uh instead
of your sort of native agent?
>> Yeah. So it's almost Yeah. I I think
like an agent is almost like a employee
and then like employee doesn't only work
in one app, right? You should talk to
them from like, you know, Slack or
wherever you guys work, you know?
>> Yeah. I I think I think everyone's kind
of moving that direction now, right,
with especially with the latest releases
uh that that everyone's kind of putting
out.
>> Do you think Kyrie like chef chat has a
tier when like um if if like everyone's
using linear through the through the
agent or the MCP instead of like all all
the beautiful buttons and the UI that
exists?
>> Honestly, no. I I I I don't I don't
think I don't think he minds at uh minds
at all. And like I think you know linear
is like also very um naturally like a
multiplayer system. So different players
are going to use it in different ways,
right? Like we we have like for example
a lot of uh you know customer support
agents, right? People right who are
doing customer support will um use
linear by escalating things out of
Zenesk or intercom, right? And like
that's their entire linear surface usage
is is there's there's a there's some
controls in the in the plug-in at
intercom for them to escalate and pick a
template or describe the issue or
whatever it is, right? Like that's and
that's it. that's their linear usage,
but it's it's very valuable because
that's the input stream for everyone
else to actually do their work.
>> Yeah, that's a good point. Yeah. So, I
guess because Kari talks about being
like opinionated about the product,
right? But like I guess you need to let
people use it from what whatever
workflow or service they they want, you
know? You can't be too opinionated.
Yeah.
>> Yeah. Yeah. Exactly.
>> Yeah. Cool. All right, guys. Well, I
mean, uh I guess what what what's next
for Linear Agent? Where can people want
to learn more about it?
uh what's next is you know I think we're
we're we're definitely introducing some
aspects of proactivity and uh and and
sort of longer running memory right
those are those are the two areas where
you know we're really kind of focused on
um if you think about like uh like let's
say you're you're you're building like a
project at your company and that project
can last it could be a very short
project over in a few days or it could
like last a whole quarter and you know
the agent should be very well aware of
everything that happened throughout the
lifetime of that project and be able to
kind of naturally push it forward,
right? There's a lot of different
moments where you have to coordinate
people, you have to make sure that, you
know, documents are kept up to date and
all those kinds of things. And that's
where we we want that to just be
something you can take for granted.
>> Yeah. Like I I run a project in linear.
I can take for granted that it's
wellrun, right? Like that's that's the
that's the goal that we're looking for.
>> That makes sense. Yeah. Whenever I have
like a long Slack thread with an
engineer and they're like, "Okay, let's
go up to the PRD." I'm like, "I'm just
too lazy to update the PRD. [laughter]
I don't up to the PRD." So yeah, just be
able to assign it to linear. That'll be
super useful. Yeah. Cool. All right,
guys. Well, thanks so much, man. Thanks
for giving us an inside look at how
linear agent works. And uh yeah, I I I
wish you guys the best of luck. I think
it's a very interesting problem to
solve. Thanks, Peter. Thanks.
Ask follow-up questions or revisit key timestamps.
This video features a deep dive with Naan and Jacob from Linear on building and evolving their 'Linear Agent.' They discuss the agent's architecture—calling an LLM in a loop with tools—and how it has evolved from managing simple tasks like creating tickets to acting as a versatile agent capable of end-to-end product development. The conversation highlights their philosophy of using 'skills' for specialized tasks, the importance of providing good entry points (on-ramps) for users rather than just relying on a chatbot, and the shift towards proactive, long-term project management.
Videos recently processed by our community