HomeVideos

Factory's Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself

Now Playing

Factory's Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself

Transcript

1737 segments

0:00

Bezos at Amazon it's customer obsession.

0:02

But in our mind that's an input metric.

0:04

Like you don't want to measure input

0:05

metrics. It doesn't matter if you're

0:07

customer obsessed. Like you could be

0:08

customer obsessed and they file a

0:09

restraining order against you cuz they

0:11

don't like what it is that you're doing.

0:13

Like our job is to build something so

0:15

good that our customers themselves

0:17

become obsessed with us. That is our

0:18

job. It's like you know, the analogy is

0:21

if you're a coach of a basketball team,

0:23

you don't want to tell your players

0:24

before they come out there like, "Hey

0:25

guys, make sure to sweat." It's like,

0:27

"What?" Like no, like score points. Like

0:29

we need to score points. And in doing

0:31

so, yeah, you're probably going to

0:32

sweat. And I think similarly to create

0:34

obsessed customers, you probably need to

0:36

be really obsessed yourself with the

0:38

customers, but the output is what

0:40

matters.

0:46

>> [music]

0:58

>> We're here in the studio with Matan from

1:00

Factory. This is our second time with

1:03

Matan.

1:04

>> Thanks for having me.

1:04

>> You're in the small and elite group of

1:06

second time training data attendees, so

1:08

thank you.

1:09

>> Matan is the co-founder and CEO of

1:11

Factory, which makes droids, which are

1:14

autonomous agents for the art of

1:16

software development.

1:17

>> Yes, indeed.

1:17

>> And Matan, we're going to jump right in

1:19

because I think you guys are a little

1:20

bit of a dark horse candidate in this

1:22

world of software development. It is a

1:24

market that is absolutely taken off.

1:26

There are folks like Claude Code and

1:28

Cognition and others who who have

1:30

but you guys are coming up strong. Talk

1:33

about the competitive dynamics and what

1:35

makes Factory special.

1:36

>> It's been a wild ride. We started

1:38

Factory three three and a half years ago

1:39

now. So, in April of 2023, when the

1:43

world and the enterprise in particular

1:45

was barely ready for GitHub Copilot, let

1:49

alone fully autonomous agents. And so,

1:51

um I think the first two years it was

1:53

kind of

1:55

our journey in the desert is is how I

1:57

like to refer to it because we were

1:58

focused on fully autonomous agents, but

2:01

engineers weren't ready, procurement

2:03

teams at the enterprise weren't ready,

2:04

and so

2:06

I think retrospectively we really like

2:09

honed our craft and learned a lot about

2:11

how to build for developers in the

2:14

enterprise, but um

2:16

you know, it it it took a lot of time to

2:18

actually come around to when they were

2:20

ready to receive it. And so we're kind

2:22

of now emerging much more in some of

2:24

these other players like Anthropic or

2:26

OpenAI who have a ton of distribution

2:28

um are going in and you know, bringing

2:30

their incredible tools like Claude Code

2:31

or Codex.

2:33

The thing that enterprises are really

2:35

caring about that we have learned

2:36

through those two years is they do not

2:39

want anyone to kind of be their single

2:41

point of failure. They do not want

2:43

anyone to kind of control their fate.

2:46

And so something that really matters is

2:48

model independence.

2:50

Everyone learned from cloud where, you

2:53

know, back in the cloud days it was like

2:55

AWS or or Azure being like, "Hey, you

2:57

know, come on in, sign this three-year

2:59

contract. It's going to be so cheap.

3:01

We're going to subsidize it. It'll be

3:02

great."

3:04

And then a couple years later when it

3:05

came time to renewal, they would 10x the

3:07

the the contract.

3:08

>> Haha, didn't gravity. We got you now.

3:10

>> got you. What are you going to do a

3:11

two-year migration to go to someone

3:13

else? Like no way.

3:14

Everyone has scars from that now. And so

3:16

everyone knows, look, Claude Code is

3:18

fantastic. Uh Codex from OpenAI is

3:20

fantastic. We cannot put our fate in any

3:23

one of these model providers' hands.

3:26

Also, like you just look at the risk

3:28

profiles of the model labs versus the

3:30

cloud providers. What's the last piece

3:32

of drama that came out of a one of the

3:35

cloud providers versus like the model

3:37

labs, it seems like there's kind of

3:38

always some sort of chaos of, you know,

3:40

internal fighting or getting in spats

3:41

with the government or, you know, any

3:43

other entities. And so uh if you're

3:46

going to, you know,

3:47

build this very important part of your

3:49

business, you want to make sure that

3:50

you're robust to any of these changes.

3:53

And that's something

3:54

that we've learned over those kind of

3:55

initial 2 years is like

3:57

developers really care about things

3:59

being modular. They want to know that

4:01

they can customize it to what they want.

4:03

They want to know that if there's a new

4:04

model that comes out that's faster or

4:06

cheaper or more performant, they can

4:08

kind of hot swap it in. And that's I

4:10

think one of the biggest reasons why a

4:12

lot of the largest enterprises are

4:13

taking the momentum that they had from a

4:15

Codex or Cloud Code and then are

4:17

carrying that into Factory because they

4:20

get that performance from these

4:21

fantastic models, but they do it without

4:23

the vendor lock-in that, you know, the

4:25

Model Ops direct model provides.

4:26

>> if I'm the enterprise, I'm going to be

4:27

like, "Wait a minute, am I not just

4:29

getting locked into Factory?" What's the

4:30

answer to that?

4:31

>> So, the good It's a really good question

4:32

cuz that is something that you might

4:34

think of like, "Okay, wait. So, we're

4:35

just switching the the lock-in point."

4:37

All of the modularity that we build is

4:39

such that if at some point you wanted to

4:40

say, "Hey, you know what? Factory's not

4:42

staying at the frontier anymore."

4:43

Whether it's like the automations that

4:45

you build or the skills registry that we

4:47

help you create, the work that we've

4:49

done stays in your code base and any of

4:51

the automations that we've created, the

4:53

artifacts also live in your code base.

4:55

In other words, there aren't really

4:57

things that we're saying like

5:01

our tribal knowledge about your org that

5:02

we're keeping on our side and not giving

5:04

to you. Um and that's part of the

5:06

relationship that we have with customers

5:07

is like we similarly want to make sure

5:09

we're providing the best experience

5:11

possible. If we help you arbitrage

5:14

between different models to get cost

5:16

optimization, we're giving you that

5:18

optimization. We're not taking that away

5:20

from you. Um and I think that's a really

5:21

important part of the trust that we're

5:22

building with with these enterprises.

5:24

>> You and I were talking probably a couple

5:26

months ago at this point and I was

5:29

trying to give you credit for having the

5:30

right vision for this market 2 3 years

5:32

ago and you responded with something

5:34

along the lines of

5:36

thank you, but being 2 3 2 or 3 years

5:39

early is the same as being wrong.

5:40

>> Yes.

5:41

>> Um which I thought that was a wonderful

5:43

response in so many ways. Can you talk

5:45

about like that those 2 years in the

5:47

desert? How did it feel to have this

5:50

vision that turned out to be right that

5:51

nobody appreciated for a year or two?

5:54

Can you just talk about like that

5:55

journey and what it has done to the DNA

5:57

of of your company?

5:58

>> Yeah, I mean, in the moment it's really

6:00

really difficult because, you know,

6:04

I hadn't had a job before. I dropped out

6:06

of my PhD to start this company and, you

6:08

know, over the course of those two years

6:09

convinced, you know, 20 of the smartest

6:11

people that I've ever met to quit what

6:14

it was that they were doing and, you

6:16

know, join Factory and join us on this

6:17

mission.

6:19

And these are people with families.

6:20

These are people with kids who are like

6:22

dedicating years of their lives to this

6:25

problem and going, you know, customer

6:28

after customer and they like they

6:30

weren't ready for agents. They didn't

6:31

get it. Also, the models weren't as

6:33

performant, but I think a lot of it was

6:34

behavioral. And I mean, even just a fun

6:37

anecdote of like giving developers an

6:39

NPS survey. If you ever are giving a

6:42

developer an NPS survey, they do not

6:45

like whatever it is that you're giving

6:46

it to them. Because like developers,

6:48

they vote with their feet. They are very

6:50

clear what they like and what they don't

6:51

like. And if you're like, "Hmm, I wonder

6:53

if they like it?" They definitely don't.

6:55

Um and uh

6:57

but during that time, I think there were

6:58

there was a lot that we were learning.

7:00

There was a lot that I myself was like

7:02

I'd never had a job before. Enterprise

7:04

sales is not something that comes

7:05

obvious to to a physicist. Um

7:08

but

7:09

at the end of the day, it doesn't it

7:11

doesn't matter. There's no you don't get

7:12

any, you know, bonus points for being

7:15

early

7:16

because like who cares? Like there's no

7:18

consolation prize. It's either you do

7:20

the thing or you don't do the thing and

7:22

that's all that matters. And for the

7:24

team, it's really tough. There were

7:25

points where

7:26

uh

7:27

we ended up getting good at enterprise

7:29

sales, but the product still wasn't

7:30

good.

7:31

And that's a very tricky position to be

7:33

in because we ended up, you know,

7:36

getting to a point where we were like

7:38

just under 2 million in revenue and the

7:40

product was not good.

7:41

And there was a point in time where we

7:43

realized this. Cuz if you're really good

7:44

at sales, you can sign contracts. That's

7:46

like you can definitely do that.

7:48

But if you're doing that and the

7:49

developers don't like your product, it's

7:50

like a ticking time bomb. Because it's

7:53

eventually they're going to turn and

7:54

it's going to be really, really bad.

7:56

We realized this

7:57

and we proactively gave all of those

8:00

customers their money back. And I

8:02

remember that uh

8:04

was one of the most difficult

8:07

decisions to make. Because not only is

8:09

there a uh you know, a group of you

8:11

know, 20 people who are getting

8:12

ridiculous offers from all the labs.

8:15

They have these huge, you know,

8:15

financial incentives to go elsewhere.

8:17

There are all these other companies that

8:18

are doing well and they decided to do

8:19

this.

8:20

And then we're going to say, "Oh yeah,

8:21

hey, by the way, that, you know, little

8:23

bit of revenue we managed to get, we're

8:24

actually going to give it back cuz we

8:25

don't think product is making their

8:26

developers happy." We also had to

8:29

>> you make that decision?

8:30

>> You know, we

8:32

sold them on a good vision

8:34

and convinced them that, you know, this

8:35

is the the right team to work with and

8:37

that we were going to deliver the

8:38

solution for them. But we realized that

8:40

the way that we had sold them on it and

8:43

the product that we were delivering was

8:44

not up to snuff in a way that I don't

8:47

think it would hold true to one of our

8:48

operating principles. And one of our

8:50

operating principles that I really like

8:52

is create obsessed customers.

8:54

>> Yeah.

8:54

>> This kind of like flips over Bezos'

8:57

thing. Where Bezos at Amazon it's

8:58

customer obsession. But in our mind,

9:00

that's an input metric.

9:02

And like input metrics are like you

9:03

don't want to measure input metrics.

9:05

It doesn't matter if you're customer

9:06

obsessed. Like you could be customer

9:08

obsessed and they file a restraining

9:09

order against you.

9:10

>> [laughter]

9:10

>> Cuz they don't like what it is that

9:11

you're doing. Like our job is to build

9:13

something so good that our customers

9:15

themselves become obsessed with us.

9:17

That is our job. It's like, you know,

9:19

the analogy is

9:20

if you're a coach of a basketball team

9:22

you don't want to tell your players

9:24

before they come out there like, "Hey

9:25

guys, make sure to sweat."

9:27

It's like, "What?" Like, "No, like score

9:28

points. Like we need to score points."

9:30

And in doing so, yeah, you're probably

9:31

going to sweat. And I think similarly,

9:33

to create obsessed customers, you

9:35

probably need to be really obsessed

9:37

yourself with the customers,

9:39

but the output is what matters. And I

9:41

think that it coming back to this,

9:44

the product that we were delivering was

9:45

not creating obsessed customers. And we

9:47

wanted to make sure like this was a

9:49

group of the smartest people I've ever

9:50

met. We were getting there. Like we were

9:52

getting a lot of intuition. Things were

9:54

starting to come together internally.

9:56

Like we could see internally we were

9:57

starting to become a lot more

10:00

agent native in how we were doing things

10:02

in the product was kind of scratching

10:03

that itch.

10:05

But we were kind of ahead of our

10:06

customers and we wanted to maintain

10:08

trust with our customers so that when it

10:11

does hit, we can come back to them and

10:13

say, "Hey guys, this is the real deal. I

10:14

promise." And to build that credibility,

10:16

we had to say, "Hey look,

10:18

you know, even though you were maybe

10:19

happy to continue,

10:20

we're going to give you this back and

10:21

say, "Three months from now, I think

10:23

it'll be ready. Give us some time and I

10:25

promise we will knock your socks off."

10:27

>> How did your customers react when you

10:28

had that conversation?

10:30

>> Some of them were like, "Oh great.

10:31

Sounds good." cuz I think it wasn't

10:33

something that they, you know, were

10:34

obsessed with.

10:36

Some of them were a little bit confused.

10:37

Um but I think gen- generally it's

10:40

especially enterprises, they're not used

10:41

to these things. A lot of times

10:42

enterprise budget, once it's gone it's

10:44

gone and no one really cares.

10:45

>> Yeah.

10:46

>> And so some of them didn't even know if

10:47

they had a mechanism by which to take

10:49

back money.

10:49

>> [laughter]

10:50

>> Um but you know, it's a difficult thing

10:53

to tell also like investors who believe

10:54

in you. Like, you know, I remember

10:56

having the conversation with Shawn. Um

10:58

I think Shawn obviously he's stayed

11:00

really close with the company so he was

11:01

very like on the same page. But it's

11:03

kind of a scary thing. Be like, "Hey by

11:05

the way, you know, remember all those

11:06

updates and you're saying, 'Hey look,

11:07

the you know, revenue's going up. It's

11:09

about to go down to zero.'" Um it was a

11:12

scary thing. Uh and I think it was kind

11:14

of a leap of faith of like

11:16

we see the signal internally early of

11:19

like this is the direction we need to

11:20

go. We need to kind of pivot the

11:21

approach on the product. But I remember

11:24

that all hands where we told the whole

11:25

team,

11:27

it was like, oh my god. That was like

11:28

one of the worst months of my life. Like

11:30

I was just

11:31

cuz no like

11:32

not everyone was going to say like,

11:34

"What the hell is this? What's going

11:35

on?" But it's kind of the looks on their

11:36

faces where they kind of go a little bit

11:38

pale and they're like, "Oh boy, like is

11:40

this just the early signs and we're

11:41

about to sink completely?" Um

11:43

>> How did you keep the team together

11:44

through that?

11:46

>> I think honestly

11:48

the only reason the team stayed together

11:50

is we were so ruthless about hiring

11:52

early on where it was like

11:54

people that are genuinely really, really

11:56

obsessed with the mission, which our

11:58

mission is to bring autonomy to software

12:00

engineering. And like really, really

12:02

caring about that, making sure everyone

12:05

um

12:06

was also like very clear feedback loops

12:09

as to like

12:10

this the fate is in our hands. It's not

12:13

like this is like, "Oh, something that I

12:14

go do." It's like we all have a part to

12:16

play in, you know, making this work. And

12:17

I think embracing how much it sucked was

12:20

also I think something that was very

12:21

valuable.

12:22

>> Just being honest about it.

12:23

>> Being super honest about like, "Yeah,

12:24

this sucks. Like you know, look look at

12:25

those competitors, their revenue's going

12:27

up like crazy. Like this is not good.

12:29

Like we're in a very bad position. Like

12:31

we just have to give back all of our

12:32

revenue. Like we need to really get our

12:34

together. Um

12:36

and in the moment, I think

12:38

retrospectively, those are the moments

12:39

where really the deepest bonds are made.

12:42

Like if you talk to people who were like

12:44

athletes or even like

12:46

academics or whatever, whenever you're

12:48

in the like stressful period, whether

12:50

it's like cramming before finals or you

12:52

know

12:53

an intense like, you know, we have some

12:55

some rowers on our team and I think

12:56

that's an example we always go to. Like

12:58

that sport sucks.

12:59

>> sport.

13:00

>> a pain sport.

13:00

>> It's a pain. It's literally just there's

13:02

one number that quantifies your

13:03

performance. It's just what is your time

13:05

on your 2K or your time in that. Um

13:08

but like embracing that is what creates

13:10

those enduring bonds such that

13:12

afterwards like

13:13

we know what it's like to be at rock

13:15

bottom. We know what it's like to lose.

13:17

We know what it's like, I mean, when we

13:18

first started the company, our valuation

13:20

was 5 million.

13:22

Like a lot of our competitors, a lot of

13:24

the companies out there these days, they

13:25

don't know what it's like to not be a

13:27

unicorn. That's like manifestly, that is

13:29

what they are day one.

13:30

Whereas like we have been there kind of

13:32

in those dark moments and not a single

13:34

person left.

13:35

>> Yeah.

13:35

>> That makes us so resilient and so strong

13:38

that, you know, going forward, things

13:40

are going a lot better now.

13:42

But there are going to be really bad

13:43

times, but we have that resiliency in

13:45

our DNA that I'm not sure some of these

13:47

other companies do.

13:48

>> I love that.

13:49

Um so talk us talk to us about what

13:51

changed. And I'm curious your comment

13:53

from earlier that the models getting

13:55

better is not important thing that

13:57

happened. Cuz at least in my mind, the

13:58

model getting better is the most

13:59

important thing that happened. So just

14:00

help me understand.

14:01

>> Yeah, so so a couple things. So one is

14:03

the interaction pattern that we were

14:05

building for before was too ambitious.

14:07

Like to your point, we were right in

14:09

that what we were building for was fully

14:11

autonomous agents, but it was 2 years

14:13

too early, which makes it wrong. And

14:15

fully autonomous agents require a

14:17

complete change in behavior from the

14:18

developer. And we were trying to do that

14:20

out of the box before they were even

14:21

using tools like Copilot.

14:23

It was just too much of a leap. It was

14:24

too much of a step function jump. So

14:27

it's an important day. September 26th,

14:29

2025 was when we first put out um

14:32

basically the the Droid CLI. And the

14:35

Droid CLI met developers where they were

14:38

in a manner that previously these fully

14:40

autonomous agents did not. Um and also

14:43

its performance was like completely

14:44

state of the art and it was model

14:45

agnostic. So it could use every model

14:47

that was out there. September 26th was

14:49

also 2 years after we initially started.

14:52

So the world had gotten much more used

14:54

to using things like auto complete. Like

14:56

by by late 2025, most engineers were

14:59

using an auto complete tool, and many

15:02

were starting to

15:04

um

15:04

at the time use like a chat interface to

15:07

ask an agent to go do changes like

15:09

wholesale. So like the more agentic

15:10

interaction. However,

15:13

what we see is that like if you go back

15:15

now and use in this like agentic

15:17

interaction some of these older models,

15:19

they're still good.

15:20

So the biggest thing that changed was

15:23

developers, and in particular in the

15:25

enterprise, like being open-minded to

15:27

this new way of working. In particular,

15:29

you know, developers,

15:31

they've established their workflows over

15:32

the last 30 years. They can be stubborn.

15:35

A lot of them were like, "No, no, no,

15:36

like my craft could never be done by an,

15:39

you know, an AI tool." So, a lot of it

15:40

was just like understanding how to work

15:42

with these tools and having the

15:43

willingness to go in and try, and also

15:45

the intuition about what are the

15:46

guardrails that you need to provide in

15:48

order for it to succeed.

15:49

>> Yeah.

15:50

>> Um

15:51

And so, I think it was a combination of

15:52

both of these things. The model's

15:53

getting better, so you need to do less

15:55

in the way of providing guardrails, but

15:57

also developers lowering their guard and

15:59

being like, "Okay, you know what? Let me

16:00

go try and do these things. It's going

16:01

to go do things I don't like." And then

16:03

also,

16:05

there's a certain degree to which when

16:07

Andrej Karpathy tweets about something,

16:09

then every engineer suddenly is like,

16:11

"Okay, you know, maybe this is true."

16:12

And

16:14

Andrej started to tweet about these

16:15

agentic work. Early on, he wasn't as

16:17

open to it. And then him being more open

16:19

to it genuinely just changed some

16:21

people's minds, uh which is funny, but

16:23

that's some of the things that go into

16:24

behavior changes. Like,

16:26

you hear it from people you trust, you

16:28

start seeing it, you know, from people

16:30

within your organization who are maybe a

16:31

little bit more agent native, but that's

16:33

that's kind of these things together is

16:35

what what changed that.

16:36

>> And now we're all going to be on Slack.

16:37

But

16:38

>> We might we might be on Slack we might

16:40

be pushing the limits of Slack, which I

16:41

think is going to be another interesting

16:43

thing, but yeah.

16:44

>> Okay, so September 2025, you launched

16:46

the Droid CLI. You said frontier

16:48

performance soda. What does that mean

16:50

for you?

16:50

>> There's like the benchmarks, which have

16:52

a very short half-life. Like, anytime

16:54

there's a good benchmark, it gets

16:55

bench-maxed within like 3 to 6 months.

16:58

>> Yeah.

16:59

>> At the time, I think the one that we

17:01

kind of championed when we launched, and

17:03

kind of it ended up becoming a pretty

17:05

good benchmark, was Terminal Bench. So,

17:07

prior to that, the one that was kind of

17:08

leading was Sweet Bench, which was um

17:11

kind of took some open-source projects

17:13

and some examples of issues that were

17:14

then solved.

17:15

The problem with that was it was very

17:18

focused on like Python and like

17:20

scripting or like individual file

17:23

changes.

17:24

Whereas terminal bench was more when it

17:26

was in the terminal setting, so it was

17:27

things like

17:29

scheduling runs and things that were not

17:30

just like changing the code file, but

17:32

general software development tasks. And

17:34

that was something that we ended up you

17:36

know, having really frontier performance

17:37

on. Now, it's like

17:40

benchmarked

17:41

to the extreme to where it's like I

17:43

think, you know, models that come out

17:45

now are like 90% on it. And I think um

17:48

there's a very short time horizon from

17:50

putting out a good benchmark to then it

17:51

being kind of in the training data.

17:53

>> What goes into building a great and is

17:54

it that is the great harness? And it

17:56

seems like there's almost a lot of fud

17:57

in the ecosystem of my harness is better

17:59

than your harness and

18:01

you know, you need to own the model to

18:02

have a good harness or actually you you

18:03

have a better harness if you don't own

18:04

the model. Like what's your mental model

18:06

for for

18:07

>> Yeah.

18:07

>> you know, benchmark maxing aside, what

18:10

keeps you at the frontier?

18:11

>> Yeah, so a couple of so some general

18:13

things that matter are um the way you do

18:15

caching. So, you know, cash tokens end

18:18

up being like a tenth as expensive. And

18:20

so one big piece of performance for a

18:22

given harness is what what is your like

18:25

rate of of token caching. Um another

18:28

example would be how do you perform well

18:30

in compression or compaction?

18:33

So typically when you're dealing with a

18:34

a long session,

18:36

you're going to exceed the context limit

18:38

of the model itself. And so the harness

18:40

will do some sort of, you know,

18:41

summarization, compression, compaction,

18:43

whatever you want to call it. And the

18:45

way that you perform during that

18:47

compaction is a big determining factor

18:48

of how good your harness is.

18:50

Um and tests that they do for that are

18:51

like, you know, they call it needle in

18:53

the haystack where you have some long

18:55

thread and maybe there's one piece of

18:56

information that's really important. How

18:58

often will your harness preserve that

19:00

through compaction? Other examples are

19:02

like tool use or how does it use the

19:03

environment to validate whatever work

19:05

that it's doing? Um

19:07

these are things that you can kind of

19:09

have individual metrics on and that we

19:11

kind of have our own internal benchmarks

19:12

to measure how do the out of out of the

19:14

box agents do versus how does factory

19:16

perform? I think one thing that um

19:20

naively everyone believed initially was

19:23

if you train the model and you build the

19:24

harness, you're going to make them

19:25

better together.

19:26

>> Yeah.

19:27

>> And much to the chagrin of many of my

19:29

friends at OpenAI and Anthropic, this is

19:31

not true.

19:32

If you build a harness that supports

19:35

different models, that harness will be

19:37

better.

19:37

>> What's the like my intuition would be

19:39

model harness co-design makes you

19:41

better.

19:41

>> Yes.

19:42

>> What's the intuition for why it's

19:43

actually not?

19:43

>> It's very analogous to

19:46

the idea maybe like I don't know 10

19:49

years ago of if you were to be like,

19:51

"Hey, I want to train my personal AI

19:52

back in like like ML days before like

19:55

GPT-3, I want to train my personal AI,

19:57

I'm going to give it all of my data cuz

19:59

I want it to know me."

20:00

Turns out the answer was, "Train it on

20:01

the whole internet and it'll be so much

20:03

better for you than if it were just

20:05

trained on your data." So there's a sort

20:07

of analog that emerges where it's

20:09

what data is to a model,

20:12

models are to a harness.

20:15

Where the more models you expose to a

20:16

harness, you avoid over fitting that

20:18

harness to the nuances of that model in

20:20

particular.

20:22

And there are certain intricacies about

20:24

different models that you can learn from

20:26

and then improve different models

20:28

performance in your own harness. And

20:29

this was why for example, we kind of

20:31

stopped doing it because Terminal Bench

20:32

got so

20:33

uh bench maxed, but initially when like

20:36

every new Opus or GPT model would come

20:38

out, it would perform better on Terminal

20:40

Bench in droid than it would in Claude

20:43

Coder Codex. Um

20:45

which is what like and this is something

20:46

that you know

20:48

I think was somewhat frustrating to cuz

20:50

I deal like from a lab perspective you

20:52

ideally want it so that it's better

20:53

together because then it means you have

20:55

to use their harness and you can't use a

20:57

different one, but I think the reality

20:58

is it it it's uh

21:00

you know, having that

21:02

multi-model harness ends up getting kind

21:04

of frontier on on all of axes.

21:05

>> Is there a good like example or

21:07

illustration of that? Conceptually it

21:08

makes sense. Is there a like an easy way

21:10

to illustrate it?

21:11

>> Maybe maybe a good example of it is like

21:13

if you're familiar with the different

21:14

behaviors of

21:16

uh

21:17

Opus and GPT 5.6 right now.

21:19

>> I am. He's not.

21:20

>> Okay. Opus tends to

21:21

>> [laughter]

21:23

>> I mean, loosely loosely. I mean, to be

21:25

fair, honestly these days I'm not doing

21:26

it as much either. Uh but I will say

21:28

this. Loosely, Opus is kind of like that

21:32

super friendly colleague where you're

21:33

like, "Hey, I want to go do these 20

21:35

tasks." And they're like, "Okay, cool.

21:37

Hey, by the way, five of those tasks I

21:38

realized we didn't need to do it. Don't

21:40

worry about it. I got other of these

21:41

done. Did it this way.

21:43

>> not a good time. Let's pick it up in the

21:44

morning."

21:44

>> Yeah, like and let's go let's go get a

21:46

beer afterwards and hang out, whatever.

21:48

Meanwhile, like GPT 5.6 is like

21:51

absolutely, I will do every single one

21:52

of those and nothing will stop me. I'm

21:54

not going to sleep until those are It's

21:55

a kind of very OCD and you know,

21:57

meticulous. But sometimes you know, you

22:00

want one where it's like it actually

22:01

realizes, "Hey, that list of 20 that you

22:03

gave me actually here's a better way of

22:04

doing it anyway." You know, 5.6 is more

22:07

methodical. If you build a harness for

22:09

each of those, they're actually

22:10

different things that that harness will

22:12

then be good or bad at. So, for example,

22:15

um

22:15

one thing that, you know, typically

22:17

agents will do is they'll

22:18

they'll have a to-do list of like if you

22:20

have a task, it'll go and generate a

22:21

to-do list. Um and the Claude code

22:24

harness can in some cases or and this is

22:27

maybe less relevant now, but I think

22:29

earlier this is a just a more

22:30

illustrative example. Earlier, it was

22:32

really strict to make sure it would

22:35

stick to the to-do list cuz the model

22:37

itself would typically wander.

22:39

Meanwhile, Codex wouldn't do that

22:41

because the model itself was really

22:43

really OCD about that. But if you're a

22:45

user, you want to have the same

22:46

experience regardless. Like you want to

22:48

make sure if you switch to a different

22:49

model, you're not going to suddenly lose

22:51

track of whatever things that you are

22:52

working on. And so, there are certain

22:54

things where like

22:56

maybe in some cases you really want

22:57

robust tool use. And there are tools

23:00

that you use to do these to-do lists.

23:02

You want really robust tool use and you

23:04

want to make sure that no matter what if

23:05

I'm a user I want to see my to do list

23:07

there. Like there were some cases where

23:08

it would just like not have the to do

23:09

list. And so that these are things that

23:11

kind of improve the general performance

23:13

and that the to do list matters because

23:15

we're doing some crazy migration and you

23:17

don't have the to do list and then

23:18

you're in this long session where

23:19

there's compaction that might get lost

23:22

in the summarization and then now you

23:23

forgot what your seventh step was and

23:25

that could be one of the failure modes.

23:26

That's kind of an example of of how they

23:27

do it. Yeah.

23:28

>> Yeah, yeah, yeah.

23:29

>> Good example.

23:30

>> Okay, so we talked about one type of

23:31

maxing, benchmark maxing. Let's talk

23:32

about token maxing cuz it seems like the

23:34

world has changed a lot. We've gone from

23:36

token maxing to now cost

23:38

rationalization. What does that mean for

23:39

Vector?

23:40

>> Yeah, so um maybe I'll I'll lay this out

23:43

just to so we're all on the same page of

23:45

like the way that we see what what's led

23:47

us to this token maxing. So

23:49

loosely there was like this phase one

23:52

where maybe phase zero was like no one

23:54

believes in AI. Then phase one everyone

23:55

believes in AI and then boards were like

23:57

Mr. CEO, what are you doing about AI?

24:00

What's your AI strategy? And Mr. CEO is

24:02

like I don't know like what's our

24:04

AI strategy? CTO like make sure everyone

24:06

goes and uses AI. And so then phase two

24:09

is you know CTO is like okay we

24:11

got to make sure everyone uses AI. Let's

24:12

start putting it in performance reviews.

24:14

Let's make public like bench or public

24:16

like rankings of who's using tokens the

24:18

most cuz everyone's stubborn. No one

24:20

wants to use this stuff. They're all

24:21

skeptical.

24:22

And then we enter phase three which is

24:24

everyone sees these ratings. They see

24:26

that it's part of their perf reviews and

24:28

they're like okay, I'm going to use AI

24:29

for everything. And that's kind of phase

24:31

three. It's this token maxing where

24:35

people are using like Opus for literally

24:37

everything. Like what's the weather in

24:39

SF? Opus, tell me. I don't know. Like

24:41

there are banks that we are working with

24:44

where they are spending literally

24:46

hundreds of thousands of dollars a month

24:47

on people asking things like literally

24:50

what is the weather.

24:51

Or like tell me about Python. Like

24:53

trivial questions that you could Google,

24:55

people are asking Opus. Um

24:57

and the reality is this happened because

24:59

we were so worried about adoption that

25:00

we overcorrected and we're like adoption

25:02

by any means necessary.

25:04

And I think that's actually it's like a

25:05

decent approach. Like

25:07

it's probably faster to do that and then

25:09

curb uses or make usage more responsible

25:11

than it is to start limiting and be like

25:14

you know, you can only use it for this

25:15

thing. Cuz when you have people that are

25:16

stubborn, first you want to just prove

25:18

that it works and then you can get kind

25:20

of more mature about it. Where factory

25:22

fits in, I think one of the most

25:23

important things that we do is that we

25:25

have the factory router which allows you

25:28

to dynamically route to different models

25:30

based on the task that you're doing. So,

25:32

you know, if you're asking what the

25:33

weather is you probably don't need the

25:35

very frontier of human intelligence to

25:37

answer that for you.

25:38

>> Or or you really do.

25:39

>> I mean, it depend I don't know, it

25:41

depends on what kind of answer you're

25:42

looking for. Um you know, giving you

25:44

like a full like down to the like

25:46

molecular level of what's happening.

25:48

But um

25:49

uh you know, allowing that but also more

25:51

importantly for every enterprise,

25:53

something that no one's dealing with yet

25:55

but 12 months from now is going to be

25:56

the case is um

25:59

not everyone needs the same tokens.

26:02

Having a blanket kind of token cap for

26:05

every individual in some large bank,

26:07

let's say, makes no sense.

26:09

So, every CIO is going to need to answer

26:11

for every incremental token, where do we

26:13

put it? And right now it is super not

26:15

obvious how you would do that. Like

26:18

right now we're saying, oh you know, the

26:19

PMs who are like vibe coding dashboards

26:21

get the same token limits as like the

26:22

engineers who are building like critical

26:24

infrastructure. That's probably not the

26:25

best thing to do. Or similarly

26:28

you might be dealing with COBOL code

26:30

bases where Opus is not the best model

26:32

to use but instead maybe some fine-tuned

26:34

model on that code base in particular.

26:37

The point of the router is that we can

26:39

kind of accommodate these different

26:40

constraints where maybe you say, you

26:41

know what, this part of the org, they're

26:43

just vibe coding, they can use Gemini

26:44

Flash. This part of the org, they're

26:46

doing COBOL, we fine-tuned this great

26:49

model to to on COBOL, let's route to

26:51

that when we're working on that part of

26:52

the code base. Maybe this other part

26:55

we really care about reliability, so

26:57

let's generate the code with OpenAI,

26:59

test it with Anthropic, review it with

27:01

like Gemini.

27:03

Things like that. And we can actually

27:04

take in your routing procedure

27:06

instructions in natural language. So,

27:09

you could even say things like it's not

27:10

purely deterministic, it can even be

27:12

like hey, you know, Pat, I don't know,

27:14

like I don't know what he's doing, like

27:15

>> If Pat jumps in my flash

27:16

>> given flash, like I don't know.

27:18

>> [laughter]

27:19

>> Or, you know, I think we really need to

27:22

avoid having them use open models

27:24

because, you know, whatever reason we

27:25

don't like the way open models perform

27:27

here. And we'll do internal benchmarking

27:29

to know which models are better at which

27:31

of these tasks.

27:32

>> How close are the open models at this

27:33

point? Which one's the best?

27:35

>> GLM 5.2 is incredible.

27:37

It's at the point where internally we

27:38

have no token limits for our engineers.

27:41

And

27:42

like half of our tokens are open to open

27:44

models.

27:45

>> Wow.

27:45

>> Yeah.

27:46

Cuz they're just faster and they're

27:47

cheaper. They're just as performant. And

27:50

I think the thing that everyone gets

27:51

wrong is everyone is comparing

27:53

like GLM 5.2 to the latest model like

27:56

Opus 4.8 or GPT 5.6.

27:59

But really they should be compared to

28:00

Opus 4.7 or GPT 5.5. Um

28:04

>> Why?

28:05

>> Because generally the open models come

28:06

later and they're they're kind of a

28:07

generation behind.

28:09

And that's kind of the the frontier

28:11

models will be frontier. The question is

28:14

are the open models getting as good as

28:15

like frontier minus one?

28:17

And the answer is unequivocally yes.

28:19

Which I think is a really, really

28:21

interesting outcome. It's great for

28:23

consumers. And by consumers I don't mean

28:25

like individuals, I mean the consumers

28:26

of the APIs. Because

28:29

if you're a, you know, a business that

28:31

is doing in our like software

28:33

engineering

28:34

your job is at a very high level to

28:36

solve problems.

28:38

And if we can allow you to solve those

28:39

problems faster and with cheaper models

28:42

that are just as performant, that means

28:44

you can solve more problems. Like that

28:45

is a good thing. And it is a very good

28:47

world where there is not like a monopoly

28:50

on intelligence, but instead kind of a a

28:52

garden of intelligence that you can pick

28:54

and choose um you know, when you'd like.

28:57

It's something that we joke about is

28:58

like

28:59

you know, on this intelligence

29:01

allocation thing. Um

29:03

if you're if you're trying to get a a

29:04

tutor for your daughter in algebra,

29:07

you can probably find someone cheaper

29:08

than Albert Einstein to be that tutor.

29:11

Now, it might be that she eventually

29:12

goes and becomes like a leading, you

29:14

know, physicist or something, in which

29:15

case yeah, maybe let's let's get Albert

29:17

Einstein in there. But most likely you

29:19

can get, you know, a high school student

29:21

or something like that. Um and it's

29:23

probably much more cost-effective for

29:24

you as well to do so. So.

29:26

>> Since you guys do the model routing,

29:28

like if you look at the you know, if

29:29

there's a pie chart that shows the

29:30

complexion of models being used by your

29:32

customer base today,

29:34

what did it look like a few months ago?

29:36

What does it look like today? What do

29:37

you think it'll look like in a year?

29:38

>> Yeah. I will caveat this with saying

29:40

that right now

29:42

enterprises haven't gone too opinionated

29:44

yet into the routing procedures.

29:46

>> Okay.

29:46

>> This is something that will happen over

29:47

the next 6 to 12 months, but right now

29:49

they're just going from no router to

29:50

router. That's kind of the first change.

29:52

Then it's going to be like the exact

29:54

nature of the of the routing. At the

29:56

beginning of the year,

29:57

there's less than 1% of tokens

30:00

went to open models.

30:02

In the first quarter, it became a

30:04

single-digit percent.

30:06

It is now crossed into being a

30:07

double-digit percent of tokens. Um

30:10

now, percent of tokens is not always the

30:11

same as percent of cost cuz the open

30:13

tokens are cheaper. Um but

30:15

uh it's pretty crazy to see the the

30:16

growth there.

30:17

>> What's your forecast?

30:19

>> [sighs]

30:20

>> My sense is that

30:21

we will asymptote towards vast majority

30:24

being open just because it provides you

30:26

more optionality and it's cheaper. Um

30:29

but

30:30

that doesn't mean they're going to be

30:32

like that's of token share, not

30:33

necessarily of leverage share. Cuz maybe

30:36

there are 1% of tokens that are

30:37

incredibly incredibly valuable um

30:40

and are like very key decision-making,

30:42

and then the rest are more like

30:43

implementation tokens or kind of

30:46

uh lower stakes, if you will. I don't

30:48

think there's going to be a world in

30:49

which like

30:50

it's ever going to be 100%.

30:52

>> Yeah.

30:52

>> I think the frontier of intelligence

30:54

will inherently always be valuable for

30:56

every business, could just cuz the

30:57

stakes are going to get higher in the

30:59

kind of intricacy with which you think

31:00

is going to be more important, but we'll

31:02

be better at offloading certain tasks.

31:04

And this is like you can loosely think

31:06

of this

31:07

uh already with the way orgs are

31:08

structured, where you know, in general,

31:11

engineering leaders are more tenured

31:13

engineers who in theory have like more

31:15

wisdom, and each kind of minute of their

31:18

brain power is higher leverage, in

31:19

theory. Um, and even, you know, you can

31:22

also imagine like consider a human

31:25

engineer and try mapping over the course

31:27

of their day like how much brain power

31:30

they're using. And like, you know, it's

31:32

probably going to be really low for a

31:34

lot of it, but then there're going to be

31:35

some moments where they're like going

31:37

pretty hot, like they're deeply

31:38

concentrating and thinking about some

31:40

you know, systems design problem or

31:41

whatever. All of those low leverage

31:43

moments, we want to automate away.

31:46

And like we want to like those like very

31:47

high leverage moments sometimes it's

31:49

like, you know, we're referring to them

31:50

as like the eureka moments, or the

31:52

moments where they're like doing

31:53

something that's very high leverage.

31:55

What if those aren't just moments, but

31:56

what if those are like

31:57

hours at a time?

31:59

Because you don't have to deal with all

32:00

the other stuff. And I think that's kind

32:01

of the the way to think about

32:03

intelligence allocation is if you're an

32:05

engineer and you're writing docs, that

32:07

is such a low leverage use of your time.

32:08

Like you've become an expert in your

32:10

craft.

32:11

And you used to spend hours writing

32:12

docs. Like I remember it was actually

32:14

valuable. Like I remember Stripe had so

32:17

much alpha for just having incredible

32:20

docs.

32:21

But imagine all the other stuff those

32:22

incredible engineers could do if it

32:23

wasn't writing documentation. Like we

32:25

should live in a world where everyone

32:27

can have docs as good as Stripe, and

32:29

that is like strictly beneficial for

32:31

everyone. And then the question is,

32:33

okay, what do those really smart

32:34

engineers do do their time once they

32:35

don't have to do that?

32:37

>> Maybe this is a good time to talk about

32:38

business model, given that, you know,

32:40

especially with the rise of open weight

32:41

models, the cost differential.

32:44

Um I imagine that means very different

32:45

things for your for your cost structure,

32:48

but very similar value delivered to

32:49

customers. How do you think about uh

32:52

business model and pricing?

32:53

>> Yeah, this is more what our customers

32:55

want and need as opposed to what we want

32:57

and need. So, for example, I think right

32:58

now usage-based is clearly the way to

33:00

go. We want to be aligned with like what

33:02

they are doing and what we are doing. I

33:03

think seat-based doesn't make sense at

33:05

least for what we are doing. My sense is

33:07

that eventually we will change to

33:09

outcome-based.

33:11

Now,

33:12

I don't think the enterprise is ready

33:13

for that and we've learned our lesson

33:14

from those first 2 years. We are not

33:16

going to impose things, right?

33:17

[laughter]

33:18

Um but my suspicion is that

33:21

you know, in the 2030s, things will

33:23

probably look more like outcome-based.

33:25

>> Yeah. What does outcome-based mean for

33:27

your market? What would be the

33:28

definition of an outcome?

33:29

>> So, maybe here's a way to put it. So,

33:31

right now we charge we we are

33:32

usage-based. Like the more tokens you

33:33

use, you know, the more you pay, the

33:35

more we get. Um

33:37

now,

33:39

since we are model-independent, we kind

33:42

of with our router, we are kind of

33:44

pointing a token cannon at either

33:47

OpenAI, Anthropic, AWS, GCP, you know,

33:49

any one of these people.

33:51

Uh to a certain degree, this is like a

33:54

really

33:56

dumbed-down version of a marketplace.

33:59

Where right now, there is a a buyer the

34:01

buy side is an engineer who wants a task

34:03

done. And then, you have the model

34:05

providers who are saying like

34:07

either in benchmarks right now, they're

34:09

like, we perform at this cost and this

34:11

performance. And then we determine who

34:13

we go to for that given task. Yeah.

34:15

There's a world in which,

34:17

you know, if it's so important to get

34:19

these tokens, they might kind of

34:22

unaf- like bid in a certain way of

34:24

saying like, look, here is our cost for

34:26

this task. We will get this task done at

34:29

this cost no matter what, but they're

34:30

pricing it such that, you know, they

34:32

hope that they can make a margin there.

34:33

If they price it wrong, they're at a

34:35

negative margin. If they price it right

34:36

and win the bid, then they get the

34:37

positive margin. And the way you

34:40

determine if the task was successful

34:42

is by some validation loops.

34:44

Cuz no one is using these tools anymore

34:45

where it's just like, "Write me code.

34:47

Great. Thank you." It's generally,

34:48

"Write me code and here's how I know it

34:50

was done well."

34:51

And similarly,

34:53

if you are like a model lab and you are

34:55

given, "Here's a task. Here's the

34:56

validation criteria." You'll be able to

34:58

say roughly how much you think you would

35:00

be willing to pay to get those tokens. I

35:02

say and you know, you want to have some

35:04

some margin on that. And then in that

35:06

world, that's basically that's a way to

35:08

that you kind of dynamically shift from

35:10

usage-based to outcome-based. Um I think

35:12

that there's so many questions with this

35:14

and this is very much forward-looking.

35:16

Um but I think there's a lot of

35:17

questions about how do you subdivide

35:19

tasks? You know, divvying that up I

35:21

think is something that's not obvious.

35:22

>> Yeah.

35:22

>> But as these tools get better, doing

35:24

things like that actually become way

35:25

easier.

35:25

>> Yeah, that's fascinating.

35:26

>> Yeah.

35:27

>> Yeah, if you can scope a task and then

35:29

create a competitive market for a place,

35:30

that'd be a fascinating version of the

35:31

future.

35:32

>> Yes. And as a user, it then creates an

35:33

incentive to be very thorough in your

35:35

validation criteria.

35:36

>> Yeah.

35:36

>> Cuz like, you know, there are stories of

35:37

like, you know, you ask an agent to like

35:39

fix my code and it deletes your code.

35:41

>> Yeah.

35:42

>> It's like, you know, the solution is

35:43

just get rid of it all cuz it had

35:44

>> that Silicon Valley episode. It's crazy

35:46

how precious it was.

35:47

>> Yeah. But like, so you need to make sure

35:48

your tests are very thorough cuz

35:50

technically it could hit all of your

35:51

validation criteria.

35:52

>> will go rogue?

35:53

>> Yeah, exactly. Yeah, yeah, [laughter]

35:54

yeah. Exactly. Yeah, yeah, yeah. Yeah.

35:56

That happens.

35:56

>> Um maybe zooming out a little bit, you

35:59

named the company Factory. Actually, you

36:00

named it Droid before before Factory.

36:03

But you named it Factory before this

36:05

concept took off. And now it feels like

36:07

everybody wants to build a software

36:08

factory. Where do you think we are today

36:11

in in terms of the building of software

36:12

factories and how close are we to the

36:14

ultimate vision of a software factory?

36:16

>> Yeah, everyone has a software factory

36:18

whether they know it or not. It's just a

36:19

very inefficient one. So it's kind of

36:21

like it feels like, you know,

36:23

pre-industrialization

36:24

where like, you know, people were

36:25

manually like, you know, sewing things

36:27

together or like woodworking or whatever

36:29

it might be. And

36:31

these things are very inefficient. Like

36:33

right now, if you go to an organization

36:34

that has more than 10,000 people, and

36:37

you were to ask about the process by

36:39

which they decide and release a feature,

36:43

there is like hundreds or maybe

36:44

thousands of people in that process, and

36:47

most likely they couldn't even draw it

36:49

for you. Like, there's very low

36:50

likelihood that they would know what

36:52

what that process looks like. Um that is

36:54

not because they think that is the right

36:56

way of doing things. That is just kind

36:58

of the nature of building large software

37:02

as it is kind of today.

37:04

But with these systems,

37:06

so much tribal knowledge can be

37:07

codified. So much of this stuff that

37:09

typically would require, oh, we need to

37:10

ask this guru who's been here for 30

37:12

years who has the wisdom. Oh, we then

37:14

need this approval and that approval.

37:15

Oh, and I forgot there was some doc that

37:17

said we always have to do this

37:18

checklist. Um and it relies so much on

37:20

kind of human behavior and like

37:22

redundancy,

37:24

so much of that can be automated and

37:26

refocused on like what actually moves

37:28

the needle for our business. And I think

37:29

this move towards software factories is

37:31

a move towards how do we figure out what

37:34

are the actual inputs that determine

37:36

what features we need to build? And that

37:38

might be inputs from the customers,

37:40

inputs from the market, inputs from

37:42

like, you know, product leaders at the

37:44

company.

37:45

And let's be very clear. These are the

37:47

signals, the inputs that we are taking

37:48

in here. Okay, great. We have those

37:49

signals. Then what is the process by

37:51

which we build this? Um and really like

37:53

mapping out the like assembly lines of

37:56

how you are building software is really

37:58

important because then

37:59

you get to close the loop and say,

38:01

did this actually deliver outcome for

38:03

our business?

38:04

Talking before about the tokenomics,

38:06

if you're that CIO and you're faced with

38:08

that question of where do you put every

38:10

incremental token, really the question 2

38:12

years from now is going to become where

38:13

do you put every incremental dollar?

38:16

And so, you're going to have to be be

38:17

asked,

38:19

do you put that incremental dollar

38:20

towards headcount or towards tokens? And

38:22

if tokens, to where in the org.

38:24

And these are things that you can only

38:26

really know when you have these kind of

38:28

feedback loops that give you examples of

38:30

like, "Hey, by the way, we made those

38:31

decisions based on this data, and it did

38:33

not matter at all. We added these new

38:35

features and no one cared. It didn't

38:37

create more retention, it didn't create

38:38

more usage, or whatever metrics that

38:40

business is looking to optimize." And

38:42

the only way to do this is like you need

38:44

kind of more rigor and more process.

38:46

It almost feels like

38:48

like 10 years from now, we're going to

38:49

look back at this previous era of

38:51

software, and it's going to feel like

38:53

businesses in like

38:55

ancient times where they didn't do

38:57

accounting.

38:58

It is like

38:58

>> to be like it's going to be like

38:59

marketing in the day of Mad Men, right?

39:01

Where it's like all creative and you

39:02

have no idea what's actually working.

39:04

>> It makes no like it's like, "Oh, yeah,

39:05

let's ship that feature. Oh, I think it

39:07

went well. Like, yeah, we had I got some

39:09

metrics on that." It's like, "No." If

39:11

you guys read the the blog post that

39:14

Jack Dorsey put out about how every

39:16

company is like an AGI,

39:17

>> Yeah.

39:18

>> there's also this degree to which if

39:20

your company is an AGI, you want to

39:22

optimize the weights.

39:23

>> Yeah.

39:23

>> You want to figure out what nodes are

39:24

doing the what things, which are

39:26

load-bearing, which are not, which need

39:28

more tokens, where do you need more

39:30

nodes.

39:31

And in order to do it like you don't

39:33

train a model by vibes. I mean, okay,

39:35

actually you kind of do, but

39:37

>> [laughter]

39:37

>> you don't I guess more importantly you

39:38

don't do backprop in a model by vibes.

39:40

Like you are running those actual like

39:42

calculations, and you are seeing when we

39:44

change this node, what happens. Now, you

39:46

might be making bets on how to change

39:47

the model by vibes, but you like you're

39:50

it's pretty like mathematical in what

39:52

you were doing. Meanwhile, at companies,

39:55

you know, people are determining token

39:57

budgets just by shooting from the hip.

39:59

People are laying people off by shooting

40:01

from the hip and just being like, "Oh,

40:02

yeah, like 20,000 There is no way there

40:05

is science to laying off 20,000 people."

40:07

That is just like, "Here is a chunk, and

40:09

let's just see what happens." Instead, I

40:11

think in these organizations, the way

40:12

they can do things is much more

40:14

mathematical of like

40:16

this part of the business matters a lot

40:18

and does better if we give it more

40:19

tokens. It doesn't actually matter if we

40:21

give it more humans. So, let's give them

40:23

more tokens. There might be other parts

40:24

of the business where actually giving

40:26

them more tokens doesn't matter, but

40:28

more people matter because if we build

40:30

more relationships with our customers

40:31

and deeper relationships with our

40:32

customers, that matters. But, these are

40:34

things that we're going to need like

40:36

quantitative insight on

40:38

and you need a software factory to do

40:39

that. Otherwise, you're just like

40:40

shooting from the hip and just guessing,

40:42

which

40:43

won't work as well.

40:44

>> limit

40:46

how much do you think people will spend

40:47

on tokens versus on engineering head

40:50

count?

40:50

>> It'll depend on the business.

40:51

>> Mhm.

40:52

>> I think every business will have a

40:53

balance and it just depends on like like

40:56

they're just going to be like an easy

40:58

example is generally sales people, they

41:00

probably don't need that many tokens if

41:02

they're good sales people. Cuz generally

41:04

the where they provide the most alpha is

41:05

like when they're in the seat

41:07

face-to-face with their customers

41:09

talking about the customers' problems,

41:10

understanding, you know, how they build

41:12

software in our case, and how we can

41:14

make that you know, more efficient, more

41:15

productive. They can use tokens a little

41:17

bit of like, oh, whatever, generate them

41:19

some, you know, AI debrief, take some

41:21

notes, like help them with the

41:22

follow-up, but like it's so minimal the

41:25

number of tokens it basically doesn't

41:26

matter. Like if you add more tokens to

41:28

the sales team, it probably won't change

41:30

their output. If you add more humans to

41:32

the sales team, it probably will.

41:34

Meanwhile, engineering teams are pretty

41:35

different where engineering teams

41:37

generally it seems like

41:39

the you want people to own an outcome

41:41

end-to-end, but then if you give them

41:43

more tokens, they can produce a lot

41:45

more. And so, it seems like there and

41:48

then there's a lot of kind of places in

41:49

between of like operations, finance,

41:51

marketing. These are places where are

41:52

neither here nor there where I think

41:53

they're they're somewhere in between and

41:55

it kind of depends on your business.

41:56

But, I think every business is going to

41:58

have to ask, like what is our core

42:00

competency?

42:01

Something that we see a lot in the

42:02

market or we used to see and now they

42:03

finally kind of hit reality. What we

42:05

used to see is

42:06

oh, like we're going to build our own

42:08

like software development agents.

42:10

And we're like, okay, like you're a like

42:12

a consumer

42:14

uh like logistics company.

42:16

Like are you sure you want to do that?

42:17

They're like, yeah, yeah, yeah, we're

42:17

going to This is a core We have to do

42:19

this. And so I'm like, okay. And then 6

42:20

months later it's like, wait, actually,

42:22

this is not a core competency for our

42:23

business. We don't want to hire, you

42:25

know, AI engineers to be doing this. Our

42:27

core competency is, you know, consumer

42:29

logistics. That's what we want to focus

42:31

on. And I think this is an opportunity

42:33

for every business to double down on

42:35

their core competency and what matters

42:36

for them and then procure externally

42:39

whatever it is that doesn't matter for

42:40

them. Like a trivial example of this is

42:42

like,

42:43

I don't know, in the days of the early

42:44

internet, you probably had to be a

42:46

programmer to build a website. And like

42:49

websites generally help if you're a

42:51

pizza shop cuz you want to have, you

42:52

know, people come to your pizza shop,

42:54

they want to be able to or like

42:55

whatever. At that time, would you say it

42:57

was a core competency of like a pizza

42:59

shop to have engineers?

43:01

Like certainly not. Like that is kind of

43:03

a byproduct of like a brief moment in

43:05

time, but then there were companies out

43:06

there that help you build a website, you

43:08

don't need to be technical, and then

43:10

this is why we live in a world where

43:11

like most pizza shops don't have an

43:13

engineering department, which I think is

43:14

probably a good thing. Um and I think

43:16

similarly, a lot of businesses have

43:17

dealt with the reality of if you want to

43:20

do XYZ other thing, you have to

43:23

bring in people of this type of role,

43:25

but I think that's been like something

43:26

you had to do not because it's a core

43:28

competency of the business.

43:30

And allowing businesses to focus and

43:31

double down on the things that they are

43:32

best at, I think it's going to be good

43:34

for the consumers of their business. And

43:35

so I think we're just going to see like

43:37

a lot like ruthless refocusing on what

43:39

actually matters,

43:41

um

43:42

which is going to be cool to see. Well,

43:43

on that so, you know, every company kind

43:44

of has to go through this process of

43:46

reinvention. You know, 10 or 20 years

43:48

ago, people talked about digital

43:49

transformation.

43:50

And I don't know if anybody's given it a

43:52

buzzword now, but AI transformation,

43:54

something of that sort. Um couple years

43:56

ago, you ran into a bunch of

43:58

organizations that just weren't ready to

44:00

deal with autonomous agents.

44:02

Things you've seen your customer start

44:04

to change. And so the question is,

44:07

when you look at your customers as they

44:08

kind of go up this maturity curve and

44:10

sort of reinvent themselves for the

44:11

future,

44:12

um any good like tricks or techniques

44:15

that you've seen them use to

44:17

repot themselves a bit?

44:19

>> Yeah, I mean, I think um

44:22

surprisingly, like the companies that

44:24

have been doing like company-wide

44:25

hackathons

44:27

really end up doing well. It seems like

44:29

relatively trivial, but like just

44:30

setting aside a day

44:32

for everyone in the workforce is just

44:33

like

44:34

build with AI.

44:36

It really sets the tone and sets the

44:38

pace.

44:39

>> Certainly has given me a look as she's

44:41

>> I I tried to force him to build

44:42

something [laughter] with Coding Agent.

44:44

Didn't go so well.

44:45

>> We'll work on it. We'll do after this

44:46

one.

44:46

>> I you know, I we gave it a great effort.

44:49

>> But that's it. Like it literally just

44:50

setting aside the time to like do it.

44:52

And like even if it fails miserably,

44:54

like it's fine. And also like the orgs

44:56

that are okay with failing. Like it feel

44:58

like it feels like there are some who

44:59

are like, "We need to do it exactly

45:01

right. We need to make the right

45:02

decision from day one. No

45:03

>> Yeah.

45:04

>> Like you're going to make mistakes.

45:05

Everyone is going to. And the orgs who

45:07

are kind of

45:08

leaning into it and embracing it to a

45:10

certain degree, I think are succeeding.

45:12

Like one of our largest customers is EY.

45:14

EY is not necessarily known to be like

45:16

at the absolute frontier of AI, but I

45:18

think for them, they were just like,

45:19

"Look, this matters. We were kind of

45:22

There have been other trends and

45:23

transformations that we were late to.

45:24

We're not going to be late to this. Like

45:26

we're just going to go in. We might mess

45:27

up, but like obviously respecting like

45:30

secure The things that you're not

45:31

allowed to mess up. You can put those

45:32

aside. But like

45:34

let's go and get our engineers to mess

45:35

around and build this stuff and see

45:37

where it breaks and understand what they

45:38

like and what they don't like. Um

45:41

I think that really matters a lot in the

45:44

ones that we're seeing succeed. And also

45:45

the ones who are like

45:48

pretty bold in reinventing the processes

45:50

that they've put in place. And just

45:52

saying like, "Hey, it's a There's no

45:53

sacred cows. Like

45:55

let's let's put this aside, try

45:56

something out. If it doesn't work, we'll

45:57

put that sacred cow right back. Um and I

46:00

think that's that's been kind of a

46:01

determining factor there. Um and when it

46:03

comes from within, if it comes from the

46:05

board, probably not going to go well.

46:06

>> Yeah.

46:07

>> If it comes from within like the tech

46:08

team or the the ICs or the leadership,

46:11

that's when we see it go better.

46:12

>> Hm. Do you have any predictions for

46:15

the most important changes that are

46:16

going to happen in your space over the

46:17

next, call it, 12 months?

46:19

>> A lot of AI consumption's going up like

46:21

crazy. And everyone's super, super

46:22

excited because our revenue's going

46:24

wild. Like a lot of this is synchronous

46:26

usage. In other words, like if everyone

46:29

woke up sick tomorrow, like

46:31

a lot of Claude code usage would be

46:33

zero. Cuz it's all just, "Hey, Claude

46:35

code." Or, "Hey, Codex." Or, "Hey,

46:36

Droid." Right? I think in 12 to 24

46:39

months, like 90% of tokens will be

46:41

asynchronous tokens. So, these are going

46:43

to be, you know, droids on their own

46:46

autonomously being like, "Hey, here's

46:47

some signal that I found from a

46:49

customer. Let's go fix it." Or, "Let's

46:50

go create a first-pass solution to

46:52

this." And I think that is going to be

46:54

where

46:55

the real like agent-native stuff begins.

46:58

Cuz right now we're still kind of in

46:59

like co-pilot mode. Like if you're going

47:01

to an agent say, "Hey, go do this for

47:02

me." It is more agentic because it's not

47:04

going to come back and ask you a ton.

47:06

But it's still like you are kicking it

47:07

off. Like yeah, if you guys have ever

47:09

been to Tesla's factories, which is one

47:10

of the sources of inspiration for the

47:12

name, is like

47:13

it's just robotic arms everywhere going

47:15

and doing stuff. Like it's not like

47:16

there are people there like going and,

47:17

you know, attaching the widget to the

47:19

thing. Um and this idea of like a dark

47:21

factory where like the lights are off

47:23

and things are just happening, that is

47:25

where software development's going.

47:26

That's kind of where the the name came

47:28

from is like, you know, Elon was always

47:30

talking about the factory is the machine

47:31

that builds the machine.

47:32

>> Yeah.

47:32

>> And that's been something that we took

47:34

to heart. Um and I guess also that

47:36

combined with his whole thing about

47:38

how you're destined to become the

47:40

opposite of your name.

47:41

>> Hm.

47:41

>> Um

47:42

and in our case, you know, factory

47:43

becomes artisanal. Which is kind of a

47:45

good uh

47:46

a good flip there. So.

47:48

>> What's your most optimistic version of

47:49

the future both for Factory and for the

47:51

world at large?

47:53

>> So, I think short-term

47:54

there's going to be a lot of turbulence

47:57

because I think

47:59

a lot of companies have misallocated

48:01

resources pretty poorly. There's been a

48:03

lot of bloat. Um and I think the

48:05

correction that's going to happen there

48:06

is going to be really painful for a lot

48:07

of people. And I think that's something

48:09

that I think every AI CEO should really

48:11

bear much more responsibility than they

48:13

currently are for. Um and also figuring

48:16

out ways to like

48:18

address and kind of ameliorate in some

48:21

way because this is something that's

48:22

going to be very painful for a lot of

48:23

people.

48:24

Now,

48:25

I also I have optimism that we can

48:28

actually address that faster than we

48:29

think. We just need to start now in

48:31

terms of addressing that. Now, the

48:32

longer term and why I think this is a

48:34

good thing is

48:36

and why I don't believe at all like, you

48:37

know, the BS that people are saying of

48:39

oh, engineers are going away. Generally,

48:41

there is a huge number of problems in

48:42

the world.

48:43

A large subset of those problems can be

48:45

solved with software.

48:47

A small subset of those problems are

48:48

currently being solved with software.

48:50

And so, in the short term, this means

48:53

that okay, first there's a given problem

48:55

that was overallocated engineering

48:57

resources. So, okay, we need to

48:59

reallocate those. Reallocate those is a

49:01

very

49:02

kind of cold way of saying some people

49:04

are going to lose their jobs. But I

49:05

think the the thing that's going to

49:07

happen in the longer term is

49:09

we need engineers. Engineers are some of

49:11

the best systems thinkers and the best

49:12

problem solvers. And there are so many

49:14

problems that can be solved with

49:15

software that are not being solved with

49:17

software. And so, that means that we are

49:18

going to take those engineers and have

49:20

them go and solve problems that

49:22

previously were not being solved. That

49:23

is such a net good for the world.

49:25

Because again, there are so many there's

49:27

problems that we are not solving. And

49:28

also, there's so many problems that we

49:30

are maybe solving but with really shitty

49:32

software. And like, this is going to

49:33

enable people to solve it with

49:34

incredible software. And, you know, the

49:37

vision for Factory is that we are kind

49:38

of the the factory that allows them to

49:41

go and build this incredible software to

49:43

solve these different problems. And

49:44

these problems range from like things

49:46

that are trivial to you know, like

49:49

government software typically is not

49:50

very good, whether it's like DMV or like

49:53

IRS web like all that stuff is generally

49:55

a pretty poor experience. Um we don't

49:57

need to live like that. Like we can all

49:58

live in we can live in a world where all

50:00

software is really fantastic. Um but

50:03

also things like you know,

50:04

pharmaceutical research. Like so much

50:06

that goes into solving diseases is not

50:09

just like a biology problem. A lot of it

50:12

requires the best software engineers in

50:13

the world. And previously those problems

50:17

haven't allocated the right dollars to

50:19

attract the best engineers.

50:21

But now because of what's happening, I

50:22

think we will be much more closely

50:24

allocated to like these are the biggest

50:26

problems. Let's get the best minds and

50:27

the best problem solvers to solve that.

50:29

Um I think it's kind of our job as an

50:32

industry to do that relocation

50:34

reallocation as quickly as possible. So

50:37

it's not 10 years, but maybe like 6

50:39

months or a year.

50:40

>> Wonderful. Maton, I think the clarity

50:41

and consistency of your vision over time

50:44

has just always been very inspiring and

50:46

then just seeing how much you've grown

50:47

as a leader and how much Factory has

50:49

grown as a company. Even since the last

50:51

time we did the Stranded Deep episode,

50:52

it's truly all inspiring. So thank you

50:55

for for joining us again to share what

50:56

you're up to.

50:57

>> I appreciate it a lot. Thank you.

50:59

>> Thank you.

51:09

>> [music]

51:25

[music]

Interactive Summary

The video features Matan, CEO and co-founder of Factory, discussing his company's journey in building autonomous software development agents. Matan emphasizes the core operating principle of creating 'obsessed customers' rather than just focusing on internal metrics. He shares insights into the challenges of being early to the market, the importance of model independence for enterprises to avoid vendor lock-in, and the shift from 'token maxing' to intelligent model routing. Matan also discusses the future of software factories, highlighting how these systems enable engineers to focus on high-leverage problem solving, and shares his optimistic outlook on using AI to solve previously unaddressed global challenges.

Suggested questions

5 ready-made prompts