HomeVideos

AI Tools in Regulated Banking: Monzo's Control Model | Suhail Patel

Now Playing

AI Tools in Regulated Banking: Monzo's Control Model | Suhail Patel

Transcript

1462 segments

0:09

Hey there, I'm Matt. I run product

0:10

design engineering at Ona, and I'm

0:12

joined by Suhaib.

0:14

Let me try that again. That didn't go

0:14

very well.

0:16

Hey there, I'm Matt. I run product

0:17

design engineering here at Ona, and I'm

0:19

joined by my good friend Suhaib, who has

0:21

worked at Monzo for many, many years

0:23

now. Uh, if you haven't heard of Monzo,

0:25

I don't know where you've been living,

0:26

probably under a rock, but Monzo is one

0:27

of the UK's like defining banks. Um, as

0:30

of 2025, we reported at 12 million

0:32

customers. Um, and Monzo's

0:34

And Suhaib's been there since

0:37

Is it 8, 9 years now?

0:38

Yeah, 8 years now. 8 years now. Long

0:40

amount of time.

0:42

Yeah, and you currently are a principal

0:44

engineer, I believe, who sits in the

0:46

platform group. Can you tell us a little

0:47

bit more about your role in the platform

0:48

group?

0:50

Yeah, absolutely. Um, so yes, hi, my

0:51

name is Suhaib. I am one of the

0:53

principal engineers at Monzo, um, and

0:55

I've been at Monzo for about 8 years

0:57

now. Um, and I help lead the platform

1:00

group, uh, as an individual contributor,

1:02

and the platform group is responsible

1:03

for building Monzo's foundational base.

1:06

So, we look after all of the

1:07

infrastructure, you know, the the

1:08

standard things that you would expect

1:10

with AWS and Google Cloud, and, you

1:12

know, all the sort of like, cloud-native

1:14

relationships and things like that, and

1:15

all the cloud-native tools that sit sit

1:17

on top of these these, uh, providers.

1:19

For example, your Kubernetes and your

1:21

Kafka and your your your Cassandras and

1:24

all your myriad of technologies that,

1:26

you know, a modern engineering

1:28

organization would adopt. But, our real

1:31

secret sauce is, uh, within the platform

1:33

group, we also own the relationship with

1:35

engineers, the developer experience that

1:37

engineers interface with on a day-to-day

1:39

basis. So, our goal is to provide a very

1:41

opinionated architecture, uh, and very

1:44

opinionated code base, um, and libraries

1:47

and tools that engineers uh,

1:50

want to adopt. Uh, you know, we we don't

1:52

sort of force them to adopt it, because

1:54

that is the way to get stuff, uh,

1:56

without friction onto the platform,

1:58

Uh and comes with a bunch of batteries

1:59

included, but you know, there's only so

2:02

much forcing you can do. Uh like, you

2:03

know, the tools have to be really,

2:04

really good for engineers to want to use

2:06

them.

2:07

Uh and use the standardized stack. Uh

2:10

and what we do is we help build and

2:12

maintain and provide really, really good

2:14

tooling and observability around that

2:15

stack, so that engineers can focus on

2:17

building the bank and not have to worry

2:19

about infrastructure.

2:21

There's so many things you said there

2:22

that I want to jump off. So, um we'll

2:24

start with um how big is the platform

2:26

group? It sounds like you've got a

2:27

really wide mandate, actually.

2:30

Yeah, absolutely. We've got um

2:32

I can't even remember the amount of

2:33

squads we have because we're adding uh

2:36

uh

2:36

bits and pieces all over the place. We

2:39

span a very wide gamut between um you

2:41

know, core platform and infrastructures,

2:43

the stuff that, you know, will interface

2:45

with the app on a day-to-day basis, to

2:47

things like um you know, our data

2:49

infrastructure as well, like in our in

2:50

our data warehousing estate, uh making

2:52

sure that we have really good modeling

2:53

tooling

2:54

uh for for data modeling and analytics

2:56

and and and business intelligence. Um

2:59

we've also uh very recently um

3:02

uh done a significant amount of

3:03

investment in machine learning. Um and

3:05

this is traditional machine learning.

3:07

Um you know, your neural nets and and

3:09

your models and

3:10

uh

3:11

and the standard stuff, you know, which

3:12

is absolutely big in in our domain and

3:15

space. Uh and a lot of really, really

3:17

good techniques that we can adopt um and

3:20

and run with. Uh so, we've a significant

3:23

amount of investment in machine

3:24

learning, and a lot of that has extended

3:25

into into a lot of the really cool stuff

3:27

with the AI and LLMs as well. Um and on

3:31

the other side, uh you know, we have

3:32

teams that deal with developer

3:33

experience. Um and also, you know, even

3:36

things that are very, very important,

3:37

things like FinOps. Um you know, make

3:39

sure that we're not spending

3:40

egregiously. Um you know, our customer

3:43

uh unit cost is uh staying within

3:46

control. Um so, yeah, a really, really

3:48

wide gamut of uh teams are are included

3:51

within within the platform group.

3:53

And for the platform group specifically,

3:55

do you find that your hiring plans are

3:58

trending up since you've seen AI

4:00

adoption come in or is it kind of going

4:01

flat? You're hiring less? Like how do

4:03

you think about hiring in a world where

4:05

AI tooling it can can can help you be

4:07

way more productive?

4:09

Yeah, that is a great question.

4:11

Obviously, there's a lot of stuff out

4:12

there in the industry which is like, you

4:13

know, all these tools are going to

4:15

replace engineers.

4:17

And you know, these tools are just

4:18

really really good. Like, you know, even

4:20

the last couple months you know, with

4:22

with with all the tools out there, all

4:24

of the foundational models and even open

4:26

models as well.

4:27

You know, the tools have been changed

4:29

the way that we work on a day-to-day

4:31

basis.

4:32

Even me like, you know, early on I was a

4:33

bit skeptical about like, you know, the

4:35

amount of revolution this would have on

4:37

on the day-to-day. But now like, you

4:39

know, I don't even open my editors like

4:41

my first port of call when I start the

4:43

day. The first thing I do is open my

4:44

terminal open open cloud code.

4:47

So like it is like fundamentally rewired

4:50

how I think about my brain.

4:52

And you know, this is a journey that

4:54

many engineers have have gone on. For

4:56

specifically at Monzo, we've just become

4:57

more ambitious.

4:59

So like, you know, ultimately our hiring

5:00

plans are remaining like, you know,

5:03

business as usual, but we've just become

5:05

more ambitious in what we want to be

5:07

able to achieve.

5:09

You know, we are definitely hiring more

5:10

like, you know, we've got a bunch of

5:11

open roles

5:13

open and continue to have them open just

5:15

because our ambition has gotten bigger

5:17

and and bolder and we want to be able to

5:21

do more as a result of that. You know,

5:24

fundamentally you know, these tools are

5:25

not going to replace all of the product

5:27

things we want to be able to build.

5:30

You know, if these tools could click a

5:31

finger and we could have all of our

5:33

product stuff built, you know, we would

5:35

have more ambition in the product that

5:36

we want to build

5:38

in general. I don't think this would be

5:39

a replacement for the for the

5:42

for the ambition that we have and for

5:44

the people that we have and will

5:45

continue to hire.

5:47

I love answer. It's exactly what we've

5:49

seen too is, you know, we we actually do

5:50

want to do more of it, not less of it,

5:52

right? Is um great engineers can now do

5:54

great things even faster. So, you know,

5:55

who who doesn't want more of that?

5:57

Um I'm yet to see an empty backlog. Um

6:00

don't know if you've seen one yet, but I

6:00

haven't seen one. No, absolutely not.

6:02

Like, you know, the backlogs um continue

6:04

to get ever bigger. And you know, just

6:06

specifically in the world of platform,

6:08

um like, you know, uh the amount of

6:10

people that we reach is, like, you know,

6:12

gotten infinitely larger. You know, our

6:13

remit, you know, uh was basically to

6:16

make engineers more productive. And now,

6:18

like, our our remit has expanded to make

6:21

the whole organization more productive.

6:23

Um you know, when the cost of writing

6:24

code is basically zero, right? Um

6:27

especially for prototyping and you know,

6:28

one-offs and things like that. You know,

6:30

you still want to have a reliable,

6:31

secure platform to go ship these things

6:33

on. Uh you know, to serve them, like,

6:35

you know, to to, you know, host them

6:37

internally to to users and to have the

6:39

right controls and security and

6:41

mechanism. And these things need to come

6:43

batteries included by default, have

6:44

observability. Uh you know, if you're

6:46

running a model and, like, you know, you

6:47

want to backtest on some data and you

6:49

want to see how that how that data

6:51

performed. Um you know, even if you've

6:53

come up with an idea and sort of live

6:55

coded a prototype before, uh like, you

6:57

know, for for anything, you still need

6:58

to have all of these mechanisms

7:00

available. So, our ambition and the

7:01

amount of people that we are serving has

7:04

gotten bigger.

7:06

That makes tons of sense to me. It's

7:07

exactly what we've seen too. How do you

7:09

think about rolling AI AI tools out to

7:12

um engineers and business users? Is

7:14

there a different uh roles or different

7:16

ways they have to operate? And what does

7:17

it take for a business user to to land a

7:19

PR in production, for example?

7:21

Yeah.

7:23

I don't think we've been perfect at

7:25

this. So, like, you know, I wouldn't

7:26

take our answer as like the the the the

7:28

pinnacle of of perfection. I think like

7:31

most organizations we're still figuring

7:32

it out, right? Um

7:34

like, you know, we've had to be very,

7:35

very deliberate about rolling these

7:37

things out. Um there's one side which is

7:39

like we want to give access to all of

7:40

these tools. We've onboarded all of

7:41

these tools, like you know, we've gone

7:43

through all of the all of the processes

7:45

and like you know, gotten access to all

7:46

of these tools and rolled them out to

7:47

engineers as quickly as possible, so we

7:49

can get it in the hands of people and

7:51

learn ourselves what these tools are are

7:53

capable of doing, right? And for the

7:56

entire industry, this is a learning

7:58

experience, right? Now, what we've been

8:00

able to rely on as we brought these

8:01

tools out to engineers is that engineers

8:03

have, for example, done their secure

8:05

development training, like you know,

8:06

have got like all all of the various

8:08

bits in mind, you know, are writing code

8:09

that is uh like aligned with our stack,

8:12

like you know, they're able to ask the

8:13

right questions and sort of interject

8:16

when the model is going off the rails,

8:18

right? Uh like you know, an engineer is

8:20

able to do that because they have done

8:23

that for, you know, all the time they've

8:24

been engineering. You know, when you're

8:26

doing PR reviews or when you're

8:27

mentoring an engineer or like you know,

8:29

when you're questioning something you

8:30

read on the internet or or like you

8:32

know, you're responding to to a Slack

8:34

message and you know, conversing with

8:35

other humans. This is a muscle memory

8:37

that they have built over time. But now,

8:40

you've got this different class of

8:41

applications that is completely being

8:43

built by non-engineers.

8:44

Um you know, and that is something that

8:46

we've uh you know, had to put

8:49

uh like you know, we've had to like

8:50

really think about how we serve this uh

8:53

these these tools to these people

8:55

because effectively, we are now making

8:57

the cost of writing code to zero. But

8:59

how do we make sure that you know, folks

9:01

are aware that they shouldn't be hosting

9:03

these tools um like you know, on the

9:04

public internet without authentication

9:06

and authorization. Uh you know, they

9:08

shouldn't be signing up to a random

9:09

third-party service uh

9:11

uh to go and host these tools or like

9:13

you know, they shouldn't be writing

9:14

stuff that has like you know, injection

9:16

capabilities or like you know, uh

9:19

uh

9:20

you know, uh writing stuff in in a in a

9:22

in a non-idiomatic language. Um you

9:25

know, if you write something in Prolog

9:27

uh and you get code code generated, you

9:29

know, that's all fine and dandy, but

9:30

like you know, then hosting it becomes

9:32

an absolute pain in the rear

9:34

uh and it becomes unmaintainable um for

9:36

anyone who doesn't really know Prolog.

9:39

So, there's all of these different like

9:40

nuances that you've got to you've got to

9:42

really think about.

9:43

Um,

9:44

as we release these tools to the rest of

9:46

the organization, we really want to

9:48

empower them, but like, you know, they

9:50

don't have the

9:52

baseline engineering mindset. And what

9:54

we don't want is we don't want them to

9:55

say, "Okay, right. What we will do is we

9:56

will deal with this problem when you

9:58

create a PR, right? Like, you know, the

10:00

PR will be the block in the pull request

10:01

will become the blocking mechanism."

10:03

Because ultimately then we're just going

10:04

to have our engineers spending all of

10:06

their time that we have helped unblock

10:09

reviewing code that has been generated

10:12

using AI.

10:14

And, you know, that's not intended to be

10:15

reviewed. Instead, how can we make sure

10:17

that the boundary of

10:19

you know, harm that can be done is

10:22

reduced as much as possible.

10:24

You know, it based on on what has been

10:26

generated.

10:28

And, you know, some of the risks are

10:30

like, you know, multi-faceted. For

10:32

example, you can host these things on

10:34

like, you know, the software container

10:35

or a VM or what have you, right? But, if

10:37

it's a HTML page and you also got to

10:39

think about the client sandbox as well,

10:41

right? You're hosting this page to a

10:42

human. You know, if that is calling out

10:44

to a random like third party

10:47

and, you know, exfiltrating data that

10:48

way, then effectively you've just moved

10:50

the attack vector from the server

10:52

exfiltrating data, which you might have

10:54

sandboxed away to the client

10:55

exfiltrating data. Ultimately, your

10:57

data's still gone.

10:59

So, how do you make sure that you build

11:00

the right protection mechanisms? And the

11:02

protection mechanisms are there. Like,

11:04

you know,

11:05

uh, with the with all the tools that we

11:06

have available on our on our day-to-day

11:07

basis.

11:09

How do you make sure that you embed

11:10

those

11:11

as like a de facto

11:13

in in these tools being brought. And

11:14

that is the key thing that we are

11:16

thinking about before unlocking all of

11:18

these tools. So, right now we have

11:20

opened up access to all of these tools

11:22

across the organization,

11:24

but with caveats and limitations on the

11:26

kinds of stuff that they can build and

11:28

the kinds of stuff that they can host.

11:31

That makes tons of sense. Like I I feel

11:33

the same way and we've heard it often

11:34

quoted as the cost of generating code is

11:36

going to zero, but the cost of ownership

11:38

has never actually been higher. Like

11:39

you've got to make sure that you have

11:41

complete understanding and visibility

11:42

across all of these different

11:43

applications, types, use cases. And

11:45

previously you kind of had this natural

11:46

bottleneck in engineering, right? And

11:48

that bottleneck is disappearing in that

11:49

people can kind of go forth and deploy

11:52

things in in ways maybe we didn't even

11:53

think about and uh the you know, the

11:55

risk obviously steps up. And there's now

11:58

people who have access to AI tools.

12:00

There's been famously lots of discussion

12:01

around mythos, you know, kind of

12:03

scanning, looking for issues and stuff.

12:05

So it's more important both sides have

12:07

uh you know, really good um

12:09

uh

12:09

security appetite. Like you need to make

12:11

sure that internally you've got a great

12:13

um security story and also those folks

12:15

externally like they um

12:17

you know, don't have access to tools I

12:19

guess that that the rest of the market

12:20

doesn't. So it's a really interesting

12:21

time and for that specific problem,

12:23

right? So one thing I want to come back

12:24

to Sohaib is we talked a little bit

12:25

about you having a really opinionated

12:28

stack.

12:29

Um which I love and it's also actually

12:30

something I've always really admired

12:32

about uh about Monzo is um

12:34

you know, that there's always kind of a

12:37

a clear uh golden path if you will of

12:39

how to how to do things. Would love to

12:41

hear a little bit more about that. Like

12:43

how you think about having that

12:44

opinionated setup, how you evolve it,

12:45

and then potentially how AI has uh been

12:48

helpful or hindered you in in kind of

12:50

carrying on that journey that you've had

12:51

since very early in Monzo's engineering

12:53

days.

12:54

Yeah, that that is an excellent question

12:56

and a quite a loaded one. So I'll try

12:57

and cover all of the aspects um like you

12:59

know, as as uh as swiftly as I can. So

13:02

um yes, we have a very opinionated

13:04

stack. We have a mono repo um especially

13:06

for like our back end um but this also

13:08

applies to other technologies as well.

13:10

We have a bunch of mono repos um

13:12

around this area and effectively our

13:14

goal is to have a centralized way of

13:16

doing things. So we've invested over the

13:18

last decade in having for example like

13:20

service generators and like code

13:21

generators and like analysis checks and

13:24

like you know, just effectively having

13:26

one structure fits all, which is

13:28

absolutely fantastic for AI and LLMs to

13:30

hoover in because we can say, "Here are

13:33

3,000 We have, you know, over 3,000

13:35

microservices in production in and in

13:37

for our in our backend monorepo. Here

13:39

are 3,000 examples of how you write code

13:41

here at Monzo.

13:43

You know, don't like, you know, go and

13:45

invent a new way or an import like a

13:47

bunch of third-party dependencies

13:48

because it is very likely that this has

13:50

been solved in another part of the

13:53

codebase and you are able to the copy

13:55

the implementation or abstract the

13:56

implementation away or use an existing

13:58

vendor library or what have you, which

14:01

is absolutely fantastic because it means

14:02

that we get a lot of reusability within

14:04

our platform

14:06

while still maintaining like, you know,

14:07

our existing structure and guardrails.

14:09

And

14:10

we're still able to have like all of the

14:12

like guardrails,

14:14

you know, with static analysis checks

14:15

and authorization checks and

14:17

authentication checks like still apply,

14:19

right? You know, if we're using, for

14:20

example, a bunch of like different HTTP

14:24

serving library in

14:25

in Go, which is the predominant language

14:27

that we use in our monorepo, right? Then

14:30

we need to write static analysis checks

14:31

to make sure you're doing the right

14:32

authorization and authentication

14:34

mechanisms for all of those different

14:36

systems. Whereas we have very

14:37

opinionated HTTP library is the one that

14:40

is used across all services. We're able

14:41

to write these checks once and have

14:43

everyone benefit.

14:45

And leading on to that, this kind of

14:46

stuff really, really helps when it comes

14:48

to migrations. I think that has been a

14:50

huge unlock.

14:52

You know, for example, let's say that

14:53

you want to deprecate a method or like,

14:55

you want to move folks towards a new

14:57

implementation. Maybe the old way of

14:59

doing things was was not the recommended

15:01

way, bad patterns or like, you know,

15:03

something that is particularly missing.

15:05

You know, there would have been a lot of

15:06

like

15:07

hand-rolling you would have had to have

15:08

done to either write like a go like

15:11

migrator, like, you know,

15:13

a service migrator or what have you that

15:15

would parse the AST and regenerate new

15:17

code

15:18

or or something like that. And AI has

15:20

just reduced that barrier to entry to

15:22

zero, right? You're able to write really

15:24

complex analysis, really complex checks,

15:28

um really complex code rewriting,

15:30

uh the code mod style tools, um you

15:32

know, go native, and you know, just even

15:33

standard code mod tools.

15:35

You know, at at the at the uh click of a

15:38

finger,

15:39

uh which is absolutely revolutionary.

15:41

Like, you know, this is how I've been

15:41

using Claude on a database. Like, you

15:43

know, you have an incident, um you know,

15:46

code-related incident, and you know, you

15:48

want to go add a check for it. You know,

15:49

the cost of writing the check has now

15:50

become zero. You write the check once.

15:53

Uh you know, you you're not spending

15:55

significant amounts of time in AST code.

15:56

You're able to give it a bunch of test

15:58

cases

15:59

um of like, you know, success and

16:01

failure, and then like the whole

16:02

engineering community is able to

16:03

benefit, uh which I think is really,

16:06

really powerful.

16:07

Um

16:08

So, I think these things have uh like

16:10

really, really helped is that like AI

16:12

and LLMs are able to reference um this

16:14

like, you know, really, you know,

16:16

well-curated code base that we have

16:18

invested um significantly in. Now, the

16:22

downside of that is that again, that

16:24

opinionated nature. Um so, I take for

16:27

example like our web mono repo.

16:29

Uh our web mono repo um

16:31

you know, follows for example the

16:33

current version of like React and like,

16:35

you know, a bunch of other like

16:36

frameworks. Um and especially what I

16:39

found with LLMs, or what has been my

16:40

personal experience, is that um it is

16:42

very um

16:44

very eager to go just pull in other

16:46

dependencies or go and upgrade your your

16:48

node packages and like, you know, just

16:50

go and upgrade to like a later version

16:52

of React. And um you know, that same

16:54

level of checking is uh then sort of

16:57

falls over uh because not everything has

16:59

been brought along for the ride. Uh

17:01

especially like, you know, as the web

17:03

ecosystem moves extremely sporadically

17:06

um

17:07

and very, very rapidly as well. Like,

17:09

there's a new hot framework on the

17:10

block, you know, nearly every other week

17:12

now

17:13

um that that people want to adopt. So,

17:15

the um

17:17

like, you know, the

17:19

I think there's there's a benefit, like

17:20

you know, especially when you have like

17:21

a well-curated codebase. And you know,

17:23

we get a lot of benefits, ecosystem

17:25

benefits with like choosing Go and like,

17:27

you know, Go's type safety and the like

17:29

amount of investment that we've put in.

17:31

But then there are other areas where,

17:33

like you know, it would just go

17:34

completely off the rails and get stuck

17:36

and, you know, get get into a loop and

17:38

just decide, "Okay, right, I'm going to

17:39

go like, you know, go pull in this

17:41

third-party dependency."

17:43

And then, like you know, just cold havoc

17:45

in your in your like well-curated

17:47

garden.

17:49

And, you know, at that point it's just

17:51

spent a whole bunch of tokens and, you

17:52

know, the engineer then has to bail it

17:53

out

17:55

at some point. I think that has been one

17:56

of the um,

17:58

things that we really want to solve for

18:00

that I don't think we've got a perfect

18:01

answer for before unleashing this as

18:04

like a background capability. As, you

18:07

know, one of the things that we're

18:07

investing very heavily in is as a sort

18:09

of background agents. You know, you're

18:10

able to kick off a task, you know, go

18:12

for your commute, go for your, like

18:14

whatever, you know, you don't need to be

18:15

hands at keyboard. But if you've got

18:17

something churning away for like two

18:18

hours consuming, you know, hundreds of

18:21

thousands of tokens and it just goes

18:23

down a completely absurd path, you know,

18:26

where the engineer should have bailed it

18:27

out 5 minutes ago. So, 5 minutes in to a

18:30

2-hour long-running like, you know,

18:33

side call.

18:35

You know, that that is a problem, right?

18:36

Cuz that's just cost us a bunch of money

18:38

and, like you know, wasted a bunch of

18:39

time

18:41

that, like you know, the engineer could

18:42

have used better if we had prompted it

18:44

better.

18:46

I think this is this is an interesting

18:49

dilemma because we still want to

18:51

maintain control over our code. Whereas,

18:54

you know, you look at a lot of the

18:55

projects that, like you know, people are

18:56

using these for, like these are entirely

18:59

greenfield projects, you know, you kick

19:00

off a repo, you don't care about the

19:02

code, it's a pet project, you don't care

19:04

what version of what package it's using,

19:07

right? What you want is something

19:08

rendered on the screen.

19:09

You know, this is how I start projects

19:12

that are like completely sort of green

19:14

field. Like you know, even personal

19:16

projects as well. Which is like here's a

19:17

clean repo. I don't care what you do. I

19:19

don't care what libraries you import as

19:21

long as you know, I'm not going to get

19:22

sued for them because you put in

19:24

proprietary code.

19:26

You know, go go forth and and develop.

19:29

And then like you know, I I will go and

19:31

figure it out later.

19:33

I don't think we have that luxury when

19:34

you still want to maintain control over

19:36

your you know, overall code base and the

19:38

health of the mono repo.

19:41

Makes tons of sense. And I love that we

19:42

naturally kind of arrived at background

19:44

agents as well. And you've talked about

19:46

one of the downsides of them, right? Is

19:47

the problem with people not being at

19:49

keyboards and like longer trajectory

19:50

tasks is they also can go off piste and

19:53

spend a lot of money and not actually

19:54

achieve the outcome and you're not

19:55

necessarily watching it until it comes

19:57

back.

19:58

In the sort of our pre-discussion you

20:00

mentioned that you'd kind of taken a lot

20:01

of inspiration from some of the stuff

20:03

Uber, Ramp, and the teams publishing the

20:05

open about their journey on background

20:06

agents have done. So I'd love to kind of

20:08

hear a little bit more about like what

20:09

resonated for you about their journey

20:11

and kind of what you're thinking of

20:12

applying over at Monzo.

20:14

Yeah, that right now this is like where

20:17

where as of right now

20:19

as of when we're recording this in sort

20:20

of early May,

20:22

you know, this is still something that

20:23

we've not fully cracked yet. So we've

20:26

done the natural thing that I think most

20:27

organizations have done which is you

20:29

know, effectively have remote

20:31

execution of models and we've interfaced

20:34

with a Slack which is what we use

20:36

as an organization on a day-to-day

20:37

basis. But I think the real unlock is to

20:41

open this organizationally wide is to

20:43

effectively have like a centralized

20:45

interface. You know, some form of UI,

20:48

some web tooling or what have you,

20:50

you know, which brings in that

20:52

organizational context alongside the

20:54

agent.

20:55

And you can see that this is what most

20:57

model providers are trying to build like

20:58

you know,

20:59

Notion

21:00

has built it with Notion AI. You know,

21:02

where you can connect your whole

21:03

organizational context and do that for

21:05

documentation. Google have done that

21:07

with Gemini. Um you know, um and what I

21:09

take strong inspiration is that you look

21:11

at Rampen and Uber, they have built that

21:14

for their organizational context. Um

21:16

this is uh one of the reasons why, you

21:19

know, we take strong inspiration from

21:20

off-the-shelf products, you know, Notion

21:22

AI and Gemini and things like that. But

21:24

effectively, they are built for a

21:25

general-purpose market. Whereas, you

21:27

know, we have significant investment in

21:29

the way that we do things

21:30

organizationally. And you know, Uber and

21:33

Ramp and and a bunch of other companies

21:35

Stripe, um like you know, they have

21:36

spent significant amounts of time

21:38

building the chrome and bringing them

21:40

building the interface so that it serves

21:43

their organizational needs. And it's

21:45

been a real unlock for these

21:46

organizations. When I speak to people on

21:48

the ground, um you know, if they want to

21:50

go and build prototypes, they've been

21:51

able to integrate the way that they

21:53

develop these systems and abstract uh

21:55

these things away so that, you know,

21:57

effectively, you have a lovable style

21:58

interface uh which will go and generate

22:00

code and package it up and go and deploy

22:03

it onto their infrastructure seamlessly

22:05

uh and bringing in all of that

22:07

organizational context rather than

22:08

plugging in a pure third-party platform

22:12

uh and you know, calling that your your

22:14

interface. So, I think this is what we

22:16

are working towards and what we will be

22:18

working towards

22:19

um is sort of owning that that interface

22:22

space. Now, of course, you know, we will

22:23

be relying on

22:25

third parties like model serving and

22:27

like, you know,

22:28

uh like, you know, uh actually running

22:29

some of these capabilities. But

22:31

ultimately, that is infrastructure

22:33

detail, right? Like, you know, it's a

22:34

very important infrastructure detail,

22:36

but that is infrastructure detail. But

22:37

what the user sees, the user effectively

22:40

gets a feel of a very um like holistic

22:43

view of how we do AI and LLMs with all

22:46

the organizational context built in. Um

22:49

I look to sort of think of it akin to

22:50

like, you know, how our consumer

22:52

product, like our Monzo app um like, you

22:54

know, when you when you sort of like

22:55

working with payments and you have a

22:57

debit card and a credit card and a and a

22:58

bunch of other stuff, right? Like, you

23:00

know, a bunch of stuff goes in behind

23:01

the scenes, but the consumer doesn't

23:03

need to think about, you know, the

23:04

complexities [snorts]

23:05

of MasterCard as have TPs and

23:07

reconciliation and maintaining a ledger

23:09

and stuff. All this stuff is super

23:10

critical, right? It's like absolute

23:12

foundational.

23:13

Uh and that's what we need to build is

23:15

that foundational capability, you know,

23:17

build and buy, um you know, there's

23:19

there's a mixture of both uh for those

23:21

capabilities.

23:22

Uh and then like provide our opinionated

23:25

interface on top that ties all of these

23:27

tools together

23:28

um rather than just adopt tools uh from

23:31

third parties uh and just plug them in

23:33

and say that we're one and done.

23:35

That all makes tons of sense.

23:37

One thing I've always admired about um

23:39

Stripe, Ramp, and also Monzo is, you

23:41

know, you're heavily regulated, but

23:43

still always at the forefront of tech.

23:45

And one thing I guess I'd love for you

23:46

to maybe give a bit of a flavor for is,

23:49

you know, if for a lot of companies

23:50

adopting background agents, it's

23:51

probably going to be quite easy. Maybe

23:53

they're a small company, they don't have

23:54

a whole bunch of regulation. I might

23:56

just be adding a couple of things in

23:57

Slack and, you know, a very simple open

23:59

source or off-the-shelf solution might

24:00

work for them. Can you give a little bit

24:02

of a flavor of why that's not the case

24:04

for you and some of the challenges that

24:05

you have to deal with in your space that

24:07

people may not be aware of?

24:09

Yeah, absolutely. Um

24:11

there's nothing regulation that stops

24:13

you from adopting AI. Like, you know,

24:15

and I think like, you know,

24:17

you know, that is a statement that I say

24:18

like, you know, for all things again,

24:19

there's nothing in regulation that stops

24:21

you from adopting cloud, tech, like, you

24:23

know, an open source framework, you

24:24

know, a third-party framework or like,

24:26

you know, AI, anything, right? Like, you

24:28

know, regulation doesn't stop you from

24:31

doing anything specific. Regulation

24:33

isn't there to stop you. It provides

24:34

guidance and guardrails, right? Like, it

24:36

provides controls of what you know, what

24:38

you should be mindful of uh and like,

24:41

you know, with penalties if you're not

24:42

mindful about these things. So, a really

24:44

good example is what you do with data,

24:46

right? Like, you know,

24:48

uh if you're for example adopting an AI

24:49

model, it's very easy to then transcend

24:52

that so that, you know, you connect

24:54

maybe your data warehouse using MCP, and

24:56

now your data is being processed by the

24:58

AI model, right? Now, if that AI model

25:00

is then using that data for training,

25:02

right? Like, you know, effectively you

25:04

have now had data exfiltration that is

25:06

being used as part of a training set,

25:08

right? And that is where like, you know,

25:10

penalties will apply, right? And that is

25:12

where like, you know, most organizations

25:14

have become scared and then just blanket

25:16

said, "No, we will do no MCP and no none

25:18

of this stuff." Um [snorts] our

25:20

philosophy and approach is is very very

25:22

different. Instead, we want to enable

25:24

these tools, but have controls about

25:27

where these tools are like, you know,

25:29

hovering up data. Like, you know, what

25:31

are you able to ingest? What are you

25:32

able to put into put into these models

25:35

uh on a on a daily basis? For example,

25:38

you know, we'll say, "Okay, right. You

25:39

know, these tools should be used for

25:40

these things." And we will add controls

25:42

and monitoring and stuff for these

25:43

tools, you know, at at the sort of

25:45

access layer.

25:47

Uh and, you know, we will restrict away

25:48

things like direct big query access or

25:50

like, you know, access to um like, you

25:53

know, uh

25:54

particular like a random third party or

25:56

what have you. Like, you know, we'll say

25:57

like, "You're not allowed to call out to

25:59

to like, you know, web search for, you

26:00

know, this particular term that you

26:02

might have found in a in a in a data

26:04

set."

26:05

Um

26:06

You know,

26:06

there are controls that you can add

26:09

um

26:10

that like, you know, will will will stop

26:13

the harmful stuff from happening. I sort

26:15

of like to think of the flip side, which

26:17

is

26:18

you know, if you don't enable these

26:20

tools for the organization, people will

26:23

find a way, right? People will find a

26:25

way to side skirt

26:27

um you know, getting access to these

26:28

tools. Um even to the point where they

26:30

will pay for it out of their own wallet,

26:33

right?

26:34

Um

26:34

and

26:36

by that fact is a losing battle because

26:38

then you have completely lost control.

26:40

Um

26:41

you know, if someone is like randomly

26:42

signed up to, you know, a a third party

26:44

LLM provider, which one crops up

26:46

literally every week,

26:48

uh you know, they you know, proxies

26:49

access to Claude or whatever

26:51

uh because that is the way that they

26:52

have side skirt. Effectively, you've now

26:54

got shadow IT.

26:56

So, this stance that oh, we can't adopt

26:59

background agents, AI LLMs, whatever the

27:02

the stance is.

27:04

You know, you've got to meet the people

27:06

where they are. They want to be able to

27:08

go and use these tools.

27:09

You know, they might use these tools in

27:11

their spare time or free time or they

27:12

might have used it for like, you know,

27:15

like a side quest or maybe in a prior

27:17

organization. And by restricting these

27:19

tools, people will side skirt and find a

27:21

way,

27:22

which I think is the worst of all

27:23

outcomes. Instead, if we're able to give

27:26

access to these tools

27:27

in a controlled manner, you know, be be

27:30

like, you know, permissive where we can

27:32

be, be restrictive where we really want

27:34

to be to make sure that we maintain

27:35

control,

27:37

like, you know, security and reliability

27:39

control,

27:41

and you know, have really good guidance

27:42

on, you know, where, you know, how how

27:44

data should flow. You know, all the same

27:46

principles apply for, you know, the

27:48

standard day-to-day, like, you know, if

27:50

you get access to a Google Sheet, you

27:52

should not be exporting that Google

27:53

Sheet and sharing it with your personal

27:55

machine or your personal iPhone, right?

27:57

Like, you know, it's a standard best

27:59

practice that that you should not be

28:00

doing. And we have controls around this.

28:02

Um,

28:04

like, to make sure that this doesn't

28:05

happen and we we restrict capabilities.

28:08

Same principle applies here. Make sure

28:09

you have those restrictive capabilities

28:11

where people might be doing something

28:14

wrong, and then you can enable these

28:16

tools and and uh

28:18

you know,

28:19

allow people to go and experiment.

28:22

The thing I love about your answer and

28:23

kind of everything you're saying here is

28:25

I think people think of AI as this brand

28:27

new thing that we need to apply all

28:28

these really special thinking towards,

28:30

but actually it's not, especially in

28:32

platform groups, platform teams like

28:33

you. It's we've seen this story before,

28:35

right? There's a new technology that

28:36

people want to adopt. Maybe it's an

28:38

editor, maybe it's

28:40

a productivity tool, maybe it's a

28:41

terminal, whatever it might be. And you

28:44

know, great tools always have a way of

28:46

kind of bubbling up. People are going to

28:47

find a way to use them. And actually

28:49

much better to meet people where they

28:50

are to support them rather than trying

28:51

to block everything and either cause

28:53

them to go around it or you know, even

28:55

worse scenario, the great people leave.

28:57

They go and work somewhere that will

28:58

support them with you know, that

28:59

productivity workflow. So,

29:01

although there is you know, new

29:04

capabilities that come with AI, like the

29:06

actual story of how you adopt these

29:07

things isn't too different from things

29:09

we've seen before, right?

29:10

Yeah, absolutely. Yeah, like you know,

29:13

the same has applied for like adopting

29:15

anything.

29:16

You know,

29:17

technology, like a third-party supplier,

29:19

what have you? Like you know, the same

29:21

principle applies here.

29:23

Yeah, so keeping that thread running,

29:25

one thing you talked about a little bit

29:26

was

29:27

you're going to start looking around for

29:28

whether you'll build or buy sort of a

29:30

background agent platform.

29:32

You've done this tons of times. I know

29:33

we've spoken a few times about things

29:34

you're exploring building, buying and

29:36

you know, Monzo does tend to be a

29:37

company of builders. Like how will you

29:39

approach that conversation of building

29:41

versus buying and how will you

29:42

ultimately decide what the right thing

29:43

to do is here for for you and how can

29:45

other people

29:46

learn from your approach effectively?

29:48

Yeah, that that is a great question. I

29:50

think the the key thing is that when we

29:53

are building or buying, it's not a

29:55

boolean option, right? There'll be a

29:57

bunch of stuff that we will definitely

29:59

buy, right? Like you know, we're not

30:00

going to be training our own frontier

30:02

models

30:04

or like you know, taking an open-source

30:05

model and like you know, building our

30:07

intelligence on top of it. You know,

30:10

the space is moving extremely quickly

30:12

and you know, to be frankly honest, like

30:15

I don't think most can afford to like

30:17

you know, just with you know, all the

30:19

capacity constraints and the amount of

30:21

stuff we want to build. I don't think

30:22

it'd be a fantastic use of our time or

30:25

our capacity or resources.

30:28

You know, both from an engineering and a

30:30

financial point of view.

30:32

I don't think that would be a great use

30:33

of time.

30:34

So like you know, those things naturally

30:36

naturally like you will be we will be

30:38

buying in.

30:39

And you know, we We strong inspiration

30:41

from the industry as well. you know, the

30:43

the models that we adopt uh and like,

30:45

you know, the the tools that we we

30:47

onboard onto onto our systems. Where I

30:49

think the build equation really comes

30:51

into the foray is um

30:54

what you surface to users in your

30:56

interface contract. So, for example,

30:58

like if we said, "Okay, right. The thing

31:00

what we're going to do is we're going to

31:01

give Claude code access to everyone,

31:03

right?" Uh you know, Claude code is a

31:05

very opinionated sort of batteries

31:06

included product, right? Uh that is

31:08

constantly changing, right? Like, you

31:09

know, it's built by Anthropic. Um you

31:11

know, they they release new features and

31:13

stuff on a day-to-day basis. If that is

31:15

your model of interface, you effectively

31:17

become tied to Claude code, right? And

31:19

what we want to do is we want to try and

31:20

own that interface as much as possible.

31:23

Uh and like, you know, sort of have

31:25

optionality. So, for example, like, you

31:27

know,

31:27

uh

31:28

you know, I've heard that Codex is

31:30

really, really good nowadays for writing

31:31

code. Um you know, Codex doesn't fit

31:33

into Claude code

31:35

um as far as I'm aware. Um like, you

31:37

know, might be a way to work around

31:38

that. Um so, like, you know, what does

31:41

what does that shell look like? And I

31:42

know there's open shells out there with

31:43

open code and pie and and a bunch of

31:45

really, really cool pieces of

31:46

technology. I think that's like one

31:48

really good example. So, what does the

31:49

chrome look like? What does the

31:50

integration look like, right? I think

31:52

for us, our the thing that we will be

31:54

building is that integration layer that

31:57

we then surface to, you know, um

32:00

Monzo engineers and non-engineers alike,

32:03

right? Uh

32:04

which I think is the thing that we are

32:05

uniquely qualified to build. Um you

32:08

know, if you want to really reduce it,

32:10

effectively it is a UI, right? It is a

32:12

UI with the

32:14

background of it, the the like

32:15

infrastructure of it uh is stuff that we

32:18

have like in combination built and

32:20

bought um from from third parties and

32:23

and internally. But, we want to make

32:25

sure that we own the interface. Uh we

32:26

want to uh make sure that we own

32:29

uh you know,

32:30

for example, having all of the telemetry

32:31

and evaluation and like, you know, where

32:33

data goes in and like, you know, um that

32:36

that sort of chrome in the middle so

32:38

that we can make globally optimized

32:40

decisions rather than teams making their

32:43

own localized optimized decisions. For

32:45

example,

32:46

let's say that you're building like you

32:47

know,

32:48

an app like you know, an AI-powered app.

32:51

Like you might say, "Okay, right. Well,

32:52

I really like OpenAI nowadays and I'm

32:53

going to go build it in ChatGPT and give

32:55

a service as a ChatGPT app, right?" Now,

32:58

if OpenAI

33:00

ceases to be a good model, right? Like

33:02

you know, that doesn't like stop the

33:04

existing thing from from working like

33:07

you know, cuz it is working with the

33:08

existing models, but if it stops being

33:10

like you know, a good model or what have

33:12

you, but now we've got this app that has

33:14

accidentally become load-bearing in the

33:16

organization because someone randomly

33:18

built it. I don't think that is a good

33:20

outcome.

33:21

Like you know, it means that like you

33:23

know, we could have had something

33:24

better, you know, the the new

33:27

legion of frontier models like you know,

33:29

might have made this app really really

33:30

better, but we are tied down by having

33:32

this app deeply integrated into the

33:34

chrome of ChatGPT and its web interface,

33:37

which I don't think is is is a good way

33:39

of doing it. So, a lot of work that

33:41

we're doing behind the scenes is like

33:42

you sort of, you know, getting that LLM

33:44

gateway in and like you know, just

33:46

having access to all of these models.

33:47

LLM evaluation is a really hot topic

33:49

right now as well.

33:51

Like you just having all of these

33:52

things, all of these tools so that for

33:54

example, you can say, "Okay, right. This

33:56

thing works, you know, right now today

33:58

with Claude, but I'm able to take the

33:59

data

34:01

Yeah, with with with Opus 4 6. I'm able

34:04

to take the data and what's equally as

34:05

well with ChatGPT Codex 5.4.

34:10

And you know, you're able to switch over

34:12

transparently.

34:14

This also helps with some of the

34:15

reliability story as well,

34:17

like you know, which is also equally as

34:19

important. You know, if these tools have

34:20

become load-bearing in your

34:22

organization, you know, when the daily

34:24

like

34:26

downtime of, you know, any any sort of

34:27

third-party supplier happens, especially

34:30

AI and LLM providers,

34:32

you want to be able to fall back.

34:34

And you know, we want to make sure that

34:36

we have a really, really good story for

34:37

that and do that graceful degradation

34:39

transparently because people are going

34:41

to be relying on these tools day in and

34:43

day out.

34:44

Yeah.

34:46

I love this. We've spent a lot of time

34:47

on this. I think initially there was

34:49

lots of products that kind of just let

34:51

you plug in like bring your own key,

34:52

which makes a lot of sense, too. But I

34:54

think when something does become low

34:55

bearing and it's critical to your

34:56

organization, like I've started to think

34:58

of an LLM provider more like a database

35:00

than a API. So like it's really

35:02

important that uptime is closer to like

35:04

four, five, six nines than it is to, you

35:06

know, have um a cheap fast stream. Like

35:10

the reliability is actually becoming

35:12

core and central and I I don't feel like

35:14

the industry has acknowledged that so

35:15

strongly enough yet. I think having

35:18

every team manage their own like API

35:20

keys, tokens is is a fairly model and I

35:22

think we'll see a shift from that

35:23

personally.

35:24

Yeah, actually I really like that

35:26

analogy. Like, you know, seeing these

35:28

tools as like database providers,

35:29

infrastructure providers.

35:31

And I might steal that as an analogy

35:33

when I describe these things internally.

35:34

But you're absolutely right. Like, you

35:36

know, obviously like a lot of these

35:37

tools are really good at providing

35:39

really great interfaces. Like, you know,

35:40

your

35:42

LLM provider desktop, cloud desktop, you

35:44

know, OpenAI desktop and you know, their

35:47

their web counterparts as well and like

35:49

you know, integrations with all of these

35:50

tools front and center. But

35:53

fundamentally we want to be

35:55

buying in the technology, which is the

35:56

infrastructure behind the scenes

35:58

and effectively have our own opinionated

36:00

stack about what is in front, the crumb

36:03

that is in front, the product that is in

36:04

front, which is like, you know,

36:06

integrating our best practices and like,

36:08

you know, what we want to be able to

36:10

achieve.

36:11

Yeah. There's one last thing you touched

36:13

on a little bit that I would love to

36:15

finish off on is you mentioned a little

36:17

bit around

36:18

um

36:19

like return on investment effectively,

36:21

productivity insights on on how it's

36:24

going. This has been the conversation

36:25

for us this year.

36:26

Last year it I want AI tools, this year

36:28

is okay, prove they're worth the money,

36:30

right? What am I What am I getting from

36:31

this? These for a lot of people or a lot

36:33

of companies, this is going to be their

36:35

biggest line item on on their

36:36

infrastructure expense uh this year and

36:38

next year and every year going forward.

36:40

And so, it's getting more important than

36:41

ever to be able to say, "Okay, well, for

36:43

this you know, for this dollar you gave

36:44

in, you got $2, $5, $10 out."

36:47

Um how are you thinking about this

36:48

internally and uh what's your kind of

36:51

plans for kind of deciding whether a

36:52

tool is is is worth the value

36:54

effectively?

36:55

Yeah, that is that is a conversation

36:57

that we're having internally as well. Um

37:00

you know, I don't think we have a

37:01

perfect answer. Um and to be honest, I

37:04

would be skeptical of any organization

37:06

right now who claims that they have a

37:08

perfect answer to this. Especially like

37:10

you know, to to have a mathematical

37:12

equation say like you know, I put a

37:13

dollar in and I have gotten like an

37:15

amount of dollars out. Um

37:17

yes, okay, if you abstracted away, then

37:19

then maybe that is the case. But, a

37:21

significant amount of the dollars that

37:23

we are spending is experimentation.

37:25

Right? Like you know, is a

37:27

uh like you know, I will put some amount

37:29

of dollars in and hope, you know, pray.

37:32

Hope is the strategy. Hope [laughter]

37:34

is the strategy here. Like you know,

37:35

hope that something useful comes out of

37:37

it. Um you know, of course, there are a

37:39

bunch of areas like you know, where we

37:41

are adopting um like LM technology where

37:44

you know, there is uh a very clear like

37:46

net benefit like you know, helping with

37:48

customers and you know, uh detection and

37:50

like you know, and stuff like that. I

37:52

think for us, the key metric that we are

37:54

looking at is is the velocity increasing

37:57

like you know, the rate of pull

37:59

requests, the rate of like code change.

38:00

I'm just talking about engineering

38:01

metrics here. Uh like you know, the the

38:03

rate of code change, um the rate of like

38:06

you know, being able to ship and deliver

38:07

products. Um I think what is interesting

38:09

is that like you know, um

38:12

like most organization, you know, we

38:13

track all all the Dora metrics and you

38:15

know, we track Dora metrics per engineer

38:17

and and stuff like that. Uh you know, on

38:19

top of platforms like DX and and and and

38:22

things like that. But, effectively, all

38:25

that becomes meaningless, right? You

38:26

know, the cost of raising a PR is

38:28

basically zero, right? The token going

38:30

to do it for you. The cost of writing

38:31

the code is zero. So, fundamentally, you

38:33

know, the the the thing that we are

38:36

measuring is the impact you're able to

38:38

have. Right? Are you able to unblock

38:39

yourself? Are you able to deliver that

38:41

impact with within a faster time frame

38:44

um is fundamentally the the exam

38:46

question that most organizations are

38:48

asking. Um

38:50

and

38:51

I I I I think that is that is a good way

38:53

to frame it. Like, you know,

38:55

a Monzo customer doesn't care how many

38:57

PRs we have shipped, right? You know,

38:59

what they care about is that have you

39:00

been able to ship the PR so that you can

39:01

get new customer features or fix a a bug

39:04

or like, you know, uh you know, correct

39:06

something that, you know, has been has

39:08

been annoying you or like, you know,

39:09

unlocked a new product feature that

39:11

you've been really really excited about.

39:13

Right? That is what you care about as a

39:15

as a Monzo customer. And that is the

39:17

thing that we want to accelerate as

39:19

quickly as possible.

39:21

That makes tons of sense to me.

39:23

Do you think we'll ever get to a world

39:24

with um

39:26

with

39:27

uh AI models where we can get to like

39:28

outcome-based pricing for engineering?

39:30

Is that something you can you see ever

39:32

being possible or do you think it's way

39:33

too hard?

39:38

I'll be honest with you. I I don't know

39:39

the answer. Like, you know, uh

39:41

having with a model having some form of

39:44

predictability, right? And to know that

39:46

okay, like, you know, with these classes

39:48

of tasks, right? Like, you know, I'm

39:49

able to take it, right? Like, you know,

39:51

we have been able to measure that with

39:52

these amount of dollars we put in, you

39:54

know, these are the this is the outcome

39:56

that we we achieve. Uh

39:58

like, you know, with that, you can have

40:00

some sort of form of like outcome-based

40:01

pricing. I think the real interesting

40:03

dilemma is that when a new model comes

40:06

out, that you know, the the amount of

40:08

experimentation and training, you know,

40:10

effectively starts from

40:11

not zero, but like, you know, close to

40:12

zero again, right? Like, you know,

40:14

something has fundamentally changed, you

40:16

know, uh something has fundamentally

40:18

gotten better or worse, you know, a

40:19

particular area is now like you know,

40:21

much faster or much quicker and things

40:23

like that. And even between model

40:25

releases like you know, changes in

40:26

system prompts,

40:28

uh you know, there's a lot of like

40:28

changes happening behind the scenes. So

40:30

effectively, you are you are working on

40:32

a foundation of non-determinism, right?

40:35

Whereas if you for example, control the

40:37

full end-to-end stack, you know, you're

40:38

picking an open model and you're able to

40:40

control releases and and stuff like

40:42

that, you can have some form of like

40:44

outcome-based pricing, which is I have

40:46

got this model like you know, I use it

40:48

for these tasks, you know, for these

40:50

tasks I come in for every dollar that I

40:52

spend on GPU inference, I'm able to get

40:54

this amount of dollar out, which is why

40:55

I've been able to automate it away,

40:57

right? Um

40:58

you know, for these things I think you

41:00

can get to some outcome-level pricing,

41:01

but when it comes to like the latest and

41:04

greatest and the frontier models and you

41:06

know, stuff that you depend on from from

41:07

third parties, I think the jury's still

41:10

out just because there's so much

41:11

experimentation happening.

41:13

Uh we're we're sort of at an really

41:15

interesting point in technology and

41:17

quite a frustrating one as well. Like

41:18

you know, if you if you ask some people,

41:19

which is you know, by the time you

41:21

finally think you've got to grips with

41:23

something, something new is out there,

41:25

right? Like you know, before you you you

41:27

feel like you've figured it out,

41:29

you know, whether whether that is true

41:30

or not or whether you when you get that

41:32

feeling, okay, right, I've now mastered

41:34

it, I've now got you know, gotten

41:35

somewhere, you know, something new is

41:36

out there. Yeah, this is how I felt with

41:38

like the the releases between the Opus

41:40

4.5 to 4.6 to 4.7. Um and you know,

41:44

there's been a a bunch of stuff out

41:45

there like on on like you know, Opus

41:47

4.7's like capabilities.

41:50

Um and like you know, system prompts

41:51

changing and and like you know, it's

41:53

training and and that stuff. I

41:55

whole bunch of discourse out there. Um

41:57

you know, these are fundamentally really

41:59

fantastic capabilities, but

42:01

deterministic is not the word that I

42:03

would use to describe them. And if you

42:04

don't have something deterministic, I

42:06

don't think um you can have a good

42:10

quantification for outcome based for

42:13

pricing or utilization or telemetry.

42:17

Yeah, I think that's a great answer. And

42:18

it's why I mean it's why engineering has

42:19

always been hard to measure, right? It's

42:21

because

42:22

okay, Suhail did 10 Jira tickets this

42:25

week, Matt did five.

42:26

Who's more productive? And the answer is

42:28

I don't really know. We'd have to like

42:29

dig way into the details and figure out

42:31

all sorts of things. And by the time you

42:32

figure it out, like it doesn't matter

42:33

anymore. Move on. So I think it's a it's

42:36

a really challenging space. And I think

42:37

people have this strong desire to to get

42:39

us towards outcome based pricing, but

42:40

I'm the same as you. I I don't see a

42:42

path there, but I'm very curious to see

42:44

to see

42:45

how this plays out. I've seen I saw

42:47

Stripe publish something this week where

42:48

they were charging like $2 to rate your

42:50

API using LLMs. And then I saw Intercom

42:53

do outcome based pricing for the AI

42:55

assisted chatbot. And I see that Claude,

42:58

I think they're charging like $5 or

42:59

something to do like a deep PR review as

43:01

well. So you can see the experimentation

43:03

happening in the space. They're trying

43:04

to find the line where people are happy

43:06

to pay for an outcome even if it is

43:08

non-deterministic, but I don't really

43:09

see how we can generalize that. But I'm

43:11

excited to see others try.

43:13

Yeah, I think what is really interesting

43:15

is sort of that evaluation. Like how do

43:17

you determine what is a good outcome?

43:20

You know, do you build outcome based LLM

43:22

tools that, you know, do test the LLM

43:25

tools?

43:26

You know, how far down the stack do you

43:28

go?

43:29

You know, if you've got something that

43:30

is a binary answer, like you know, this

43:32

is good or this is bad, and you're able

43:34

to deterministically infer that, then

43:36

yes, like you can have outcome based

43:38

pricing. I don't think we are quite

43:40

there for a lot of the usages that LLM

43:42

tools unlock. You know, is this a good

43:44

answer? Is this a reliable answer? Is

43:46

this a deterministic answer?

43:48

Does it answer my question?

43:50

I think we all we have right now is

43:51

proxy metrics across the industry.

43:54

Completely agree. Awesome.

43:56

I just want to thank you so much for

43:57

your time, Suhail. This was like super

43:59

interesting.

44:00

If people want to learn a little bit

44:01

more about Monzo, more about you,

44:03

where can they find you and what do you

44:04

recommend they they pick up on? I know

44:05

you're a prolific speaker, so if there's

44:07

any talks you want to point people at

44:08

specifically as well, we'd love to hear

44:09

it.

44:10

Yeah, absolutely. Um so um I'm Sohaib

44:13

Tahir on all the things online

44:15

uh and uh if you want to get in touch,

44:17

then I am at sohaibtahir.com. Uh so

44:20

please do do drop me a message. I'm

44:22

Sohaib Tahir on everything else on X,

44:24

Blue Sky, all the social networks uh

44:26

that that you might want. Um and for

44:29

Monzo uh yeah, we are at monzo.com. Uh

44:32

we are hiring. Um if you are based in

44:34

Europe and the UK,

44:36

um so please do get in touch. Um and

44:39

yeah, like uh you know, if you've got

44:40

any sort of remarks on on what we have

44:42

talked about today or what is happening

44:44

in the industry, I would love to hear

44:45

it. Uh you'll also find me at a bunch of

44:47

conferences. Uh we'll be at Lead Dev um

44:50

in London soon uh and we'll be at QCon

44:53

as well.

44:55

Awesome. Thank you so much, Sohaib.

44:57

Yeah, thank you so much, Matt, and thank

44:58

you for having me.

45:03

>> [music]

45:09

[music]

Interactive Summary

Matt, representing product design engineering at Ona, interviews Suhaib, a principal engineer at the UK-based bank Monzo. They discuss the platform group's role at Monzo, which centers on managing foundational infrastructure and creating an opinionated architecture to improve developer experience. The conversation delves into how AI and LLMs are revolutionizing engineering, the challenges of adopting these tools within a highly regulated banking environment, and the necessity of building internal integration layers ('chrome') to maintain control, visibility, and reliability. They also address the difficulty of measuring productivity and ROI in an era where code generation costs are near zero.

Suggested questions

4 ready-made prompts