AI Tools in Regulated Banking: Monzo's Control Model | Suhail Patel
1462 segments
Hey there, I'm Matt. I run product
design engineering at Ona, and I'm
joined by Suhaib.
Let me try that again. That didn't go
very well.
Hey there, I'm Matt. I run product
design engineering here at Ona, and I'm
joined by my good friend Suhaib, who has
worked at Monzo for many, many years
now. Uh, if you haven't heard of Monzo,
I don't know where you've been living,
probably under a rock, but Monzo is one
of the UK's like defining banks. Um, as
of 2025, we reported at 12 million
customers. Um, and Monzo's
And Suhaib's been there since
Is it 8, 9 years now?
Yeah, 8 years now. 8 years now. Long
amount of time.
Yeah, and you currently are a principal
engineer, I believe, who sits in the
platform group. Can you tell us a little
bit more about your role in the platform
group?
Yeah, absolutely. Um, so yes, hi, my
name is Suhaib. I am one of the
principal engineers at Monzo, um, and
I've been at Monzo for about 8 years
now. Um, and I help lead the platform
group, uh, as an individual contributor,
and the platform group is responsible
for building Monzo's foundational base.
So, we look after all of the
infrastructure, you know, the the
standard things that you would expect
with AWS and Google Cloud, and, you
know, all the sort of like, cloud-native
relationships and things like that, and
all the cloud-native tools that sit sit
on top of these these, uh, providers.
For example, your Kubernetes and your
Kafka and your your your Cassandras and
all your myriad of technologies that,
you know, a modern engineering
organization would adopt. But, our real
secret sauce is, uh, within the platform
group, we also own the relationship with
engineers, the developer experience that
engineers interface with on a day-to-day
basis. So, our goal is to provide a very
opinionated architecture, uh, and very
opinionated code base, um, and libraries
and tools that engineers uh,
want to adopt. Uh, you know, we we don't
sort of force them to adopt it, because
that is the way to get stuff, uh,
without friction onto the platform,
Uh and comes with a bunch of batteries
included, but you know, there's only so
much forcing you can do. Uh like, you
know, the tools have to be really,
really good for engineers to want to use
them.
Uh and use the standardized stack. Uh
and what we do is we help build and
maintain and provide really, really good
tooling and observability around that
stack, so that engineers can focus on
building the bank and not have to worry
about infrastructure.
There's so many things you said there
that I want to jump off. So, um we'll
start with um how big is the platform
group? It sounds like you've got a
really wide mandate, actually.
Yeah, absolutely. We've got um
I can't even remember the amount of
squads we have because we're adding uh
uh
bits and pieces all over the place. We
span a very wide gamut between um you
know, core platform and infrastructures,
the stuff that, you know, will interface
with the app on a day-to-day basis, to
things like um you know, our data
infrastructure as well, like in our in
our data warehousing estate, uh making
sure that we have really good modeling
tooling
uh for for data modeling and analytics
and and and business intelligence. Um
we've also uh very recently um
uh done a significant amount of
investment in machine learning. Um and
this is traditional machine learning.
Um you know, your neural nets and and
your models and
uh
and the standard stuff, you know, which
is absolutely big in in our domain and
space. Uh and a lot of really, really
good techniques that we can adopt um and
and run with. Uh so, we've a significant
amount of investment in machine
learning, and a lot of that has extended
into into a lot of the really cool stuff
with the AI and LLMs as well. Um and on
the other side, uh you know, we have
teams that deal with developer
experience. Um and also, you know, even
things that are very, very important,
things like FinOps. Um you know, make
sure that we're not spending
egregiously. Um you know, our customer
uh unit cost is uh staying within
control. Um so, yeah, a really, really
wide gamut of uh teams are are included
within within the platform group.
And for the platform group specifically,
do you find that your hiring plans are
trending up since you've seen AI
adoption come in or is it kind of going
flat? You're hiring less? Like how do
you think about hiring in a world where
AI tooling it can can can help you be
way more productive?
Yeah, that is a great question.
Obviously, there's a lot of stuff out
there in the industry which is like, you
know, all these tools are going to
replace engineers.
And you know, these tools are just
really really good. Like, you know, even
the last couple months you know, with
with with all the tools out there, all
of the foundational models and even open
models as well.
You know, the tools have been changed
the way that we work on a day-to-day
basis.
Even me like, you know, early on I was a
bit skeptical about like, you know, the
amount of revolution this would have on
on the day-to-day. But now like, you
know, I don't even open my editors like
my first port of call when I start the
day. The first thing I do is open my
terminal open open cloud code.
So like it is like fundamentally rewired
how I think about my brain.
And you know, this is a journey that
many engineers have have gone on. For
specifically at Monzo, we've just become
more ambitious.
So like, you know, ultimately our hiring
plans are remaining like, you know,
business as usual, but we've just become
more ambitious in what we want to be
able to achieve.
You know, we are definitely hiring more
like, you know, we've got a bunch of
open roles
open and continue to have them open just
because our ambition has gotten bigger
and and bolder and we want to be able to
do more as a result of that. You know,
fundamentally you know, these tools are
not going to replace all of the product
things we want to be able to build.
You know, if these tools could click a
finger and we could have all of our
product stuff built, you know, we would
have more ambition in the product that
we want to build
in general. I don't think this would be
a replacement for the for the
for the ambition that we have and for
the people that we have and will
continue to hire.
I love answer. It's exactly what we've
seen too is, you know, we we actually do
want to do more of it, not less of it,
right? Is um great engineers can now do
great things even faster. So, you know,
who who doesn't want more of that?
Um I'm yet to see an empty backlog. Um
don't know if you've seen one yet, but I
haven't seen one. No, absolutely not.
Like, you know, the backlogs um continue
to get ever bigger. And you know, just
specifically in the world of platform,
um like, you know, uh the amount of
people that we reach is, like, you know,
gotten infinitely larger. You know, our
remit, you know, uh was basically to
make engineers more productive. And now,
like, our our remit has expanded to make
the whole organization more productive.
Um you know, when the cost of writing
code is basically zero, right? Um
especially for prototyping and you know,
one-offs and things like that. You know,
you still want to have a reliable,
secure platform to go ship these things
on. Uh you know, to serve them, like,
you know, to to, you know, host them
internally to to users and to have the
right controls and security and
mechanism. And these things need to come
batteries included by default, have
observability. Uh you know, if you're
running a model and, like, you know, you
want to backtest on some data and you
want to see how that how that data
performed. Um you know, even if you've
come up with an idea and sort of live
coded a prototype before, uh like, you
know, for for anything, you still need
to have all of these mechanisms
available. So, our ambition and the
amount of people that we are serving has
gotten bigger.
That makes tons of sense to me. It's
exactly what we've seen too. How do you
think about rolling AI AI tools out to
um engineers and business users? Is
there a different uh roles or different
ways they have to operate? And what does
it take for a business user to to land a
PR in production, for example?
Yeah.
I don't think we've been perfect at
this. So, like, you know, I wouldn't
take our answer as like the the the the
pinnacle of of perfection. I think like
most organizations we're still figuring
it out, right? Um
like, you know, we've had to be very,
very deliberate about rolling these
things out. Um there's one side which is
like we want to give access to all of
these tools. We've onboarded all of
these tools, like you know, we've gone
through all of the all of the processes
and like you know, gotten access to all
of these tools and rolled them out to
engineers as quickly as possible, so we
can get it in the hands of people and
learn ourselves what these tools are are
capable of doing, right? And for the
entire industry, this is a learning
experience, right? Now, what we've been
able to rely on as we brought these
tools out to engineers is that engineers
have, for example, done their secure
development training, like you know,
have got like all all of the various
bits in mind, you know, are writing code
that is uh like aligned with our stack,
like you know, they're able to ask the
right questions and sort of interject
when the model is going off the rails,
right? Uh like you know, an engineer is
able to do that because they have done
that for, you know, all the time they've
been engineering. You know, when you're
doing PR reviews or when you're
mentoring an engineer or like you know,
when you're questioning something you
read on the internet or or like you
know, you're responding to to a Slack
message and you know, conversing with
other humans. This is a muscle memory
that they have built over time. But now,
you've got this different class of
applications that is completely being
built by non-engineers.
Um you know, and that is something that
we've uh you know, had to put
uh like you know, we've had to like
really think about how we serve this uh
these these tools to these people
because effectively, we are now making
the cost of writing code to zero. But
how do we make sure that you know, folks
are aware that they shouldn't be hosting
these tools um like you know, on the
public internet without authentication
and authorization. Uh you know, they
shouldn't be signing up to a random
third-party service uh
uh to go and host these tools or like
you know, they shouldn't be writing
stuff that has like you know, injection
capabilities or like you know, uh
uh
you know, uh writing stuff in in a in a
in a non-idiomatic language. Um you
know, if you write something in Prolog
uh and you get code code generated, you
know, that's all fine and dandy, but
like you know, then hosting it becomes
an absolute pain in the rear
uh and it becomes unmaintainable um for
anyone who doesn't really know Prolog.
So, there's all of these different like
nuances that you've got to you've got to
really think about.
Um,
as we release these tools to the rest of
the organization, we really want to
empower them, but like, you know, they
don't have the
baseline engineering mindset. And what
we don't want is we don't want them to
say, "Okay, right. What we will do is we
will deal with this problem when you
create a PR, right? Like, you know, the
PR will be the block in the pull request
will become the blocking mechanism."
Because ultimately then we're just going
to have our engineers spending all of
their time that we have helped unblock
reviewing code that has been generated
using AI.
And, you know, that's not intended to be
reviewed. Instead, how can we make sure
that the boundary of
you know, harm that can be done is
reduced as much as possible.
You know, it based on on what has been
generated.
And, you know, some of the risks are
like, you know, multi-faceted. For
example, you can host these things on
like, you know, the software container
or a VM or what have you, right? But, if
it's a HTML page and you also got to
think about the client sandbox as well,
right? You're hosting this page to a
human. You know, if that is calling out
to a random like third party
and, you know, exfiltrating data that
way, then effectively you've just moved
the attack vector from the server
exfiltrating data, which you might have
sandboxed away to the client
exfiltrating data. Ultimately, your
data's still gone.
So, how do you make sure that you build
the right protection mechanisms? And the
protection mechanisms are there. Like,
you know,
uh, with the with all the tools that we
have available on our on our day-to-day
basis.
How do you make sure that you embed
those
as like a de facto
in in these tools being brought. And
that is the key thing that we are
thinking about before unlocking all of
these tools. So, right now we have
opened up access to all of these tools
across the organization,
but with caveats and limitations on the
kinds of stuff that they can build and
the kinds of stuff that they can host.
That makes tons of sense. Like I I feel
the same way and we've heard it often
quoted as the cost of generating code is
going to zero, but the cost of ownership
has never actually been higher. Like
you've got to make sure that you have
complete understanding and visibility
across all of these different
applications, types, use cases. And
previously you kind of had this natural
bottleneck in engineering, right? And
that bottleneck is disappearing in that
people can kind of go forth and deploy
things in in ways maybe we didn't even
think about and uh the you know, the
risk obviously steps up. And there's now
people who have access to AI tools.
There's been famously lots of discussion
around mythos, you know, kind of
scanning, looking for issues and stuff.
So it's more important both sides have
uh you know, really good um
uh
security appetite. Like you need to make
sure that internally you've got a great
um security story and also those folks
externally like they um
you know, don't have access to tools I
guess that that the rest of the market
doesn't. So it's a really interesting
time and for that specific problem,
right? So one thing I want to come back
to Sohaib is we talked a little bit
about you having a really opinionated
stack.
Um which I love and it's also actually
something I've always really admired
about uh about Monzo is um
you know, that there's always kind of a
a clear uh golden path if you will of
how to how to do things. Would love to
hear a little bit more about that. Like
how you think about having that
opinionated setup, how you evolve it,
and then potentially how AI has uh been
helpful or hindered you in in kind of
carrying on that journey that you've had
since very early in Monzo's engineering
days.
Yeah, that that is an excellent question
and a quite a loaded one. So I'll try
and cover all of the aspects um like you
know, as as uh as swiftly as I can. So
um yes, we have a very opinionated
stack. We have a mono repo um especially
for like our back end um but this also
applies to other technologies as well.
We have a bunch of mono repos um
around this area and effectively our
goal is to have a centralized way of
doing things. So we've invested over the
last decade in having for example like
service generators and like code
generators and like analysis checks and
like you know, just effectively having
one structure fits all, which is
absolutely fantastic for AI and LLMs to
hoover in because we can say, "Here are
3,000 We have, you know, over 3,000
microservices in production in and in
for our in our backend monorepo. Here
are 3,000 examples of how you write code
here at Monzo.
You know, don't like, you know, go and
invent a new way or an import like a
bunch of third-party dependencies
because it is very likely that this has
been solved in another part of the
codebase and you are able to the copy
the implementation or abstract the
implementation away or use an existing
vendor library or what have you, which
is absolutely fantastic because it means
that we get a lot of reusability within
our platform
while still maintaining like, you know,
our existing structure and guardrails.
And
we're still able to have like all of the
like guardrails,
you know, with static analysis checks
and authorization checks and
authentication checks like still apply,
right? You know, if we're using, for
example, a bunch of like different HTTP
serving library in
in Go, which is the predominant language
that we use in our monorepo, right? Then
we need to write static analysis checks
to make sure you're doing the right
authorization and authentication
mechanisms for all of those different
systems. Whereas we have very
opinionated HTTP library is the one that
is used across all services. We're able
to write these checks once and have
everyone benefit.
And leading on to that, this kind of
stuff really, really helps when it comes
to migrations. I think that has been a
huge unlock.
You know, for example, let's say that
you want to deprecate a method or like,
you want to move folks towards a new
implementation. Maybe the old way of
doing things was was not the recommended
way, bad patterns or like, you know,
something that is particularly missing.
You know, there would have been a lot of
like
hand-rolling you would have had to have
done to either write like a go like
migrator, like, you know,
a service migrator or what have you that
would parse the AST and regenerate new
code
or or something like that. And AI has
just reduced that barrier to entry to
zero, right? You're able to write really
complex analysis, really complex checks,
um really complex code rewriting,
uh the code mod style tools, um you
know, go native, and you know, just even
standard code mod tools.
You know, at at the at the uh click of a
finger,
uh which is absolutely revolutionary.
Like, you know, this is how I've been
using Claude on a database. Like, you
know, you have an incident, um you know,
code-related incident, and you know, you
want to go add a check for it. You know,
the cost of writing the check has now
become zero. You write the check once.
Uh you know, you you're not spending
significant amounts of time in AST code.
You're able to give it a bunch of test
cases
um of like, you know, success and
failure, and then like the whole
engineering community is able to
benefit, uh which I think is really,
really powerful.
Um
So, I think these things have uh like
really, really helped is that like AI
and LLMs are able to reference um this
like, you know, really, you know,
well-curated code base that we have
invested um significantly in. Now, the
downside of that is that again, that
opinionated nature. Um so, I take for
example like our web mono repo.
Uh our web mono repo um
you know, follows for example the
current version of like React and like,
you know, a bunch of other like
frameworks. Um and especially what I
found with LLMs, or what has been my
personal experience, is that um it is
very um
very eager to go just pull in other
dependencies or go and upgrade your your
node packages and like, you know, just
go and upgrade to like a later version
of React. And um you know, that same
level of checking is uh then sort of
falls over uh because not everything has
been brought along for the ride. Uh
especially like, you know, as the web
ecosystem moves extremely sporadically
um
and very, very rapidly as well. Like,
there's a new hot framework on the
block, you know, nearly every other week
now
um that that people want to adopt. So,
the um
like, you know, the
I think there's there's a benefit, like
you know, especially when you have like
a well-curated codebase. And you know,
we get a lot of benefits, ecosystem
benefits with like choosing Go and like,
you know, Go's type safety and the like
amount of investment that we've put in.
But then there are other areas where,
like you know, it would just go
completely off the rails and get stuck
and, you know, get get into a loop and
just decide, "Okay, right, I'm going to
go like, you know, go pull in this
third-party dependency."
And then, like you know, just cold havoc
in your in your like well-curated
garden.
And, you know, at that point it's just
spent a whole bunch of tokens and, you
know, the engineer then has to bail it
out
at some point. I think that has been one
of the um,
things that we really want to solve for
that I don't think we've got a perfect
answer for before unleashing this as
like a background capability. As, you
know, one of the things that we're
investing very heavily in is as a sort
of background agents. You know, you're
able to kick off a task, you know, go
for your commute, go for your, like
whatever, you know, you don't need to be
hands at keyboard. But if you've got
something churning away for like two
hours consuming, you know, hundreds of
thousands of tokens and it just goes
down a completely absurd path, you know,
where the engineer should have bailed it
out 5 minutes ago. So, 5 minutes in to a
2-hour long-running like, you know,
side call.
You know, that that is a problem, right?
Cuz that's just cost us a bunch of money
and, like you know, wasted a bunch of
time
that, like you know, the engineer could
have used better if we had prompted it
better.
I think this is this is an interesting
dilemma because we still want to
maintain control over our code. Whereas,
you know, you look at a lot of the
projects that, like you know, people are
using these for, like these are entirely
greenfield projects, you know, you kick
off a repo, you don't care about the
code, it's a pet project, you don't care
what version of what package it's using,
right? What you want is something
rendered on the screen.
You know, this is how I start projects
that are like completely sort of green
field. Like you know, even personal
projects as well. Which is like here's a
clean repo. I don't care what you do. I
don't care what libraries you import as
long as you know, I'm not going to get
sued for them because you put in
proprietary code.
You know, go go forth and and develop.
And then like you know, I I will go and
figure it out later.
I don't think we have that luxury when
you still want to maintain control over
your you know, overall code base and the
health of the mono repo.
Makes tons of sense. And I love that we
naturally kind of arrived at background
agents as well. And you've talked about
one of the downsides of them, right? Is
the problem with people not being at
keyboards and like longer trajectory
tasks is they also can go off piste and
spend a lot of money and not actually
achieve the outcome and you're not
necessarily watching it until it comes
back.
In the sort of our pre-discussion you
mentioned that you'd kind of taken a lot
of inspiration from some of the stuff
Uber, Ramp, and the teams publishing the
open about their journey on background
agents have done. So I'd love to kind of
hear a little bit more about like what
resonated for you about their journey
and kind of what you're thinking of
applying over at Monzo.
Yeah, that right now this is like where
where as of right now
as of when we're recording this in sort
of early May,
you know, this is still something that
we've not fully cracked yet. So we've
done the natural thing that I think most
organizations have done which is you
know, effectively have remote
execution of models and we've interfaced
with a Slack which is what we use
as an organization on a day-to-day
basis. But I think the real unlock is to
open this organizationally wide is to
effectively have like a centralized
interface. You know, some form of UI,
some web tooling or what have you,
you know, which brings in that
organizational context alongside the
agent.
And you can see that this is what most
model providers are trying to build like
you know,
Notion
has built it with Notion AI. You know,
where you can connect your whole
organizational context and do that for
documentation. Google have done that
with Gemini. Um you know, um and what I
take strong inspiration is that you look
at Rampen and Uber, they have built that
for their organizational context. Um
this is uh one of the reasons why, you
know, we take strong inspiration from
off-the-shelf products, you know, Notion
AI and Gemini and things like that. But
effectively, they are built for a
general-purpose market. Whereas, you
know, we have significant investment in
the way that we do things
organizationally. And you know, Uber and
Ramp and and a bunch of other companies
Stripe, um like you know, they have
spent significant amounts of time
building the chrome and bringing them
building the interface so that it serves
their organizational needs. And it's
been a real unlock for these
organizations. When I speak to people on
the ground, um you know, if they want to
go and build prototypes, they've been
able to integrate the way that they
develop these systems and abstract uh
these things away so that, you know,
effectively, you have a lovable style
interface uh which will go and generate
code and package it up and go and deploy
it onto their infrastructure seamlessly
uh and bringing in all of that
organizational context rather than
plugging in a pure third-party platform
uh and you know, calling that your your
interface. So, I think this is what we
are working towards and what we will be
working towards
um is sort of owning that that interface
space. Now, of course, you know, we will
be relying on
third parties like model serving and
like, you know,
uh like, you know, uh actually running
some of these capabilities. But
ultimately, that is infrastructure
detail, right? Like, you know, it's a
very important infrastructure detail,
but that is infrastructure detail. But
what the user sees, the user effectively
gets a feel of a very um like holistic
view of how we do AI and LLMs with all
the organizational context built in. Um
I look to sort of think of it akin to
like, you know, how our consumer
product, like our Monzo app um like, you
know, when you when you sort of like
working with payments and you have a
debit card and a credit card and a and a
bunch of other stuff, right? Like, you
know, a bunch of stuff goes in behind
the scenes, but the consumer doesn't
need to think about, you know, the
complexities [snorts]
of MasterCard as have TPs and
reconciliation and maintaining a ledger
and stuff. All this stuff is super
critical, right? It's like absolute
foundational.
Uh and that's what we need to build is
that foundational capability, you know,
build and buy, um you know, there's
there's a mixture of both uh for those
capabilities.
Uh and then like provide our opinionated
interface on top that ties all of these
tools together
um rather than just adopt tools uh from
third parties uh and just plug them in
and say that we're one and done.
That all makes tons of sense.
One thing I've always admired about um
Stripe, Ramp, and also Monzo is, you
know, you're heavily regulated, but
still always at the forefront of tech.
And one thing I guess I'd love for you
to maybe give a bit of a flavor for is,
you know, if for a lot of companies
adopting background agents, it's
probably going to be quite easy. Maybe
they're a small company, they don't have
a whole bunch of regulation. I might
just be adding a couple of things in
Slack and, you know, a very simple open
source or off-the-shelf solution might
work for them. Can you give a little bit
of a flavor of why that's not the case
for you and some of the challenges that
you have to deal with in your space that
people may not be aware of?
Yeah, absolutely. Um
there's nothing regulation that stops
you from adopting AI. Like, you know,
and I think like, you know,
you know, that is a statement that I say
like, you know, for all things again,
there's nothing in regulation that stops
you from adopting cloud, tech, like, you
know, an open source framework, you
know, a third-party framework or like,
you know, AI, anything, right? Like, you
know, regulation doesn't stop you from
doing anything specific. Regulation
isn't there to stop you. It provides
guidance and guardrails, right? Like, it
provides controls of what you know, what
you should be mindful of uh and like,
you know, with penalties if you're not
mindful about these things. So, a really
good example is what you do with data,
right? Like, you know,
uh if you're for example adopting an AI
model, it's very easy to then transcend
that so that, you know, you connect
maybe your data warehouse using MCP, and
now your data is being processed by the
AI model, right? Now, if that AI model
is then using that data for training,
right? Like, you know, effectively you
have now had data exfiltration that is
being used as part of a training set,
right? And that is where like, you know,
penalties will apply, right? And that is
where like, you know, most organizations
have become scared and then just blanket
said, "No, we will do no MCP and no none
of this stuff." Um [snorts] our
philosophy and approach is is very very
different. Instead, we want to enable
these tools, but have controls about
where these tools are like, you know,
hovering up data. Like, you know, what
are you able to ingest? What are you
able to put into put into these models
uh on a on a daily basis? For example,
you know, we'll say, "Okay, right. You
know, these tools should be used for
these things." And we will add controls
and monitoring and stuff for these
tools, you know, at at the sort of
access layer.
Uh and, you know, we will restrict away
things like direct big query access or
like, you know, access to um like, you
know, uh
particular like a random third party or
what have you. Like, you know, we'll say
like, "You're not allowed to call out to
to like, you know, web search for, you
know, this particular term that you
might have found in a in a in a data
set."
Um
You know,
there are controls that you can add
um
that like, you know, will will will stop
the harmful stuff from happening. I sort
of like to think of the flip side, which
is
you know, if you don't enable these
tools for the organization, people will
find a way, right? People will find a
way to side skirt
um you know, getting access to these
tools. Um even to the point where they
will pay for it out of their own wallet,
right?
Um
and
by that fact is a losing battle because
then you have completely lost control.
Um
you know, if someone is like randomly
signed up to, you know, a a third party
LLM provider, which one crops up
literally every week,
uh you know, they you know, proxies
access to Claude or whatever
uh because that is the way that they
have side skirt. Effectively, you've now
got shadow IT.
So, this stance that oh, we can't adopt
background agents, AI LLMs, whatever the
the stance is.
You know, you've got to meet the people
where they are. They want to be able to
go and use these tools.
You know, they might use these tools in
their spare time or free time or they
might have used it for like, you know,
like a side quest or maybe in a prior
organization. And by restricting these
tools, people will side skirt and find a
way,
which I think is the worst of all
outcomes. Instead, if we're able to give
access to these tools
in a controlled manner, you know, be be
like, you know, permissive where we can
be, be restrictive where we really want
to be to make sure that we maintain
control,
like, you know, security and reliability
control,
and you know, have really good guidance
on, you know, where, you know, how how
data should flow. You know, all the same
principles apply for, you know, the
standard day-to-day, like, you know, if
you get access to a Google Sheet, you
should not be exporting that Google
Sheet and sharing it with your personal
machine or your personal iPhone, right?
Like, you know, it's a standard best
practice that that you should not be
doing. And we have controls around this.
Um,
like, to make sure that this doesn't
happen and we we restrict capabilities.
Same principle applies here. Make sure
you have those restrictive capabilities
where people might be doing something
wrong, and then you can enable these
tools and and uh
you know,
allow people to go and experiment.
The thing I love about your answer and
kind of everything you're saying here is
I think people think of AI as this brand
new thing that we need to apply all
these really special thinking towards,
but actually it's not, especially in
platform groups, platform teams like
you. It's we've seen this story before,
right? There's a new technology that
people want to adopt. Maybe it's an
editor, maybe it's
a productivity tool, maybe it's a
terminal, whatever it might be. And you
know, great tools always have a way of
kind of bubbling up. People are going to
find a way to use them. And actually
much better to meet people where they
are to support them rather than trying
to block everything and either cause
them to go around it or you know, even
worse scenario, the great people leave.
They go and work somewhere that will
support them with you know, that
productivity workflow. So,
although there is you know, new
capabilities that come with AI, like the
actual story of how you adopt these
things isn't too different from things
we've seen before, right?
Yeah, absolutely. Yeah, like you know,
the same has applied for like adopting
anything.
You know,
technology, like a third-party supplier,
what have you? Like you know, the same
principle applies here.
Yeah, so keeping that thread running,
one thing you talked about a little bit
was
you're going to start looking around for
whether you'll build or buy sort of a
background agent platform.
You've done this tons of times. I know
we've spoken a few times about things
you're exploring building, buying and
you know, Monzo does tend to be a
company of builders. Like how will you
approach that conversation of building
versus buying and how will you
ultimately decide what the right thing
to do is here for for you and how can
other people
learn from your approach effectively?
Yeah, that that is a great question. I
think the the key thing is that when we
are building or buying, it's not a
boolean option, right? There'll be a
bunch of stuff that we will definitely
buy, right? Like you know, we're not
going to be training our own frontier
models
or like you know, taking an open-source
model and like you know, building our
intelligence on top of it. You know,
the space is moving extremely quickly
and you know, to be frankly honest, like
I don't think most can afford to like
you know, just with you know, all the
capacity constraints and the amount of
stuff we want to build. I don't think
it'd be a fantastic use of our time or
our capacity or resources.
You know, both from an engineering and a
financial point of view.
I don't think that would be a great use
of time.
So like you know, those things naturally
naturally like you will be we will be
buying in.
And you know, we We strong inspiration
from the industry as well. you know, the
the models that we adopt uh and like,
you know, the the tools that we we
onboard onto onto our systems. Where I
think the build equation really comes
into the foray is um
what you surface to users in your
interface contract. So, for example,
like if we said, "Okay, right. The thing
what we're going to do is we're going to
give Claude code access to everyone,
right?" Uh you know, Claude code is a
very opinionated sort of batteries
included product, right? Uh that is
constantly changing, right? Like, you
know, it's built by Anthropic. Um you
know, they they release new features and
stuff on a day-to-day basis. If that is
your model of interface, you effectively
become tied to Claude code, right? And
what we want to do is we want to try and
own that interface as much as possible.
Uh and like, you know, sort of have
optionality. So, for example, like, you
know,
uh
you know, I've heard that Codex is
really, really good nowadays for writing
code. Um you know, Codex doesn't fit
into Claude code
um as far as I'm aware. Um like, you
know, might be a way to work around
that. Um so, like, you know, what does
what does that shell look like? And I
know there's open shells out there with
open code and pie and and a bunch of
really, really cool pieces of
technology. I think that's like one
really good example. So, what does the
chrome look like? What does the
integration look like, right? I think
for us, our the thing that we will be
building is that integration layer that
we then surface to, you know, um
Monzo engineers and non-engineers alike,
right? Uh
which I think is the thing that we are
uniquely qualified to build. Um you
know, if you want to really reduce it,
effectively it is a UI, right? It is a
UI with the
background of it, the the like
infrastructure of it uh is stuff that we
have like in combination built and
bought um from from third parties and
and internally. But, we want to make
sure that we own the interface. Uh we
want to uh make sure that we own
uh you know,
for example, having all of the telemetry
and evaluation and like, you know, where
data goes in and like, you know, um that
that sort of chrome in the middle so
that we can make globally optimized
decisions rather than teams making their
own localized optimized decisions. For
example,
let's say that you're building like you
know,
an app like you know, an AI-powered app.
Like you might say, "Okay, right. Well,
I really like OpenAI nowadays and I'm
going to go build it in ChatGPT and give
a service as a ChatGPT app, right?" Now,
if OpenAI
ceases to be a good model, right? Like
you know, that doesn't like stop the
existing thing from from working like
you know, cuz it is working with the
existing models, but if it stops being
like you know, a good model or what have
you, but now we've got this app that has
accidentally become load-bearing in the
organization because someone randomly
built it. I don't think that is a good
outcome.
Like you know, it means that like you
know, we could have had something
better, you know, the the new
legion of frontier models like you know,
might have made this app really really
better, but we are tied down by having
this app deeply integrated into the
chrome of ChatGPT and its web interface,
which I don't think is is is a good way
of doing it. So, a lot of work that
we're doing behind the scenes is like
you sort of, you know, getting that LLM
gateway in and like you know, just
having access to all of these models.
LLM evaluation is a really hot topic
right now as well.
Like you just having all of these
things, all of these tools so that for
example, you can say, "Okay, right. This
thing works, you know, right now today
with Claude, but I'm able to take the
data
Yeah, with with with Opus 4 6. I'm able
to take the data and what's equally as
well with ChatGPT Codex 5.4.
And you know, you're able to switch over
transparently.
This also helps with some of the
reliability story as well,
like you know, which is also equally as
important. You know, if these tools have
become load-bearing in your
organization, you know, when the daily
like
downtime of, you know, any any sort of
third-party supplier happens, especially
AI and LLM providers,
you want to be able to fall back.
And you know, we want to make sure that
we have a really, really good story for
that and do that graceful degradation
transparently because people are going
to be relying on these tools day in and
day out.
Yeah.
I love this. We've spent a lot of time
on this. I think initially there was
lots of products that kind of just let
you plug in like bring your own key,
which makes a lot of sense, too. But I
think when something does become low
bearing and it's critical to your
organization, like I've started to think
of an LLM provider more like a database
than a API. So like it's really
important that uptime is closer to like
four, five, six nines than it is to, you
know, have um a cheap fast stream. Like
the reliability is actually becoming
core and central and I I don't feel like
the industry has acknowledged that so
strongly enough yet. I think having
every team manage their own like API
keys, tokens is is a fairly model and I
think we'll see a shift from that
personally.
Yeah, actually I really like that
analogy. Like, you know, seeing these
tools as like database providers,
infrastructure providers.
And I might steal that as an analogy
when I describe these things internally.
But you're absolutely right. Like, you
know, obviously like a lot of these
tools are really good at providing
really great interfaces. Like, you know,
your
LLM provider desktop, cloud desktop, you
know, OpenAI desktop and you know, their
their web counterparts as well and like
you know, integrations with all of these
tools front and center. But
fundamentally we want to be
buying in the technology, which is the
infrastructure behind the scenes
and effectively have our own opinionated
stack about what is in front, the crumb
that is in front, the product that is in
front, which is like, you know,
integrating our best practices and like,
you know, what we want to be able to
achieve.
Yeah. There's one last thing you touched
on a little bit that I would love to
finish off on is you mentioned a little
bit around
um
like return on investment effectively,
productivity insights on on how it's
going. This has been the conversation
for us this year.
Last year it I want AI tools, this year
is okay, prove they're worth the money,
right? What am I What am I getting from
this? These for a lot of people or a lot
of companies, this is going to be their
biggest line item on on their
infrastructure expense uh this year and
next year and every year going forward.
And so, it's getting more important than
ever to be able to say, "Okay, well, for
this you know, for this dollar you gave
in, you got $2, $5, $10 out."
Um how are you thinking about this
internally and uh what's your kind of
plans for kind of deciding whether a
tool is is is worth the value
effectively?
Yeah, that is that is a conversation
that we're having internally as well. Um
you know, I don't think we have a
perfect answer. Um and to be honest, I
would be skeptical of any organization
right now who claims that they have a
perfect answer to this. Especially like
you know, to to have a mathematical
equation say like you know, I put a
dollar in and I have gotten like an
amount of dollars out. Um
yes, okay, if you abstracted away, then
then maybe that is the case. But, a
significant amount of the dollars that
we are spending is experimentation.
Right? Like you know, is a
uh like you know, I will put some amount
of dollars in and hope, you know, pray.
Hope is the strategy. Hope [laughter]
is the strategy here. Like you know,
hope that something useful comes out of
it. Um you know, of course, there are a
bunch of areas like you know, where we
are adopting um like LM technology where
you know, there is uh a very clear like
net benefit like you know, helping with
customers and you know, uh detection and
like you know, and stuff like that. I
think for us, the key metric that we are
looking at is is the velocity increasing
like you know, the rate of pull
requests, the rate of like code change.
I'm just talking about engineering
metrics here. Uh like you know, the the
rate of code change, um the rate of like
you know, being able to ship and deliver
products. Um I think what is interesting
is that like you know, um
like most organization, you know, we
track all all the Dora metrics and you
know, we track Dora metrics per engineer
and and stuff like that. Uh you know, on
top of platforms like DX and and and and
things like that. But, effectively, all
that becomes meaningless, right? You
know, the cost of raising a PR is
basically zero, right? The token going
to do it for you. The cost of writing
the code is zero. So, fundamentally, you
know, the the the thing that we are
measuring is the impact you're able to
have. Right? Are you able to unblock
yourself? Are you able to deliver that
impact with within a faster time frame
um is fundamentally the the exam
question that most organizations are
asking. Um
and
I I I I think that is that is a good way
to frame it. Like, you know,
a Monzo customer doesn't care how many
PRs we have shipped, right? You know,
what they care about is that have you
been able to ship the PR so that you can
get new customer features or fix a a bug
or like, you know, uh you know, correct
something that, you know, has been has
been annoying you or like, you know,
unlocked a new product feature that
you've been really really excited about.
Right? That is what you care about as a
as a Monzo customer. And that is the
thing that we want to accelerate as
quickly as possible.
That makes tons of sense to me.
Do you think we'll ever get to a world
with um
with
uh AI models where we can get to like
outcome-based pricing for engineering?
Is that something you can you see ever
being possible or do you think it's way
too hard?
I'll be honest with you. I I don't know
the answer. Like, you know, uh
having with a model having some form of
predictability, right? And to know that
okay, like, you know, with these classes
of tasks, right? Like, you know, I'm
able to take it, right? Like, you know,
we have been able to measure that with
these amount of dollars we put in, you
know, these are the this is the outcome
that we we achieve. Uh
like, you know, with that, you can have
some sort of form of like outcome-based
pricing. I think the real interesting
dilemma is that when a new model comes
out, that you know, the the amount of
experimentation and training, you know,
effectively starts from
not zero, but like, you know, close to
zero again, right? Like, you know,
something has fundamentally changed, you
know, uh something has fundamentally
gotten better or worse, you know, a
particular area is now like you know,
much faster or much quicker and things
like that. And even between model
releases like you know, changes in
system prompts,
uh you know, there's a lot of like
changes happening behind the scenes. So
effectively, you are you are working on
a foundation of non-determinism, right?
Whereas if you for example, control the
full end-to-end stack, you know, you're
picking an open model and you're able to
control releases and and stuff like
that, you can have some form of like
outcome-based pricing, which is I have
got this model like you know, I use it
for these tasks, you know, for these
tasks I come in for every dollar that I
spend on GPU inference, I'm able to get
this amount of dollar out, which is why
I've been able to automate it away,
right? Um
you know, for these things I think you
can get to some outcome-level pricing,
but when it comes to like the latest and
greatest and the frontier models and you
know, stuff that you depend on from from
third parties, I think the jury's still
out just because there's so much
experimentation happening.
Uh we're we're sort of at an really
interesting point in technology and
quite a frustrating one as well. Like
you know, if you if you ask some people,
which is you know, by the time you
finally think you've got to grips with
something, something new is out there,
right? Like you know, before you you you
feel like you've figured it out,
you know, whether whether that is true
or not or whether you when you get that
feeling, okay, right, I've now mastered
it, I've now got you know, gotten
somewhere, you know, something new is
out there. Yeah, this is how I felt with
like the the releases between the Opus
4.5 to 4.6 to 4.7. Um and you know,
there's been a a bunch of stuff out
there like on on like you know, Opus
4.7's like capabilities.
Um and like you know, system prompts
changing and and like you know, it's
training and and that stuff. I
whole bunch of discourse out there. Um
you know, these are fundamentally really
fantastic capabilities, but
deterministic is not the word that I
would use to describe them. And if you
don't have something deterministic, I
don't think um you can have a good
quantification for outcome based for
pricing or utilization or telemetry.
Yeah, I think that's a great answer. And
it's why I mean it's why engineering has
always been hard to measure, right? It's
because
okay, Suhail did 10 Jira tickets this
week, Matt did five.
Who's more productive? And the answer is
I don't really know. We'd have to like
dig way into the details and figure out
all sorts of things. And by the time you
figure it out, like it doesn't matter
anymore. Move on. So I think it's a it's
a really challenging space. And I think
people have this strong desire to to get
us towards outcome based pricing, but
I'm the same as you. I I don't see a
path there, but I'm very curious to see
to see
how this plays out. I've seen I saw
Stripe publish something this week where
they were charging like $2 to rate your
API using LLMs. And then I saw Intercom
do outcome based pricing for the AI
assisted chatbot. And I see that Claude,
I think they're charging like $5 or
something to do like a deep PR review as
well. So you can see the experimentation
happening in the space. They're trying
to find the line where people are happy
to pay for an outcome even if it is
non-deterministic, but I don't really
see how we can generalize that. But I'm
excited to see others try.
Yeah, I think what is really interesting
is sort of that evaluation. Like how do
you determine what is a good outcome?
You know, do you build outcome based LLM
tools that, you know, do test the LLM
tools?
You know, how far down the stack do you
go?
You know, if you've got something that
is a binary answer, like you know, this
is good or this is bad, and you're able
to deterministically infer that, then
yes, like you can have outcome based
pricing. I don't think we are quite
there for a lot of the usages that LLM
tools unlock. You know, is this a good
answer? Is this a reliable answer? Is
this a deterministic answer?
Does it answer my question?
I think we all we have right now is
proxy metrics across the industry.
Completely agree. Awesome.
I just want to thank you so much for
your time, Suhail. This was like super
interesting.
If people want to learn a little bit
more about Monzo, more about you,
where can they find you and what do you
recommend they they pick up on? I know
you're a prolific speaker, so if there's
any talks you want to point people at
specifically as well, we'd love to hear
it.
Yeah, absolutely. Um so um I'm Sohaib
Tahir on all the things online
uh and uh if you want to get in touch,
then I am at sohaibtahir.com. Uh so
please do do drop me a message. I'm
Sohaib Tahir on everything else on X,
Blue Sky, all the social networks uh
that that you might want. Um and for
Monzo uh yeah, we are at monzo.com. Uh
we are hiring. Um if you are based in
Europe and the UK,
um so please do get in touch. Um and
yeah, like uh you know, if you've got
any sort of remarks on on what we have
talked about today or what is happening
in the industry, I would love to hear
it. Uh you'll also find me at a bunch of
conferences. Uh we'll be at Lead Dev um
in London soon uh and we'll be at QCon
as well.
Awesome. Thank you so much, Sohaib.
Yeah, thank you so much, Matt, and thank
you for having me.
>> [music]
[music]
Ask follow-up questions or revisit key timestamps.
Matt, representing product design engineering at Ona, interviews Suhaib, a principal engineer at the UK-based bank Monzo. They discuss the platform group's role at Monzo, which centers on managing foundational infrastructure and creating an opinionated architecture to improve developer experience. The conversation delves into how AI and LLMs are revolutionizing engineering, the challenges of adopting these tools within a highly regulated banking environment, and the necessity of building internal integration layers ('chrome') to maintain control, visibility, and reliability. They also address the difficulty of measuring productivity and ROI in an era where code generation costs are near zero.
Videos recently processed by our community