ByteCast Ep87: Kelly Shortridge
987 segments
This is ACM Bytecast, [music] a podcast
series from the Association for
Computing Machinery, the world's largest
education and scientific computing
society.
>> [music]
>> We talk to researchers, practitioners,
and innovators who are at the
intersection of computing research and
practice. [music]
They share their experiences, the
lessons they've learned, and their own
visions for the future of computing.
>> [music]
>> I'm your host today, Scott Hanselman.
Hi, I'm Scott Hanselman. This is another
episode of Hanselminutes in association
with the ACM Bytecast, and today I have
the honor of speaking with Kelly
Shortridge. She's a chief product
officer at Fastly. How's it going?
>> It's going very well. It's a beautiful
spring day here in New York.
>> It is beautiful spring day. I've gotten
some sunshine today and I feel a lot
better. Everything sucks, but it's just
sucks slightly less when it's sunny.
>> That is true, and things are blooming. I
can't complain.
>> Yeah, absolutely. So, you are the author
of Security Chaos Engineering,
Sustaining Resilience in Software and
Systems, and I spent the weekend reading
the book and trying to understand
where the intersection of chaos
engineering and security engineering is,
because like I remember when Chaos
Monkey was a thing, and I just got to
imagine all the Netflix people running
around pulling cables, and the monkey
was just messing up their stuff.
And now I'm trying to understand the
intersection of security engineering
with chaos engineering. I'm wondering if
you could help me understand that.
>> Yes, I think it's better characterized
by the umbrella of resilience
engineering, if anything.
I've actually had the rare and
delectable pleasure of unplugging cables
from Fastly's pops, of course. Network
continued working perfectly. It is a
thrill, though, I will admit. Part of
the title with the book is with a little
behind-the-scenes tea is chaos
engineering, especially at the time, was
a big buzzword. The book is certainly
more than just chaos engineering. That
is one tool in kind of the resilience
engineering toolkit.
When we think about resilience, it's
absolutely about how do you recover from
failure of any kind and prepare for
what's next? That what's next could be a
threat, but equally it could be a
business opportunity. It could be
massive traffic growth for good reasons
or it's a DDoS. So, security really is a
subset of
the both
surprises, stressors, opportunities, and
threats in a very broad sense that we
need to think about.
>> This might be a dumb question. It could
be a spicy question, but like why call
it security chaos engineering? Is it
because those are fun words? Because
like resilience engineering didn't
wouldn't fly off the shelves? Cuz it
seems very clear that like resilience is
really what we want, but it's just not a
sexy term.
>> I mean, I think this is a classic
tension always when you're trying to
publish, you know, a book or frankly a
movie, you have [laughter] to have some
sort of catchy name.
>> No, that's a great point. Chaos would be
an awesome movie name, but resilience is
like it's more of an A24 movie.
>> Exactly. A24 vibe. It's also, you know,
you know, talked about it at Davos and
it's in the National Association of
Corporate Directors book around
organizational resilience. It is a great
Latin root word for international
appeal, but I think chaos people are
like, "Wait a second. Chaos can be a
good thing? That doesn't sound right."
>> Absolutely.
>> I want to engineer chaos. That'll be
exciting very exciting.
>> Exactly.
>> Now, you have said that security should
be designed for failure, not for
prevention. And I think that's a really
cool way to think about that. Can you
think of an example where there's a
perfectly secure system that still
failed in the real world? Like what's an
example where Oh, it still happened and
we couldn't stop it?
>> I mean, I feel like tons. First, no such
thing as a perfectly secure system,
right? I think there's so many esoteric
failures out there. I think about the
airline industry has learned many years
ahead of software about the intricate
nature of complex systems and all the
failures that could go wrong. But the
example I always think about is the fact
that they designed I forget which
airplane it was, which model. They
designed it with safety in mind to
almost every degree except for the fact
that in a very bizarre scenario, if you
somehow exploded the coffee maker
like boiling it too hot or something, it
happened to be close enough to like the
panel with some cables that it could
cause a critical failure while the plane
was in air.
>> Oh my god.
>> Right. You wouldn't think about that as
like that is the trigger to like a
massive failure that means there has to
be like an emergency landing, but yet
there they were. So, I think looking at
real-world systems and all the just
bizarre ways that they can fall apart, I
remember back in the days with Twitter,
there was cyber squirrel where it talked
about all the power plant failures
caused by squirrels just doing things
squirrels do, you know, and how they
were almost the more threatening
advanced persistent threat because of
the damage that they wrought. I think
there's just so many examples of like
your best intentions, you know, reality
is stranger than fiction. You're not
going to be able to dream up every
scenario that's possible, so you have to
prepare for the idea of okay, things
will go wrong. How do we minimize impact
and make sure we can evolve to like meet
the moment?
>> Yeah. I think it's so important also to
remember that like
and maybe this is also a little spicy,
like it's on you to be responsible for
your own resilience. And I remember in
the early days of the cloud when we were
all trying to get five nines out of
Azure and five nines out of AWS,
it's like, okay, Azure went down. I'm
going to call somebody and yell at them.
And it's the same but how badly do you
want your site to be up
all the time? Do you want it badly
enough that you're going to put it in
both Azure and AWS?
How much do you want a copy of like how
do you make a plane that doesn't crash?
Do you fly two planes next to each
other? And then when one fails, like you
jump to the other plane? Like it is
ultimately on us, is it not? And we just
need to decide how hard to squeeze.
>> I think there is usually a trade-off if
you want to really simplify it between
cost and resilience. Um, to your point,
you know, ultimately redundancy is
multiple paths to get to the same goal
in practice. You need polyglot
applications and systems. That's pretty
expensive to pull off. I do think though
that software has a beautiful luxury we
sometimes don't leverage.
Your point about planes, sometimes you
can run two instances of a service to
like offload capacity in a way you just
can't do with physical systems. Same
thing with simulating failures, too.
Again, I think that's a very responsible
thing to do. A lot of complex systems
wish they could. For instance, you can
actually instantiate, you know, a real
kind of clone of the production system.
You can't replicate a realistic clone of
New York City to see if there's a
certain level of trash blocking sewer
drains, like what level of flooding will
cause like deaths. Like, you can't
simulate that with any degree of ethics.
But, you can in the computer world.
We're allowed to do it. So, I think
that's part of my call to action that
was big in the book is like, okay, how
do we start taking this more serious
seriously and really leveraging the
benefits that the flexibility software
begets gives us. So, I do think there's
to your point, yes, some of it is more
expensive to do, but in another sense,
maybe we should be allocating more spend
towards some of that simulation or just
understanding like the resilience
contours of our systems better when
other industries are just looking at us
shaking us like, why aren't you doing
this? We wish we could do this.
>> Yeah. I've been thinking about
resilience in my own kind of personal IT
life. I assume you have a home lab and,
you know, of various sorts. Right now,
as I talk to you, because I had an
appointment with you, I am on my backup
internet. Turns out, I'm looking at my
UniFi here, my my WAN failed over at at
4:38 a.m. and I have yet to diagnose it.
So, I'm on backup internet right now.
And when I mention that to people, like
Muggles, like regular people, they're
like, you have two internet at your
house?
I'm like, like, this is my job, bro.
Like, I'm here. Like, I got I've been
doing this at this house for 18 years.
It cost me 45 bucks for Comcast as
backup internet. I have my fiber, but
the backup for $45. It only has to fail
once, like today,
and that made it worth the money for the
year because otherwise I would have had
to cancel on you.
And that wouldn't happen.
>> Exactly.
>> Yeah. So, it's like it was a choice. And
I feel like there are teams that think
that they are resilient
until the thing happens. What's the
difference between a security team that
thinks they're resilient and maybe one
that actually is? Like, I feel like
there's a lot of false confidence and
metrics theater that happens.
>> Metrics theater, that could be an
episode in itself. The like, oh, what
percent security coverage do we have?
Nobody knows what that means. It's a
meaningless metric. That is a great
question. I think the giveaway is when
the security team feels a sense of
control, probably means that they don't
have a lot of resilience. Cuz part of
resilience is embracing the fact that
like, there will be things well outside
of your control. So, it's like, how do
you prepare for that?
If you were trying to control everything
and make things as deterministic as
possible, you've already failed, in my
view. Cuz the world is not
deterministic. Humans aren't
deterministic. We would like computers
to be deterministic, but they aren't
fully, at least. It's one of the hardest
problems in computer science is
verifying that the software works the
way that the designer of the program
intended it to. So, whenever I hear a
security team say like, well, we have
full control over the software delivery
life cycle, like,
are you sure?
>> [laughter]
>> I just saw I just imagined a meme'd
version of you and that one guy from HBO
is like, you sure about that?
You sure about that?
>> Yep.
>> Yeah, or even the opposite thing to do.
>> Yeah, you're right.
>> Exactly. Probably means you're investing
in things that make you feel good and
give you that sense of control and not
the things that minimize impact.
>> Yep. Like, it's not It's not directly
security, but that old joke of like
backups always succeed, it's restores
that fail.
>> Mhm, yes.
>> So, it makes me think about like chaos
engineering and infrastructure is about
pulling wires. And yanking wires is very
exciting. But I feel like there's a lot
of pull the wire moments that can happen
in security, but people are too scared
to try.
>> I mean, fear is pervasive in the
culture, and it's a disservice to the
industry and the mission for sure. I
think there are also cases where you
could be starting with smaller
experiments or just testing more basic
hypotheses. My favorite leveraging
actually fastly's compute, which is kind
of like a high-performance serverless,
you can think of it that way. It's just
a little function that strips out
cookies just to see like, "Hey, does
your login site
work?" Same with like off headers. It's
just those basic assumptions you hold
like, "Of course, like we're always
going to require this for the login
page." It's like, "Well, is that true?
Are you sure?"
And the
especially when you can like duplicate
the request, which this prototype did,
you know, it's pretty low impact to the
business to run that experiment. There
are of course things where it's like,
"Hey, rm -rf like the customer
database." Yeah, that's going to be a
pretty poorly designed experiment with
high consequences. But there's like such
a range in between that I think it's
very unfortunate that security
practitioners are
too hesitant to try those experiments,
especially they're very hesitant to
reach out to their peers across the
island like platform engineering and be
like, "Hey, can we conspire on
developing some of these experiments?"
Cuz there are a lot of jointly held
assumptions, too, that aren't always
poked and prodded.
>> You mentioned about determinism and how
computers and software is not as
deterministic as we would love to think
that it is. Not just because the
software pretty much always runs as you
wrote it, but whether or not your intent
was well expressed certainly is a
problem, and then the environment within
which it runs you can't always count on.
But I'm finding that people seem to be
spackling or puttying over their systems
now with what I'm calling ambiguity
loops, which are basically using an LLM
to deal with ambiguity by letting it
fill the ambiguity with randomness. And
I'm curious in your business, when now
people are like running playbooks that
aren't scripts, they are prose. Like a
markdown file is not a script, I think
you would agree.
>> Yeah.
>> Uh how do you feel about that? Is there
a place for LLMs to live in security and
in resilience and in chaos, or do they
just increase chaos and and entropy?
>> It depends. I think
they're only going to be as good as the
corpus that went into them for one. And
so unless you can verify like really
clean code went into it, it's like,
well, you know, it can maybe be a good
basis for actually in some cases chaos
experiments or, you know, specific
configurations, integration test, etc.
What I will say though is to me the more
important litmus test is is this
replacing human judgment? And if it is,
that's probably not a good case for an
LLM. I am pro human judgment and
creativity. I am pretty anti like very
repetitive work, very tedious work where
you don't need that kind of judgment
call. LLMs can be very helpful there. I
think they're also document intelligence
examples, you know, who loves going
through your compliance documents and
pulling out relevant information. Like
LLMs can shine there and that way you
can focus more on strategically, like
are we sustaining resilience? Like what
are the indicators we should be looking
at for that? But asking an LLM is our
system resilient? Probably is not going
to be a great outcome. I think the
markdown example is a little interesting
because I am also pro making security
more accessible. I think we dress it up
in a lot of arcane
key phrases and buzzwords, when really a
lot of people could benefit the industry
and contribute. So maybe simplifying how
they can enter and not requiring
scripting knowledge could be a good
thing. Else I see it as a potential foot
gun. So I feel like I'm a little mixed.
>> So now I hear you. I I'm with you 100%
on the like human judgment. Like it
cannot be overstated. I assume that
you're speaking to universities and
early in career people often and you
give your speeches and stuff. And you
talk to people like, "What should I
learn?" It's like, "You should learn how
to have good taste." How do you learn
good taste? Well, you just got to get in
there and get your hands dirty and do
the thing and start pulling wires and
figure out the system. Certainly, I
don't want to outsource things to human
judgment and toil like keeping a site up
is toil. SRE is toil, but SREs that are
really good at their job are good at
their job because of their judgment. So
there's going to be this constant
tension between the idea that you would
replace an SRE with a markdown file
makes me very nervous. But an SRE agent
that could maybe kick the node and keep
it running while I drive over there that
has value to me.
>> Yes, buying capacity and buying time I
think is a great use case as well to
your point. How do we help the human
engage better with the system? Even just
like the rubber duck problem solving
that can be a useful thing for the LLM.
That does require quite a bit of
expertise already though.
>> I love that you brought that up. I was
talking to someone recently and I gave a
whole talk and I sort of brought up
rubber duck debugging and like no one
got it. I'm like, "Am I like am I onk
suddenly?" Like no one gets. This is not
a generational thing. Like talking to
the duck. Yeah, your face is saying the
same thing. Like like those are
important moments to you're talking to
yourself in the mirror except now the
mirror can talk back and that's really
cool. I find that to be super helpful.
Have you used LLMs in that context like
figure your thoughts out and talk to
yourself?
>> Sometimes I use my cats more often for
that cuz they're they give a specially
like judgy looks that make you really
question, you know, what you're throwing
down. I think LLMs can be helpful
though. Especially I see a lot of people
struggle to get buy-in. This is kind of
getting into corporate type stuff, but
especially international companies where
it's like, "Hey, I want to get buy-in on
this resilience initiative. How is this
going to resonate across different
cultural contexts?"
>> Ooh, that's a good one.
>> Right? Like is chaos perceived
negatively in certain nations versus
others? I can tell you for instance when
I talk about deception and using that as
a technique for resilience engineering,
security engineering, American security
practitioners, not universally, tend to
result in like, "Well, we don't want to
be the bad guys." Now, the EU, they're
like,
"Tell me more. Yes, please. Like we want
the Sutter Fudge here." So, that's kind
of fascinating, right? And I feel like
that's an interesting like twist
>> Mhm.
>> um that I found really useful with LLMs.
>> I like that. That that I I didn't think
about that. Yeah, you're right. It does
It does see broader than than we do and
it can challenge your assumptions,
especially if you tell it challenge my
assumptions as opposed to telling you
that you're absolutely right.
ACM Bytecast is available on Apple
Podcast, Google Podcast, Podbean,
[music]
Spotify, Stitcher, and TuneIn. If you're
enjoying this episode, please do
subscribe and leave us a review on your
favorite platform.
You probably [music] work with big
companies like a banks and slower-moving
things, health care. They're a little
more conservative. I'm curious, is there
an example where traditional compliance
actively makes systems less secure,
where they think that they're checking
boxes, but they're actually hurting
themselves?
>> Yes, actually a frequent co-conspirator
of mine, Zulfikar Dykstra, wrote a
paper, not with me. It's an excellent
paper about that exact topic. I think
specifically covers HIPAA and maybe one
of the others that shows that it being
more compliant doesn't result in better
security outcomes.
I'm very much of the view and I've tried
to caution regulators as well as like
well-intentioned regulation in this
space very quickly calcifies and
ossifies. Like it's what helps in year
zero through maybe even
year three may end up actually eroding
resilience long-term. Great example,
I'll keep the person not a very
innovative CISO, had to explain, I think
over a few years to his auditors, like
actually it's a great thing that we
don't allow SSH access anymore cuz
that's what attackers love. They love
when you leave the door open like that.
But on the little compliance checklist
for the auditors, they're like, "Okay,
but it says you're required to have SSH
access." He's like, "Okay, but the
security outcome is now better." So,
that's where sometimes it can actually
hold companies back
even big companies who do want to
innovate by basically tying them to
their investments and their spend to
just checking those boxes, which is a
disservice to their overall mission.
>> Yeah, I always want to assert
assumptions. I feel like cuz I work at
Microsoft in my day job that my
ignorance is kind of my superpower cuz
someone will throw me in a new situation
and I'll see a checkbox like must
include SSH. I'm like, "But why?" Like
and they don't they no one knows. Like I
don't know like 13 years ago someone
wrote that checklist and now it's a
thing.
And then investors see it and compliance
people see it and that checkbox is the
thing that stands between you and some
certificate some badge.
And that's a problem.
>> It's a huge problem and actually bring
up a kind of elegant point. If you look
at what resilience means across all
sorts of complex systems but also the
ones we're talking about here,
a lot of when a system is stuck in let's
say it's not elegant but like a a less
resilient or like unresilient state or
fragile state is because a lot of the
processes and practices that they have
in place, to your point, are from an
equilibrium that no longer exists.
Right? The status quo has moved on. The
practices haven't.
And so you're just continuing to erode
resilience as you like stick to this old
world and have not adapted to the new
one and the new context. The same with I
always hear CISO to be like, "Well, once
we patch the vulnerability or fix it,
then like it's fine." It's like, "Well,
if it actually resulted in an outage or
a breach, it's not actually fine because
you still haven't addressed the
underlying impact. You've just patched
over the one way attackers got in. They
haven't adapted to that new paradigm and
that new equilibrium, which is hard.
It's updating your mental model of the
system, which is not easy. That's why it
is so important to have people who will
be like, "Well, why? Why is that?" You
know, just poke and prod.
>> Often CISOs and security teams make
dashboards cuz they want to roll things
up, and the bigger the company, the
bigger the dashboard, and then the CISO
has to really They can't know the entire
stack. The stack is now too deep. So,
what is an example of a misleading
security metric or dashboard that might
cause someone to make a mistake relying
on a dashboard, but it's maybe a
misleading metric?
>> So many metrics. Certainly that security
coverage one or risk coverage. Also, the
number of vulnerabilities
discovered. I'm trying to remember who
it was who talked about this where
actually when things started to get
better, it meant that their application
development teams were servicing more
security issues, which was a good thing.
And it meant that there was more in that
trust, mutual trust between teams, but
it looked like it was getting worse.
>> Yeah, see, that's a great example.
That's the whole thing like, you know,
like all these bugs and all these
security issues. That's this good stuff.
All of that is low-hanging fruit, but
like
they'll assume that something bad has
happened or something has changed, and
they're going to then
correlation and causation are not the
same.
>> Exactly. I think And also, a lot of the
security specific metrics don't tell the
bigger picture.
Includes I think about, you know, the
poor platform engineering teams are
handed a list of a thousand
vulnerabilities. Turns out a lot of them
are in components that aren't even
exposed to the public internet. Should
they prioritize those? Probably not. And
meanwhile, in actually there are
multiple cases of this, so I'll keep
them all anonymous, but it's surprising
actually how often this happens.
Security team will be on them and be
like, "Fix all of these, even if they're
not publicly exposed." And then the
security team actually maintains, you
know, their creds into their, you know,
like
whatever admin system or security system
that has its hooks into everything is
like in a text file on their desktop.
It's like, "Well, what do you think is
actually the bigger issue here in terms
of what attackers could leverage? So,
there's a lot of that kind of attacker
math, attacker calculus that isn't baked
in as well. And I think there's also,
even if we think about business context,
the metrics that a lot of security teams
track and even CISOs track aren't the
ones that the board wants to understand
>> Mhm.
>> or other executives need to understand
either.
>> Yeah.
If I go to fastly.com and I click on
products, you've got all the network
services, all the things that Fastly's
known for. There's a whole section on
security. And there's also, you know,
you have services and folks that you can
hire, professional services, and things
like that. But how should I think about
what security is my responsibility and
what is the responsibility of the vendor
for whom I am paying a lot of money to
make things secure?
>> I think it depends on the vendor. I
think in the case of, let's say, like
some sort of SaaS application, let's
take it sales and marketing, so it's
going to have some of your customer
data, prospect data. Feels reasonable
for the most part that, like, the
encryption and things like that should
be handled by the vendor, for sure.
>> Mhm.
>> There are cases, I'll use like Fastly,
where we have
the platform where you can basically
write code, run code, etc. It's our
responsibility, and we have done this,
to layer in like memory safety by
design. Same with like isolation models,
ensuring like safe multi-tenancy. It's
very much our responsibility. Making
sure that, like, you don't write
vulnerabilities into your own code, it's
like, well, that's probably outside of
our remit. Though, that's where
sometimes pro serve can come in. Things
like, "Hey, you have spun up like a
service on Fastly, and that connects to
a database that you have wide open
without any access control list." Not
really our responsibility cuz we don't
touch that component, right?
>> Right.
>> Um there are ways that we can help with
middleware that runs on our platform. I
do think, though, there's a fundamental
principle
though in the conversation a lot of
people miss, which is like, you have to
own your own dependencies. And so, what
you adopt, you do have to basically
think about it like, well, we have to
assume at some point something will go
wrong with it whether that's a security
issue or not. And I do see a lot of the
like hot potato game happening there.
>> Yeah, that's exactly why I asked you
that question because I think that
people pay a lot of money for a
platform, a cloud platform because they
want as they say a throat to choke. So
like who gets yelled at, right? You
know, get Kelly on the phone. I want to
know what's going on over there. But
then it's of course their thing. Then
they are bringing in who knows unknown
node packages from unknown provenance
and then they haven't thought about
their entire
secure supply chain. But at the same
time, I like your point about a sales
and marketing CRM or something like
that. Is it their job as an app to be in
charge of like AI bot management or DDoS
or even API security? Like that would be
an example where a cloud platform could
secure those endpoints
and hide that. So I like the separation
of concerns
there. But I I wonder if people who are
putting together their own systems think
about that. Like we shouldn't be in
charge of API security. Let Fastly
secure our endpoints. And then they just
have a nice clean bright line or is it
always layered and they would have to do
that? Basically you have two layers.
>> I think it really depends on the company
and their level of resourcing. There are
some companies that culturally want to
own
where things are built more of their
things and so they'll leverage us more
to be able to DIY. There's certainly
others where it's like, well, let's just
use what Fastly has, right? Especially
when it comes to security, putting in
whether it's our WAF or like you said
the AI bot kind of like insights we're
able to surface, like have that in front
of our services, apps, sites, whatever
it is. It really depends on resourcing.
There's also the element of
I'm going to say a newer worlds where I
have seen so many security leaders
thrust into a conversation now where
their CEO, their CMO, their board is
like, "Hey, what AI bots are actually
trying to scrape our stuff so we can
monetize it?" And this is like I've
never had to think about this before.
Trying to DIY that is pretty hard and
hiring that expertise is some of the
most expensive expertise out there right
now. Using a tool probably makes sense
in that case.
>> Yeah, I think that's a great point. I
mean, this is the thing. What do we do
here at the company? We do insurance.
Okay, then why are we doing AI bot
management?
>> Right.
>> That's not our job, you know? So I
always think about the business and I I
think sometimes when we are talking
about all the things we've been talking
about on this show, we don't talk about
like why did we actually make this
software? We made it to solve X business
problem. Therefore, what responsibility
is mine and what can be outsourced by
someone who actually knows what they're
doing, whether it be Fastly or Azure or
AWS. Let somebody who actually cares
about that do that while I focus on the
business problem. Honestly, I don't want
to do the stuff Fastly does. That's why
Fastly is good at it, you know what I
mean?
>> Yes, and it's laying out your own pop
infrastructure, especially in this day
and age of RAM prices being what they
are, especially if you were a small
business, it's quite unlikely that
you're going to be able to do that.
>> Yeah.
>> Nor should you cuz to your point, it's
not your core business. I think it's
I was thinking when you were talking
about the example of the duplicate or
secondary internet, part of that is
because your essentially critical
function hosting this podcast
[clears throat] is
you can record it. However, if you were
to build your own microphone, I'd be a
little bit like, is that actually your
core value add here, you know?
>> No, it's not a good example.
>> Right? It's just understanding like what
matters and what makes you unique as a
business, for sure.
>> Yeah, absolutely. This is totally random
and off topic, but we had a fishing
thing happen at work yesterday where
they they send us fishing emails, but
it's from the red team.
And I was so proud of myself. I was just
like, I don't think that's real. And I
was like, report fishing. And then it
was like, congratulations.
You're one of the better people. So,
know, there's some number of people at
the company that does that. And I'm
always impressed that there's a whole
teams out there
trying to attack us internally that I've
never even met. You know what I mean?
The red hats or the I guess they call
them blue hats at Microsoft cuz our
badges are blue.
Somehow I was just thinking about
there's people trying to create chaos
internally at the company and they tried
to catch me yesterday with a fish. And I
didn't
>> However, what I will say that has gone
wrong in the past and I've spoken
publicly before it was cool to have this
take. People were quite angry. I
remember this was many years ago where I
said, "Hey, it's maybe not a great
thing." Especially during COVID this
happened a lot. To be like
"Here's your surprise bonus plan." And
that's the fishing
like simulation email.
>> Oh, no. That would be awful.
>> Right, but that was happening. So,
that's not
>> That's just That's kind of punitive.
>> I agree.
>> Right. Click here for more money. No,
this was not that. This was more like,
"You have mail waiting for you in the
mail room."
And I'm like, "We don't have a mail
room." You know?
>> There you go. Yeah.
>> That's fair. You're right. I mean, this
is the whole like sprinkling USB keys
around the bank parking lot kind of way
of doing things. It's like, "Beyoncé's
new album." Sprinkle, sprinkle, and then
everyone plugs it in and then owns the
entire bank.
>> Yeah. Something like that. I think there
are a lot of experiments. I think it's
always keeping in mind again the human
element that you don't want to sow
distrust. But again, you can also make
the experiments collaborative
>> Mhm.
>> which that can get really fun because
I've actually You mentioned SREs. When I
talk to like real attackers, they're
generally not scared of security
engineering teams. They're scared of
SREs cuz SREs will like obsess.
Performance There was that one backdoor,
right? Was it xc utils where it was a
guy who's like "Oh, performance degraded
by I think it was less than 1%. What is
going on?" Discovered the backdoor.
>> Right.
That was awesome.
That was pretty cool.
>> Right? So, I think security teams need
to embrace like, "Hey, you may You will
have good ideas, but like there're going
to be other very clever people where if
you say, "Okay, if you got really mad at
the company, how would you attack us?"
They're probably going to have some
interesting ideas that maybe can become
experiments or clue you into some gaps
maybe you have in your current security
investments.
>> Very cool. This is you've given me a lot
to think about. I thought that chaos
engineering and security chaos
engineering and this kind of resilience
was kind of a branding exercise, but it
feels more concrete after having chatted
with you.
>> Yes, I mean, again, it was mostly a
buzzword and it's part of playing the
game that publishers have to play. I
will say though for a very long time
since I was a wee lad, as they say, I've
been obsessed with chaos theory.
>> Mhm.
>> And I do think chaos theory, which is
quite beautiful in the sense of like
systems do have an order to them, but
it's not necessarily predictable. It's
more like obviously like a fractal or
dragon curve in many cases, as any
meteorologist knows well. So, we need to
focus less on again, do we have control
over it? Are we able to predict it? The
quote I love is from Susan Elizabeth
Howe, who's a geologist, who said, "A
building doesn't care whether the
earthquake was predicted or not. It
either stays up or it doesn't."
>> That's good.
>> Right.
>> That's very good.
That's very good.
>> I feel like that's the essence of it.
>> Yeah.
I like that one.
One of my favorites in the in a similar
vein is a Babylon 5, "The avalanche has
begun. It's too late for the pebbles to
vote."
>> That is also very good. Yes. Yes.
>> pebble is like, "I don't like this This
is not a good idea. This is happening.
Sorry, this is happening. So, buckle
up."
>> Yeah. Exactly.
>> Thank you so much, Kelly Shortridge, for
chatting with me today.
>> Thank you for the great questions.
Appreciate it.
>> We have been chatting with Kelly
Shortridge, the chief product officer at
Fastly. This has been another episode of
Hanselminutes in association with the
ACM Bytecast, and we'll see you again
next week.
ACM Bytecast is a production [music] of
the Association for Computing
Machinery's Practitioner Board. To learn
more about ACM and its activities, visit
acm.org. [music]
For more information about this and
other episodes, please do visit our
website at learning.acm.org/bytecast.
[music]
That's b y t e c a [music] s t
learning.acm.org/bytecast.
Ask follow-up questions or revisit key timestamps.
In this episode of Hanselminutes, Scott Hanselman speaks with Kelly Shortridge, Chief Product Officer at Fastly and author of 'Security Chaos Engineering'. They discuss the evolution of resilience engineering, emphasizing a shift from a focus on prevention to designing systems for failure. They explore how security is often misunderstood as a quest for absolute control, whereas true resilience involves preparing for the unpredictable nature of complex systems. The conversation covers the role of LLMs in security, the importance of human judgment, the pitfalls of 'metrics theater' in compliance, and why organizational resilience requires moving beyond outdated operational models.
Videos recently processed by our community