HomeVideos

ByteCast Ep87: Kelly Shortridge

Now Playing

ByteCast Ep87: Kelly Shortridge

Transcript

987 segments

0:01

This is ACM Bytecast, [music] a podcast

0:03

series from the Association for

0:05

Computing Machinery, the world's largest

0:07

education and scientific computing

0:08

society.

0:09

>> [music]

0:10

>> We talk to researchers, practitioners,

0:11

and innovators who are at the

0:12

intersection of computing research and

0:14

practice. [music]

0:15

They share their experiences, the

0:17

lessons they've learned, and their own

0:18

visions for the future of computing.

0:20

>> [music]

0:20

>> I'm your host today, Scott Hanselman.

0:25

Hi, I'm Scott Hanselman. This is another

0:26

episode of Hanselminutes in association

0:28

with the ACM Bytecast, and today I have

0:31

the honor of speaking with Kelly

0:33

Shortridge. She's a chief product

0:35

officer at Fastly. How's it going?

0:38

>> It's going very well. It's a beautiful

0:40

spring day here in New York.

0:41

>> It is beautiful spring day. I've gotten

0:43

some sunshine today and I feel a lot

0:45

better. Everything sucks, but it's just

0:47

sucks slightly less when it's sunny.

0:49

>> That is true, and things are blooming. I

0:51

can't complain.

0:52

>> Yeah, absolutely. So, you are the author

0:54

of Security Chaos Engineering,

0:56

Sustaining Resilience in Software and

0:58

Systems, and I spent the weekend reading

1:00

the book and trying to understand

1:02

where the intersection of chaos

1:04

engineering and security engineering is,

1:07

because like I remember when Chaos

1:08

Monkey was a thing, and I just got to

1:10

imagine all the Netflix people running

1:12

around pulling cables, and the monkey

1:14

was just messing up their stuff.

1:16

And now I'm trying to understand the

1:17

intersection of security engineering

1:19

with chaos engineering. I'm wondering if

1:20

you could help me understand that.

1:22

>> Yes, I think it's better characterized

1:24

by the umbrella of resilience

1:26

engineering, if anything.

1:28

I've actually had the rare and

1:30

delectable pleasure of unplugging cables

1:32

from Fastly's pops, of course. Network

1:35

continued working perfectly. It is a

1:36

thrill, though, I will admit. Part of

1:38

the title with the book is with a little

1:40

behind-the-scenes tea is chaos

1:43

engineering, especially at the time, was

1:45

a big buzzword. The book is certainly

1:47

more than just chaos engineering. That

1:49

is one tool in kind of the resilience

1:50

engineering toolkit.

1:52

When we think about resilience, it's

1:53

absolutely about how do you recover from

1:56

failure of any kind and prepare for

1:58

what's next? That what's next could be a

2:00

threat, but equally it could be a

2:02

business opportunity. It could be

2:04

massive traffic growth for good reasons

2:06

or it's a DDoS. So, security really is a

2:08

subset of

2:10

the both

2:11

surprises, stressors, opportunities, and

2:14

threats in a very broad sense that we

2:16

need to think about.

2:17

>> This might be a dumb question. It could

2:19

be a spicy question, but like why call

2:22

it security chaos engineering? Is it

2:23

because those are fun words? Because

2:25

like resilience engineering didn't

2:27

wouldn't fly off the shelves? Cuz it

2:29

seems very clear that like resilience is

2:31

really what we want, but it's just not a

2:33

sexy term.

2:35

>> I mean, I think this is a classic

2:36

tension always when you're trying to

2:38

publish, you know, a book or frankly a

2:41

movie, you have [laughter] to have some

2:42

sort of catchy name.

2:44

>> No, that's a great point. Chaos would be

2:46

an awesome movie name, but resilience is

2:47

like it's more of an A24 movie.

2:50

>> Exactly. A24 vibe. It's also, you know,

2:53

you know, talked about it at Davos and

2:56

it's in the National Association of

2:57

Corporate Directors book around

2:59

organizational resilience. It is a great

3:01

Latin root word for international

3:03

appeal, but I think chaos people are

3:05

like, "Wait a second. Chaos can be a

3:06

good thing? That doesn't sound right."

3:08

>> Absolutely.

3:10

>> I want to engineer chaos. That'll be

3:12

exciting very exciting.

3:13

>> Exactly.

3:14

>> Now, you have said that security should

3:16

be designed for failure, not for

3:18

prevention. And I think that's a really

3:20

cool way to think about that. Can you

3:22

think of an example where there's a

3:24

perfectly secure system that still

3:26

failed in the real world? Like what's an

3:27

example where Oh, it still happened and

3:29

we couldn't stop it?

3:31

>> I mean, I feel like tons. First, no such

3:34

thing as a perfectly secure system,

3:36

right? I think there's so many esoteric

3:38

failures out there. I think about the

3:40

airline industry has learned many years

3:42

ahead of software about the intricate

3:45

nature of complex systems and all the

3:47

failures that could go wrong. But the

3:48

example I always think about is the fact

3:50

that they designed I forget which

3:51

airplane it was, which model. They

3:53

designed it with safety in mind to

3:55

almost every degree except for the fact

3:58

that in a very bizarre scenario, if you

4:00

somehow exploded the coffee maker

4:03

like boiling it too hot or something, it

4:05

happened to be close enough to like the

4:07

panel with some cables that it could

4:09

cause a critical failure while the plane

4:12

was in air.

4:13

>> Oh my god.

4:14

>> Right. You wouldn't think about that as

4:16

like that is the trigger to like a

4:18

massive failure that means there has to

4:20

be like an emergency landing, but yet

4:22

there they were. So, I think looking at

4:23

real-world systems and all the just

4:25

bizarre ways that they can fall apart, I

4:27

remember back in the days with Twitter,

4:29

there was cyber squirrel where it talked

4:31

about all the power plant failures

4:33

caused by squirrels just doing things

4:35

squirrels do, you know, and how they

4:38

were almost the more threatening

4:40

advanced persistent threat because of

4:42

the damage that they wrought. I think

4:44

there's just so many examples of like

4:45

your best intentions, you know, reality

4:48

is stranger than fiction. You're not

4:50

going to be able to dream up every

4:51

scenario that's possible, so you have to

4:52

prepare for the idea of okay, things

4:54

will go wrong. How do we minimize impact

4:57

and make sure we can evolve to like meet

4:59

the moment?

4:59

>> Yeah. I think it's so important also to

5:01

remember that like

5:03

and maybe this is also a little spicy,

5:05

like it's on you to be responsible for

5:08

your own resilience. And I remember in

5:09

the early days of the cloud when we were

5:11

all trying to get five nines out of

5:12

Azure and five nines out of AWS,

5:14

it's like, okay, Azure went down. I'm

5:16

going to call somebody and yell at them.

5:18

And it's the same but how badly do you

5:19

want your site to be up

5:21

all the time? Do you want it badly

5:22

enough that you're going to put it in

5:23

both Azure and AWS?

5:26

How much do you want a copy of like how

5:28

do you make a plane that doesn't crash?

5:29

Do you fly two planes next to each

5:31

other? And then when one fails, like you

5:33

jump to the other plane? Like it is

5:35

ultimately on us, is it not? And we just

5:37

need to decide how hard to squeeze.

5:39

>> I think there is usually a trade-off if

5:41

you want to really simplify it between

5:43

cost and resilience. Um, to your point,

5:46

you know, ultimately redundancy is

5:48

multiple paths to get to the same goal

5:50

in practice. You need polyglot

5:52

applications and systems. That's pretty

5:54

expensive to pull off. I do think though

5:57

that software has a beautiful luxury we

5:59

sometimes don't leverage.

6:01

Your point about planes, sometimes you

6:03

can run two instances of a service to

6:06

like offload capacity in a way you just

6:08

can't do with physical systems. Same

6:11

thing with simulating failures, too.

6:13

Again, I think that's a very responsible

6:15

thing to do. A lot of complex systems

6:17

wish they could. For instance, you can

6:19

actually instantiate, you know, a real

6:21

kind of clone of the production system.

6:24

You can't replicate a realistic clone of

6:27

New York City to see if there's a

6:29

certain level of trash blocking sewer

6:31

drains, like what level of flooding will

6:32

cause like deaths. Like, you can't

6:35

simulate that with any degree of ethics.

6:38

But, you can in the computer world.

6:39

We're allowed to do it. So, I think

6:41

that's part of my call to action that

6:43

was big in the book is like, okay, how

6:45

do we start taking this more serious

6:47

seriously and really leveraging the

6:48

benefits that the flexibility software

6:51

begets gives us. So, I do think there's

6:54

to your point, yes, some of it is more

6:56

expensive to do, but in another sense,

6:58

maybe we should be allocating more spend

7:00

towards some of that simulation or just

7:03

understanding like the resilience

7:04

contours of our systems better when

7:07

other industries are just looking at us

7:09

shaking us like, why aren't you doing

7:11

this? We wish we could do this.

7:13

>> Yeah. I've been thinking about

7:14

resilience in my own kind of personal IT

7:17

life. I assume you have a home lab and,

7:19

you know, of various sorts. Right now,

7:21

as I talk to you, because I had an

7:23

appointment with you, I am on my backup

7:24

internet. Turns out, I'm looking at my

7:26

UniFi here, my my WAN failed over at at

7:30

4:38 a.m. and I have yet to diagnose it.

7:33

So, I'm on backup internet right now.

7:35

And when I mention that to people, like

7:37

Muggles, like regular people, they're

7:38

like, you have two internet at your

7:40

house?

7:41

I'm like, like, this is my job, bro.

7:43

Like, I'm here. Like, I got I've been

7:45

doing this at this house for 18 years.

7:47

It cost me 45 bucks for Comcast as

7:50

backup internet. I have my fiber, but

7:53

the backup for $45. It only has to fail

7:55

once, like today,

7:57

and that made it worth the money for the

7:59

year because otherwise I would have had

8:00

to cancel on you.

8:02

And that wouldn't happen.

8:03

>> Exactly.

8:03

>> Yeah. So, it's like it was a choice. And

8:05

I feel like there are teams that think

8:07

that they are resilient

8:09

until the thing happens. What's the

8:11

difference between a security team that

8:12

thinks they're resilient and maybe one

8:14

that actually is? Like, I feel like

8:15

there's a lot of false confidence and

8:17

metrics theater that happens.

8:19

>> Metrics theater, that could be an

8:21

episode in itself. The like, oh, what

8:23

percent security coverage do we have?

8:25

Nobody knows what that means. It's a

8:27

meaningless metric. That is a great

8:29

question. I think the giveaway is when

8:32

the security team feels a sense of

8:35

control, probably means that they don't

8:37

have a lot of resilience. Cuz part of

8:39

resilience is embracing the fact that

8:41

like, there will be things well outside

8:43

of your control. So, it's like, how do

8:45

you prepare for that?

8:46

If you were trying to control everything

8:48

and make things as deterministic as

8:50

possible, you've already failed, in my

8:52

view. Cuz the world is not

8:54

deterministic. Humans aren't

8:55

deterministic. We would like computers

8:57

to be deterministic, but they aren't

8:59

fully, at least. It's one of the hardest

9:01

problems in computer science is

9:02

verifying that the software works the

9:04

way that the designer of the program

9:06

intended it to. So, whenever I hear a

9:08

security team say like, well, we have

9:09

full control over the software delivery

9:11

life cycle, like,

9:13

are you sure?

9:15

>> [laughter]

9:16

>> I just saw I just imagined a meme'd

9:18

version of you and that one guy from HBO

9:20

is like, you sure about that?

9:22

You sure about that?

9:23

>> Yep.

9:24

>> Yeah, or even the opposite thing to do.

9:26

>> Yeah, you're right.

9:27

>> Exactly. Probably means you're investing

9:29

in things that make you feel good and

9:30

give you that sense of control and not

9:32

the things that minimize impact.

9:34

>> Yep. Like, it's not It's not directly

9:35

security, but that old joke of like

9:37

backups always succeed, it's restores

9:39

that fail.

9:41

>> Mhm, yes.

9:42

>> So, it makes me think about like chaos

9:43

engineering and infrastructure is about

9:45

pulling wires. And yanking wires is very

9:47

exciting. But I feel like there's a lot

9:49

of pull the wire moments that can happen

9:51

in security, but people are too scared

9:53

to try.

9:54

>> I mean, fear is pervasive in the

9:55

culture, and it's a disservice to the

9:57

industry and the mission for sure. I

9:59

think there are also cases where you

10:02

could be starting with smaller

10:03

experiments or just testing more basic

10:05

hypotheses. My favorite leveraging

10:08

actually fastly's compute, which is kind

10:09

of like a high-performance serverless,

10:11

you can think of it that way. It's just

10:12

a little function that strips out

10:13

cookies just to see like, "Hey, does

10:15

your login site

10:16

work?" Same with like off headers. It's

10:19

just those basic assumptions you hold

10:20

like, "Of course, like we're always

10:22

going to require this for the login

10:24

page." It's like, "Well, is that true?

10:27

Are you sure?"

10:28

And the

10:30

especially when you can like duplicate

10:32

the request, which this prototype did,

10:34

you know, it's pretty low impact to the

10:36

business to run that experiment. There

10:38

are of course things where it's like,

10:39

"Hey, rm -rf like the customer

10:41

database." Yeah, that's going to be a

10:43

pretty poorly designed experiment with

10:45

high consequences. But there's like such

10:47

a range in between that I think it's

10:49

very unfortunate that security

10:50

practitioners are

10:52

too hesitant to try those experiments,

10:54

especially they're very hesitant to

10:55

reach out to their peers across the

10:57

island like platform engineering and be

10:58

like, "Hey, can we conspire on

11:00

developing some of these experiments?"

11:02

Cuz there are a lot of jointly held

11:03

assumptions, too, that aren't always

11:06

poked and prodded.

11:07

>> You mentioned about determinism and how

11:09

computers and software is not as

11:10

deterministic as we would love to think

11:12

that it is. Not just because the

11:14

software pretty much always runs as you

11:16

wrote it, but whether or not your intent

11:18

was well expressed certainly is a

11:20

problem, and then the environment within

11:21

which it runs you can't always count on.

11:24

But I'm finding that people seem to be

11:27

spackling or puttying over their systems

11:29

now with what I'm calling ambiguity

11:31

loops, which are basically using an LLM

11:33

to deal with ambiguity by letting it

11:36

fill the ambiguity with randomness. And

11:38

I'm curious in your business, when now

11:41

people are like running playbooks that

11:43

aren't scripts, they are prose. Like a

11:47

markdown file is not a script, I think

11:48

you would agree.

11:49

>> Yeah.

11:50

>> Uh how do you feel about that? Is there

11:52

a place for LLMs to live in security and

11:55

in resilience and in chaos, or do they

11:57

just increase chaos and and entropy?

12:00

>> It depends. I think

12:03

they're only going to be as good as the

12:04

corpus that went into them for one. And

12:07

so unless you can verify like really

12:09

clean code went into it, it's like,

12:10

well, you know, it can maybe be a good

12:13

basis for actually in some cases chaos

12:15

experiments or, you know, specific

12:18

configurations, integration test, etc.

12:21

What I will say though is to me the more

12:22

important litmus test is is this

12:24

replacing human judgment? And if it is,

12:28

that's probably not a good case for an

12:29

LLM. I am pro human judgment and

12:32

creativity. I am pretty anti like very

12:35

repetitive work, very tedious work where

12:37

you don't need that kind of judgment

12:39

call. LLMs can be very helpful there. I

12:41

think they're also document intelligence

12:43

examples, you know, who loves going

12:45

through your compliance documents and

12:48

pulling out relevant information. Like

12:50

LLMs can shine there and that way you

12:51

can focus more on strategically, like

12:54

are we sustaining resilience? Like what

12:56

are the indicators we should be looking

12:58

at for that? But asking an LLM is our

13:00

system resilient? Probably is not going

13:01

to be a great outcome. I think the

13:03

markdown example is a little interesting

13:05

because I am also pro making security

13:07

more accessible. I think we dress it up

13:09

in a lot of arcane

13:11

key phrases and buzzwords, when really a

13:14

lot of people could benefit the industry

13:15

and contribute. So maybe simplifying how

13:18

they can enter and not requiring

13:19

scripting knowledge could be a good

13:20

thing. Else I see it as a potential foot

13:22

gun. So I feel like I'm a little mixed.

13:25

>> So now I hear you. I I'm with you 100%

13:27

on the like human judgment. Like it

13:28

cannot be overstated. I assume that

13:30

you're speaking to universities and

13:31

early in career people often and you

13:33

give your speeches and stuff. And you

13:35

talk to people like, "What should I

13:36

learn?" It's like, "You should learn how

13:37

to have good taste." How do you learn

13:39

good taste? Well, you just got to get in

13:41

there and get your hands dirty and do

13:42

the thing and start pulling wires and

13:44

figure out the system. Certainly, I

13:46

don't want to outsource things to human

13:48

judgment and toil like keeping a site up

13:52

is toil. SRE is toil, but SREs that are

13:55

really good at their job are good at

13:56

their job because of their judgment. So

13:59

there's going to be this constant

14:00

tension between the idea that you would

14:01

replace an SRE with a markdown file

14:03

makes me very nervous. But an SRE agent

14:05

that could maybe kick the node and keep

14:07

it running while I drive over there that

14:10

has value to me.

14:12

>> Yes, buying capacity and buying time I

14:14

think is a great use case as well to

14:16

your point. How do we help the human

14:19

engage better with the system? Even just

14:21

like the rubber duck problem solving

14:24

that can be a useful thing for the LLM.

14:26

That does require quite a bit of

14:27

expertise already though.

14:28

>> I love that you brought that up. I was

14:30

talking to someone recently and I gave a

14:32

whole talk and I sort of brought up

14:33

rubber duck debugging and like no one

14:35

got it. I'm like, "Am I like am I onk

14:37

suddenly?" Like no one gets. This is not

14:39

a generational thing. Like talking to

14:41

the duck. Yeah, your face is saying the

14:43

same thing. Like like those are

14:45

important moments to you're talking to

14:47

yourself in the mirror except now the

14:48

mirror can talk back and that's really

14:50

cool. I find that to be super helpful.

14:52

Have you used LLMs in that context like

14:54

figure your thoughts out and talk to

14:55

yourself?

14:56

>> Sometimes I use my cats more often for

14:58

that cuz they're they give a specially

15:00

like judgy looks that make you really

15:03

question, you know, what you're throwing

15:04

down. I think LLMs can be helpful

15:06

though. Especially I see a lot of people

15:08

struggle to get buy-in. This is kind of

15:10

getting into corporate type stuff, but

15:12

especially international companies where

15:14

it's like, "Hey, I want to get buy-in on

15:15

this resilience initiative. How is this

15:17

going to resonate across different

15:19

cultural contexts?"

15:20

>> Ooh, that's a good one.

15:21

>> Right? Like is chaos perceived

15:23

negatively in certain nations versus

15:25

others? I can tell you for instance when

15:27

I talk about deception and using that as

15:29

a technique for resilience engineering,

15:31

security engineering, American security

15:33

practitioners, not universally, tend to

15:35

result in like, "Well, we don't want to

15:36

be the bad guys." Now, the EU, they're

15:40

like,

15:41

"Tell me more. Yes, please. Like we want

15:44

the Sutter Fudge here." So, that's kind

15:46

of fascinating, right? And I feel like

15:47

that's an interesting like twist

15:50

>> Mhm.

15:51

>> um that I found really useful with LLMs.

15:53

>> I like that. That that I I didn't think

15:55

about that. Yeah, you're right. It does

15:56

It does see broader than than we do and

15:59

it can challenge your assumptions,

16:00

especially if you tell it challenge my

16:02

assumptions as opposed to telling you

16:03

that you're absolutely right.

16:07

ACM Bytecast is available on Apple

16:09

Podcast, Google Podcast, Podbean,

16:11

[music]

16:11

Spotify, Stitcher, and TuneIn. If you're

16:14

enjoying this episode, please do

16:15

subscribe and leave us a review on your

16:17

favorite platform.

16:20

You probably [music] work with big

16:21

companies like a banks and slower-moving

16:23

things, health care. They're a little

16:25

more conservative. I'm curious, is there

16:27

an example where traditional compliance

16:30

actively makes systems less secure,

16:31

where they think that they're checking

16:33

boxes, but they're actually hurting

16:34

themselves?

16:35

>> Yes, actually a frequent co-conspirator

16:38

of mine, Zulfikar Dykstra, wrote a

16:40

paper, not with me. It's an excellent

16:43

paper about that exact topic. I think

16:45

specifically covers HIPAA and maybe one

16:48

of the others that shows that it being

16:50

more compliant doesn't result in better

16:52

security outcomes.

16:53

I'm very much of the view and I've tried

16:55

to caution regulators as well as like

16:57

well-intentioned regulation in this

16:59

space very quickly calcifies and

17:02

ossifies. Like it's what helps in year

17:06

zero through maybe even

17:08

year three may end up actually eroding

17:11

resilience long-term. Great example,

17:13

I'll keep the person not a very

17:14

innovative CISO, had to explain, I think

17:17

over a few years to his auditors, like

17:20

actually it's a great thing that we

17:21

don't allow SSH access anymore cuz

17:24

that's what attackers love. They love

17:26

when you leave the door open like that.

17:28

But on the little compliance checklist

17:29

for the auditors, they're like, "Okay,

17:31

but it says you're required to have SSH

17:33

access." He's like, "Okay, but the

17:34

security outcome is now better." So,

17:38

that's where sometimes it can actually

17:39

hold companies back

17:41

even big companies who do want to

17:42

innovate by basically tying them to

17:45

their investments and their spend to

17:47

just checking those boxes, which is a

17:49

disservice to their overall mission.

17:51

>> Yeah, I always want to assert

17:52

assumptions. I feel like cuz I work at

17:54

Microsoft in my day job that my

17:56

ignorance is kind of my superpower cuz

17:58

someone will throw me in a new situation

17:59

and I'll see a checkbox like must

18:01

include SSH. I'm like, "But why?" Like

18:03

and they don't they no one knows. Like I

18:04

don't know like 13 years ago someone

18:05

wrote that checklist and now it's a

18:07

thing.

18:08

And then investors see it and compliance

18:10

people see it and that checkbox is the

18:12

thing that stands between you and some

18:14

certificate some badge.

18:16

And that's a problem.

18:17

>> It's a huge problem and actually bring

18:19

up a kind of elegant point. If you look

18:20

at what resilience means across all

18:23

sorts of complex systems but also the

18:24

ones we're talking about here,

18:26

a lot of when a system is stuck in let's

18:29

say it's not elegant but like a a less

18:31

resilient or like unresilient state or

18:33

fragile state is because a lot of the

18:36

processes and practices that they have

18:37

in place, to your point, are from an

18:40

equilibrium that no longer exists.

18:42

Right? The status quo has moved on. The

18:44

practices haven't.

18:46

And so you're just continuing to erode

18:48

resilience as you like stick to this old

18:50

world and have not adapted to the new

18:52

one and the new context. The same with I

18:54

always hear CISO to be like, "Well, once

18:56

we patch the vulnerability or fix it,

18:57

then like it's fine." It's like, "Well,

18:59

if it actually resulted in an outage or

19:02

a breach, it's not actually fine because

19:03

you still haven't addressed the

19:05

underlying impact. You've just patched

19:06

over the one way attackers got in. They

19:09

haven't adapted to that new paradigm and

19:11

that new equilibrium, which is hard.

19:13

It's updating your mental model of the

19:14

system, which is not easy. That's why it

19:16

is so important to have people who will

19:18

be like, "Well, why? Why is that?" You

19:20

know, just poke and prod.

19:22

>> Often CISOs and security teams make

19:24

dashboards cuz they want to roll things

19:26

up, and the bigger the company, the

19:27

bigger the dashboard, and then the CISO

19:29

has to really They can't know the entire

19:31

stack. The stack is now too deep. So,

19:34

what is an example of a misleading

19:35

security metric or dashboard that might

19:38

cause someone to make a mistake relying

19:40

on a dashboard, but it's maybe a

19:41

misleading metric?

19:43

>> So many metrics. Certainly that security

19:45

coverage one or risk coverage. Also, the

19:48

number of vulnerabilities

19:51

discovered. I'm trying to remember who

19:53

it was who talked about this where

19:56

actually when things started to get

19:57

better, it meant that their application

19:59

development teams were servicing more

20:01

security issues, which was a good thing.

20:04

And it meant that there was more in that

20:05

trust, mutual trust between teams, but

20:07

it looked like it was getting worse.

20:08

>> Yeah, see, that's a great example.

20:10

That's the whole thing like, you know,

20:12

like all these bugs and all these

20:13

security issues. That's this good stuff.

20:14

All of that is low-hanging fruit, but

20:16

like

20:17

they'll assume that something bad has

20:19

happened or something has changed, and

20:20

they're going to then

20:22

correlation and causation are not the

20:24

same.

20:24

>> Exactly. I think And also, a lot of the

20:27

security specific metrics don't tell the

20:28

bigger picture.

20:30

Includes I think about, you know, the

20:31

poor platform engineering teams are

20:33

handed a list of a thousand

20:34

vulnerabilities. Turns out a lot of them

20:36

are in components that aren't even

20:37

exposed to the public internet. Should

20:39

they prioritize those? Probably not. And

20:41

meanwhile, in actually there are

20:43

multiple cases of this, so I'll keep

20:45

them all anonymous, but it's surprising

20:47

actually how often this happens.

20:49

Security team will be on them and be

20:50

like, "Fix all of these, even if they're

20:52

not publicly exposed." And then the

20:54

security team actually maintains, you

20:55

know, their creds into their, you know,

20:57

like

20:58

whatever admin system or security system

21:01

that has its hooks into everything is

21:02

like in a text file on their desktop.

21:04

It's like, "Well, what do you think is

21:06

actually the bigger issue here in terms

21:08

of what attackers could leverage? So,

21:10

there's a lot of that kind of attacker

21:11

math, attacker calculus that isn't baked

21:13

in as well. And I think there's also,

21:16

even if we think about business context,

21:20

the metrics that a lot of security teams

21:21

track and even CISOs track aren't the

21:23

ones that the board wants to understand

21:26

>> Mhm.

21:26

>> or other executives need to understand

21:28

either.

21:28

>> Yeah.

21:29

If I go to fastly.com and I click on

21:31

products, you've got all the network

21:33

services, all the things that Fastly's

21:34

known for. There's a whole section on

21:36

security. And there's also, you know,

21:38

you have services and folks that you can

21:40

hire, professional services, and things

21:41

like that. But how should I think about

21:43

what security is my responsibility and

21:46

what is the responsibility of the vendor

21:49

for whom I am paying a lot of money to

21:51

make things secure?

21:53

>> I think it depends on the vendor. I

21:55

think in the case of, let's say, like

21:57

some sort of SaaS application, let's

21:59

take it sales and marketing, so it's

22:00

going to have some of your customer

22:01

data, prospect data. Feels reasonable

22:03

for the most part that, like, the

22:04

encryption and things like that should

22:05

be handled by the vendor, for sure.

22:07

>> Mhm.

22:08

>> There are cases, I'll use like Fastly,

22:09

where we have

22:10

the platform where you can basically

22:12

write code, run code, etc. It's our

22:14

responsibility, and we have done this,

22:16

to layer in like memory safety by

22:18

design. Same with like isolation models,

22:21

ensuring like safe multi-tenancy. It's

22:23

very much our responsibility. Making

22:25

sure that, like, you don't write

22:27

vulnerabilities into your own code, it's

22:29

like, well, that's probably outside of

22:31

our remit. Though, that's where

22:33

sometimes pro serve can come in. Things

22:35

like, "Hey, you have spun up like a

22:37

service on Fastly, and that connects to

22:39

a database that you have wide open

22:40

without any access control list." Not

22:42

really our responsibility cuz we don't

22:44

touch that component, right?

22:45

>> Right.

22:46

>> Um there are ways that we can help with

22:47

middleware that runs on our platform. I

22:49

do think, though, there's a fundamental

22:51

principle

22:53

though in the conversation a lot of

22:54

people miss, which is like, you have to

22:56

own your own dependencies. And so, what

22:58

you adopt, you do have to basically

23:00

think about it like, well, we have to

23:03

assume at some point something will go

23:04

wrong with it whether that's a security

23:06

issue or not. And I do see a lot of the

23:08

like hot potato game happening there.

23:11

>> Yeah, that's exactly why I asked you

23:13

that question because I think that

23:15

people pay a lot of money for a

23:17

platform, a cloud platform because they

23:18

want as they say a throat to choke. So

23:21

like who gets yelled at, right? You

23:23

know, get Kelly on the phone. I want to

23:25

know what's going on over there. But

23:26

then it's of course their thing. Then

23:28

they are bringing in who knows unknown

23:30

node packages from unknown provenance

23:32

and then they haven't thought about

23:33

their entire

23:34

secure supply chain. But at the same

23:36

time, I like your point about a sales

23:38

and marketing CRM or something like

23:40

that. Is it their job as an app to be in

23:42

charge of like AI bot management or DDoS

23:45

or even API security? Like that would be

23:47

an example where a cloud platform could

23:49

secure those endpoints

23:51

and hide that. So I like the separation

23:53

of concerns

23:55

there. But I I wonder if people who are

23:57

putting together their own systems think

23:59

about that. Like we shouldn't be in

24:01

charge of API security. Let Fastly

24:03

secure our endpoints. And then they just

24:05

have a nice clean bright line or is it

24:08

always layered and they would have to do

24:10

that? Basically you have two layers.

24:12

>> I think it really depends on the company

24:14

and their level of resourcing. There are

24:15

some companies that culturally want to

24:18

own

24:19

where things are built more of their

24:20

things and so they'll leverage us more

24:22

to be able to DIY. There's certainly

24:25

others where it's like, well, let's just

24:26

use what Fastly has, right? Especially

24:29

when it comes to security, putting in

24:31

whether it's our WAF or like you said

24:32

the AI bot kind of like insights we're

24:34

able to surface, like have that in front

24:36

of our services, apps, sites, whatever

24:38

it is. It really depends on resourcing.

24:41

There's also the element of

24:43

I'm going to say a newer worlds where I

24:46

have seen so many security leaders

24:47

thrust into a conversation now where

24:49

their CEO, their CMO, their board is

24:52

like, "Hey, what AI bots are actually

24:54

trying to scrape our stuff so we can

24:56

monetize it?" And this is like I've

24:58

never had to think about this before.

25:00

Trying to DIY that is pretty hard and

25:02

hiring that expertise is some of the

25:03

most expensive expertise out there right

25:05

now. Using a tool probably makes sense

25:08

in that case.

25:09

>> Yeah, I think that's a great point. I

25:10

mean, this is the thing. What do we do

25:12

here at the company? We do insurance.

25:14

Okay, then why are we doing AI bot

25:17

management?

25:18

>> Right.

25:18

>> That's not our job, you know? So I

25:20

always think about the business and I I

25:22

think sometimes when we are talking

25:23

about all the things we've been talking

25:25

about on this show, we don't talk about

25:27

like why did we actually make this

25:29

software? We made it to solve X business

25:31

problem. Therefore, what responsibility

25:34

is mine and what can be outsourced by

25:36

someone who actually knows what they're

25:37

doing, whether it be Fastly or Azure or

25:39

AWS. Let somebody who actually cares

25:42

about that do that while I focus on the

25:43

business problem. Honestly, I don't want

25:45

to do the stuff Fastly does. That's why

25:46

Fastly is good at it, you know what I

25:47

mean?

25:48

>> Yes, and it's laying out your own pop

25:50

infrastructure, especially in this day

25:52

and age of RAM prices being what they

25:55

are, especially if you were a small

25:57

business, it's quite unlikely that

25:59

you're going to be able to do that.

26:01

>> Yeah.

26:01

>> Nor should you cuz to your point, it's

26:03

not your core business. I think it's

26:05

I was thinking when you were talking

26:06

about the example of the duplicate or

26:08

secondary internet, part of that is

26:10

because your essentially critical

26:12

function hosting this podcast

26:13

[clears throat] is

26:15

you can record it. However, if you were

26:17

to build your own microphone, I'd be a

26:18

little bit like, is that actually your

26:21

core value add here, you know?

26:24

>> No, it's not a good example.

26:25

>> Right? It's just understanding like what

26:27

matters and what makes you unique as a

26:29

business, for sure.

26:31

>> Yeah, absolutely. This is totally random

26:34

and off topic, but we had a fishing

26:36

thing happen at work yesterday where

26:38

they they send us fishing emails, but

26:40

it's from the red team.

26:42

And I was so proud of myself. I was just

26:43

like, I don't think that's real. And I

26:45

was like, report fishing. And then it

26:47

was like, congratulations.

26:48

You're one of the better people. So,

26:50

know, there's some number of people at

26:51

the company that does that. And I'm

26:53

always impressed that there's a whole

26:55

teams out there

26:57

trying to attack us internally that I've

26:59

never even met. You know what I mean?

27:00

The red hats or the I guess they call

27:02

them blue hats at Microsoft cuz our

27:03

badges are blue.

27:05

Somehow I was just thinking about

27:06

there's people trying to create chaos

27:09

internally at the company and they tried

27:10

to catch me yesterday with a fish. And I

27:13

didn't

27:13

>> However, what I will say that has gone

27:15

wrong in the past and I've spoken

27:17

publicly before it was cool to have this

27:19

take. People were quite angry. I

27:21

remember this was many years ago where I

27:23

said, "Hey, it's maybe not a great

27:24

thing." Especially during COVID this

27:26

happened a lot. To be like

27:28

"Here's your surprise bonus plan." And

27:30

that's the fishing

27:32

like simulation email.

27:35

>> Oh, no. That would be awful.

27:37

>> Right, but that was happening. So,

27:38

that's not

27:38

>> That's just That's kind of punitive.

27:41

>> I agree.

27:41

>> Right. Click here for more money. No,

27:43

this was not that. This was more like,

27:45

"You have mail waiting for you in the

27:46

mail room."

27:48

And I'm like, "We don't have a mail

27:49

room." You know?

27:50

>> There you go. Yeah.

27:52

>> That's fair. You're right. I mean, this

27:53

is the whole like sprinkling USB keys

27:56

around the bank parking lot kind of way

27:58

of doing things. It's like, "Beyoncé's

28:00

new album." Sprinkle, sprinkle, and then

28:02

everyone plugs it in and then owns the

28:03

entire bank.

28:05

>> Yeah. Something like that. I think there

28:07

are a lot of experiments. I think it's

28:08

always keeping in mind again the human

28:10

element that you don't want to sow

28:11

distrust. But again, you can also make

28:14

the experiments collaborative

28:16

>> Mhm.

28:16

>> which that can get really fun because

28:18

I've actually You mentioned SREs. When I

28:20

talk to like real attackers, they're

28:23

generally not scared of security

28:24

engineering teams. They're scared of

28:25

SREs cuz SREs will like obsess.

28:28

Performance There was that one backdoor,

28:29

right? Was it xc utils where it was a

28:31

guy who's like "Oh, performance degraded

28:33

by I think it was less than 1%. What is

28:35

going on?" Discovered the backdoor.

28:38

>> Right.

28:38

That was awesome.

28:40

That was pretty cool.

28:41

>> Right? So, I think security teams need

28:44

to embrace like, "Hey, you may You will

28:46

have good ideas, but like there're going

28:47

to be other very clever people where if

28:49

you say, "Okay, if you got really mad at

28:51

the company, how would you attack us?"

28:54

They're probably going to have some

28:55

interesting ideas that maybe can become

28:57

experiments or clue you into some gaps

29:00

maybe you have in your current security

29:02

investments.

29:03

>> Very cool. This is you've given me a lot

29:04

to think about. I thought that chaos

29:07

engineering and security chaos

29:09

engineering and this kind of resilience

29:10

was kind of a branding exercise, but it

29:12

feels more concrete after having chatted

29:15

with you.

29:15

>> Yes, I mean, again, it was mostly a

29:17

buzzword and it's part of playing the

29:19

game that publishers have to play. I

29:22

will say though for a very long time

29:24

since I was a wee lad, as they say, I've

29:26

been obsessed with chaos theory.

29:28

>> Mhm.

29:28

>> And I do think chaos theory, which is

29:31

quite beautiful in the sense of like

29:33

systems do have an order to them, but

29:34

it's not necessarily predictable. It's

29:36

more like obviously like a fractal or

29:38

dragon curve in many cases, as any

29:41

meteorologist knows well. So, we need to

29:43

focus less on again, do we have control

29:44

over it? Are we able to predict it? The

29:47

quote I love is from Susan Elizabeth

29:49

Howe, who's a geologist, who said, "A

29:52

building doesn't care whether the

29:53

earthquake was predicted or not. It

29:54

either stays up or it doesn't."

29:56

>> That's good.

29:57

>> Right.

29:57

>> That's very good.

29:59

That's very good.

30:00

>> I feel like that's the essence of it.

30:01

>> Yeah.

30:02

I like that one.

30:04

One of my favorites in the in a similar

30:05

vein is a Babylon 5, "The avalanche has

30:08

begun. It's too late for the pebbles to

30:09

vote."

30:11

>> That is also very good. Yes. Yes.

30:13

>> pebble is like, "I don't like this This

30:15

is not a good idea. This is happening.

30:16

Sorry, this is happening. So, buckle

30:18

up."

30:19

>> Yeah. Exactly.

30:20

>> Thank you so much, Kelly Shortridge, for

30:21

chatting with me today.

30:22

>> Thank you for the great questions.

30:24

Appreciate it.

30:25

>> We have been chatting with Kelly

30:26

Shortridge, the chief product officer at

30:28

Fastly. This has been another episode of

30:29

Hanselminutes in association with the

30:31

ACM Bytecast, and we'll see you again

30:34

next week.

30:39

ACM Bytecast is a production [music] of

30:41

the Association for Computing

30:43

Machinery's Practitioner Board. To learn

30:45

more about ACM and its activities, visit

30:48

acm.org. [music]

30:49

For more information about this and

30:51

other episodes, please do visit our

30:53

website at learning.acm.org/bytecast.

30:54

[music]

30:57

That's b y t e c a [music] s t

31:00

learning.acm.org/bytecast.

Interactive Summary

In this episode of Hanselminutes, Scott Hanselman speaks with Kelly Shortridge, Chief Product Officer at Fastly and author of 'Security Chaos Engineering'. They discuss the evolution of resilience engineering, emphasizing a shift from a focus on prevention to designing systems for failure. They explore how security is often misunderstood as a quest for absolute control, whereas true resilience involves preparing for the unpredictable nature of complex systems. The conversation covers the role of LLMs in security, the importance of human judgment, the pitfalls of 'metrics theater' in compliance, and why organizational resilience requires moving beyond outdated operational models.

Suggested questions

4 ready-made prompts