HomeVideos

People Are Mad They're Told to Learn

Now Playing

People Are Mad They're Told to Learn

Transcript

326 segments

0:00

So, I was, you know, surfing on the old

0:02

Twitters and I saw this tweet said wrote

0:05

a practical introduction to SIMD using

0:08

real examples from Ghostie and I

0:09

thought, "Hey, that's pretty cool." But,

0:12

here's the thing is that often my

0:14

projects I don't really feel like need

0:16

SIMD anywhere. In fact, I took a flame

0:19

graph of the game I'm making and only

0:21

0.83% of the entire time I'm actually in

0:24

the update function. I don't really feel

0:26

like I have a huge need, you know, to

0:29

learn the old SIMDs. Oh, whatever. Hey,

0:31

cool post though. But, then the next day

0:34

this got posted. My SIMD post was on

0:36

Hacker News and the general hostility

0:38

towards understanding how computers work

0:39

in owning your own outcomes is really

0:41

depressing. A large amount of people

0:42

seemed to simply have the attitude that

0:44

they hope someone else will pay their

0:45

bills for them. It can't be that bad,

0:48

right? The vast majority of developers

0:50

have zero need for learning SIMD. Why

0:53

mislead them and make them feel like to

0:56

be a real developer you have to know it.

0:58

What?

1:00

So, an article about SIMD is making

1:03

people feel like what am I reading? So,

1:06

obviously at this point I had to go and

1:08

actually learn SIMD because, you know,

1:10

whether you like this or not, I've never

1:12

actually used SIMD in any sort of

1:13

practical use case. I've understood what

1:16

it was. I've called functions that have

1:18

been SIMD themselves. I have just never

1:20

personally hand wrote anything to do

1:22

with it. So, I thought, "Hey, given the

1:24

fact that Hacker News hates it, I

1:26

probably need to give this a shot and

1:28

learn about it." So, today you're going

1:30

to A have to learn about it, but B we

1:32

got some hot takes, okay? So, I got to

1:34

talk about this whole thing because,

1:36

honestly, why are you misleading them,

1:38

Mitchell? Also, just in case anyone's

1:41

wondering, I'm not going to make the

1:42

joke that everybody thinks I'm going to

1:45

make involving SIMD, okay?

1:47

Not going to happen. And I'd also like

1:49

to say thank you to the sponsors. Vibe

1:51

coders, silence. Someone who's worked in

1:54

enterprises speaking. You know, simple

1:56

OAuth sign-in's not that bad to program.

1:58

For real enterprise applications that

2:00

need single sign-on, passkeys, user

2:03

management, roles and permission,

2:04

enterprise readiness, and passwords, and

2:07

recovery,

2:08

uh that's a little bit more difficult.

2:10

And that is where WorkOS comes in.

2:13

WorkOS is trusted by a lot of the top

2:14

companies, OpenAI, Anthropic, Cursor,

2:17

and a whole lot more. But not only that,

2:19

it's free for your first million users.

2:22

So that means for you, that is 999,999

2:26

free seats available. Pretty much no

2:28

matter what back-end you have chosen,

2:30

there is an SDK available for you. So go

2:32

check out WorkOS at workos.com. Links in

2:36

the description. So the article's called

2:38

"Everyone Should Know SIMD". Now,

2:40

obviously this became the immediate

2:42

focal point. The fact that he said

2:43

everyone should know SIMD, people just

2:46

got pissed off. How dare you say

2:49

everyone? It's just like, okay, here's a

2:51

great concept, and honestly it would be

2:54

really good the more people know when

2:56

programming. Right? I don't think anyone

2:58

should ever argue that. Everybody should

3:00

try and attempt to know as much as

3:02

possible so they can make the best

3:03

quality software out there. And so when

3:06

someone says, "Hey, everybody should

3:07

know this", it's like, okay, I I mean,

3:09

maybe I should know this. But instead,

3:11

people were genuinely upset about that.

3:14

Good article, but I wouldn't start off

3:16

with a bold sentence such as SIMD can be

3:17

simple to understand. And there you go.

3:19

That's the second big one. People hate

3:22

the fact that he calls it simple. Ooh,

3:25

just read the article though. "SIMD has

3:27

a reputation for being complex. I've met

3:29

many a good software engineers who

3:30

dismiss it as something too complex to

3:32

learn or a niche optimization meant only

3:34

for the highest performance software.

3:36

Not useful in everyday programming. I

3:38

think that's wrong. SIMD can be simple

3:40

to understand." And there's there's a

3:41

lot of versions of that comment. It was

3:43

all over Twitter as well. Like, how

3:45

could you say that everybody should know

3:47

something? Or how could you say it was

3:48

easy or simple to understand? It's

3:50

complex. And it's funny because by

3:52

saying it that way, his first line was

3:55

in fact kind of made obvious. So, me who

3:59

who's never really understood what SIMD

4:01

was, I figured, "You know what? Why

4:03

don't I learn what SIMD is? Why don't I

4:05

give it a try, come up with a quick

4:06

problem, and see if I can make something

4:09

disproportionately faster, and then I

4:11

can judge. Is SIMD difficult? Is it some

4:14

magical beast? Or really was that all

4:16

that bad at all?" All right, so let's go

4:17

TO THE WHITEBOARD. OH GOSH. OH, GOD

4:20

DARIO. DARIO, stop doing that, Dario.

4:22

Don't do that. Okay, here. Hold on.

4:24

Let's wrong one, wrong one. So, I

4:26

thought, "What could be a good SIMD

4:28

project?" Well, SIMD you want to do some

4:30

action that you need to do over a long

4:32

for loop, but you have to do like the

4:34

same comparison a bunch or some same

4:36

operation over and over again. At least

4:38

that's what I thought in my head. So, I

4:39

thought, "Okay, a good project would

4:41

probably be I have these long log files

4:44

that I want to be able to find the

4:45

individual logs out, and I want to be

4:47

able to go over this whole list, and you

4:49

know, you just generally have to keep

4:51

doing this over and over and over again.

4:53

And it doesn't really matter how long

4:54

this line is. It just matters that yes,

4:57

we just have to do this over and over

4:58

again up until the you know, {slash} n.

5:00

And again, I'm not going to make the

5:02

joke that everybody thinks I'm going to

5:04

make. So, I figured I would just create

5:06

a program that would find each one of

5:08

these new lines and be able to parse out

5:10

this delicious JSON. I dropped the

5:12

actual parsing of JSON, but instead just

5:14

looked for the new lines, cuz I wanted

5:16

to see how much faster could I make the

5:18

new line looking. Next, I actually had

5:20

to learn what SIMD was. In the most

5:21

simple sense, effectively you can

5:24

operate instead of on a singular value,

5:26

say X, you operate on a vector, right?

5:29

So, pretend you have a vector of X Y Z,

5:32

you could operate on all of these at

5:34

once. Let's just say you wanted to

5:35

compare, say, finding a new line. Well,

5:39

what I'd have to do is I'd have to load

5:40

up a vector with three new lines, kind

5:42

of like my example, and then in one

5:44

single CPU operation. Now, this is not a

5:47

single CPU cycle. CPUs are like magic to

5:49

me. I have no idea what's going on

5:51

underneath there. But what I'm just

5:52

going to call it an operation, a single

5:54

operation can compare this vector to

5:57

this new line. And then it can produce

5:59

out for you a new vector that's going to

6:02

say, "Hey, this is uh the new lines. No

6:05

new line in the first position. There

6:06

was a new line in the second position.

6:07

There was no new line in the third

6:09

position." So, I can actually take that

6:10

and go, "Okay, was there a new line? Oh,

6:13

there was." That means I need to find

6:16

out where the new line was within my

6:18

vector, and then I can return that. And

6:21

that's because SIMD literally stands for

6:22

single instruction multiple data. And

6:25

you can see that right here. You can see

6:27

the fact that I have multiple data, and

6:29

I'm doing a single instruction. Are

6:31

these equal? So, I didn't really know

6:33

how to do that, but I just followed

6:34

exactly what was inside the blog post.

6:37

First, I have to create the constants

6:39

and initialize any vector stuff. I have

6:41

to loop over the input one vector chunk

6:43

at a time. I perform the comparison. I

6:46

reduce it down to something that I can

6:47

actually use. And then I just handle the

6:49

remaining bits at the very very end.

6:51

That actually feels, well, pretty

6:53

straightforward. So, this is the naive

6:55

implementation. This is how I would have

6:57

implemented it if I did not use SIMD.

7:00

I'm going to go over the entire data

7:01

starting at some point, and I'm going to

7:04

go and just say, "Hey, are you a new

7:05

line? If you are, then this is the

7:07

spot." Else, I say, "Hey, no new line

7:10

found. The end." This is your very

7:12

classic index of approach. Then, of

7:15

course, I just created something for all

7:17

the new lines and just go over all of

7:19

them. I'm sure there's better way of

7:20

testing the code. This was good enough

7:21

for me. And when I had run the program,

7:23

as you can see here, it would take about

7:25

1.54 seconds. Okay, not that long. And

7:29

this is the SIMD version. Yes, it is a

7:32

lot larger. It's a little bit more

7:33

tedious to do, but when you break it

7:35

down into the five steps, it's actually

7:36

pretty simple. Step one, the broadcast

7:38

part. Broadcast just means I need to

7:39

create that constant. Remember how I

7:41

talked about this new line constant?

7:44

Well, I do it right here. I go, "Okay, I

7:46

need to create effectively a byte array

7:49

that's lanes big." Lanes in this case is

7:51

how many bytes can I fit in my SIMD

7:53

operations? SIMD it turns out to be kind

7:56

of CPU dependent. My CPU can do 256 bits

8:00

of SIMD operations, so I can do

8:03

effectively 16 of them at a time if

8:06

they're byte wide, meaning they're 8

8:07

bits wide. So, this is my little array

8:09

16 8-bit items. Then I just simply

8:12

create 16 new line bytes. And whatever

8:15

value that is, I think it's hex A. Hex

8:17

A, new line. Okay, I'm a genius. Actual

8:20

genius of useless information, of

8:22

course. Number two, instead of

8:23

iterating, say, one byte at a time, I

8:26

iterate lanes byte at a time. Lanes, of

8:28

course, is how big my SIMD operation is.

8:30

In this particular case, it's 16. So,

8:32

I'm doing 16 bytes at a time instead of

8:35

one byte at a time. From there, I

8:37

actually just create my little value

8:39

array. I just slice off 16 bytes from my

8:41

array and go, "Okay, you are now a SIMD

8:44

vector." From there, I just simply see

8:46

are these equal? I take all my new lines

8:48

and compare to all my values in like one

8:51

single operation. Instead of doing that

8:53

little tight for loop, it's just I'm

8:54

doing 16 operations at once. It's kind

8:56

of cool. Extract most significant bits

8:58

just really checks for any values that

8:59

has a one in it. Since all my values

9:01

should equal zero if there's no new

9:02

lines. If there is a new lines, one of

9:04

them will equal a one. It will grab out

9:06

that most significant bit. And by doing

9:08

a cardinality check, if it's not zero,

9:11

if they're not all zeros, then I have at

9:12

least one one, then I need to go there

9:15

and find that one one. Then I have to

9:16

manually go through my little SIMD

9:18

vector and I have to go, "Okay, which

9:19

one of you are one?" And then the first

9:21

one I find, that's where the new line

9:23

is. If I don't find a new line, I just

9:25

add lanes to the end of end. So, I'm

9:27

just incrementing it 16 at a time. So,

9:30

at the very end you have to do this,

9:31

which is kind of funny. It's called the

9:32

scalar tail, which makes a lot of sense

9:34

if you think about it. Let's just say

9:36

you have seven bytes remaining. Well,

9:38

you can't create something that's 16

9:40

long from seven bytes. Okay, so then

9:43

instead I have to just manually walk the

9:45

remaining seven. And by doing it that

9:47

way, if I simply just run this, you can

9:49

see it's like massively faster. And so

9:52

yes, am I going to be using SIMD anytime

9:55

soon? Probably not. I have virtually no

9:58

problems with performance inside the

9:59

game I'm making, but now I kind of know

10:02

the shape I need to have and now I see

10:04

why I'm going to want to use structure

10:06

of arrays versus array of structures,

10:08

because there is a huge possibility I

10:10

can have some mega wins in some future

10:12

performance optimizations. And honestly,

10:14

it kind of feels cool. I I have to say,

10:17

it does feel kind of neat. So, should

10:20

everybody know about SIMD? Yeah, why

10:23

not? It It It took me an hour to go

10:25

really learn everything and program up

10:27

an example and debug it a little bit.

10:29

Not that bad. You can do Chat GPT can

10:31

tell you all about what you need to do.

10:33

It's actually pretty straightforward.

10:35

You can ask infinity questions. At the

10:37

end of the day, which one would you

10:38

rather have? Maybe a little too much

10:41

knowledge, knowledge that you don't

10:42

often use or don't need to use and you

10:44

just simply kind of disregard it, but

10:46

you have it kind of floating in the back

10:48

of your head. Like Dijkstra's shortest

10:49

algorithm just floating in the back of

10:51

my head. Okay, priority queue. Or would

10:53

you rather be this guy? The vast

10:55

majority of developers have zero need

10:57

for learning SIMD. Why mislead them? At

10:59

the end of the day, I'd rather work with

11:01

the person that wants to know more. I'd

11:04

rather be the person that wants to know

11:06

more. And if I never use SIMD and

11:09

knowing this was completely useless

11:11

throughout my entire life, I'm happy I

11:13

know it. The name is did you know that

11:16

Go 1.6 right here, as you can see right

11:18

here, new experimental SIMD arch SIMD

11:20

package? Did you know that Go even Go

11:22

would now have SIMD support? Obviously,

11:25

you know what that means, right? That

11:26

means we can SIMD these nuts on your

11:29

face, Agen.

Interactive Summary

The video explores the topic of SIMD (Single Instruction, Multiple Data) in programming, sparked by controversial online discussions on whether every developer should learn it. The presenter, who initially felt they had no need for SIMD in their game development projects, decides to learn it after witnessing hostile community reactions to a blog post claiming 'Everyone Should Know SIMD.' By implementing a practical example—parsing a file for newlines—the presenter demonstrates that while SIMD is more complex than standard scalar programming, it is achievable and provides significant performance improvements. Ultimately, the presenter argues that despite its limited day-to-day use for many, acquiring the knowledge is a worthwhile pursuit for any software developer.

Suggested questions

3 ready-made prompts