HomeVideos

Notational Intelligence, Linus Lee | Compile 26

Now Playing

Notational Intelligence, Linus Lee | Compile 26

Transcript

386 segments

0:00

I want to begin with an observation.

0:03

The observation is that in this venue, there's probably in the order of

0:08

a thousand, maybe a couple thousand computers.

0:10

There's some in your pockets.

0:11

You brought some with you. There's computers running various systems

0:14

in the

0:15

buildings. There are a few thousand computers.

0:17

But if you think about instances of notations of symbols, graphics, signage,

0:21

maps, things that people have written down,

0:24

notes in your notebook.

0:25

There are probably at least 100 times, if not a thousand times more instances

0:29

of notations of ways that you write things down to communicate,

0:32

to remember

0:32

things, to work on them in this venue.

0:35

And my observation

0:39

is that even though we may not be conscious of it at all times,

0:45

the impact of notations,

0:48

how we write things down to communicate and represent things, is potentially

0:52

greater than the impact of machines that we build to augment our

0:57

intelligence. They're both ways to make us smarter,

1:00

make us more effective at

1:01

what we want to do, but notations possibly has a greater

1:04

impact, even though we

1:05

may not be conscious of it. And I call this idea notational

1:08

intelligence.

1:09

And for the rest of our time together today, I want to talk through first what

1:12

makes good notations, why do some ways of writing things down

1:16

feel more

1:16

satisfying and more useful,

1:18

and dive into some attributes of good notations.

1:20

And then I want to think through how we might be able to use some of their

1:23

modern tools of deep learning to think about new kinds of notations we can

1:26

create. A great example of what I call notational

1:29

intelligence is when you do

1:30

long addition. If I ask you to add three numbers

1:34

or four-digit numbers, let's

1:35

say 217 and 868.

1:40

Maybe some of you can do it in your head.

1:42

I can't really do it that fast in your head.

1:43

But then you do this.

1:45

You can move some numbers around. I hope I'm not making a mistake.

1:48

That's embarrassing.

1:52

It's not this. It's this. Okay. And in this calculation,

1:57

I didn't really do the calculation purely in my head.

2:00

It's not just that I became more intelligent in my brain.

2:02

It's also not that the chalkboard is a machine or a calculator.

2:05

There's nothing behind this. There's something in the way

2:09

that we represent

2:10

information that allows us, the combined system of my brain

2:13

and the things that

2:14

I'm writing down, the shapes that I'm putting on the board,

2:16

to be smarter.

2:17

And so that's what I call notational intelligence.

2:19

And there are different kinds of notation.

2:21

I want to just walk through a few desirable properties that I

2:26

think make certain kinds of notation useful and satisfying.

2:28

And the first is abstraction.

2:34

Abstraction is the idea that you can write down one symbol

2:38

or one kind of

2:39

shape, and it can stand for a whole set of things

2:41

that have similar properties.

2:42

So a good example of this is in algebra, you might say, y = mx +

2:47

b,

2:48

and b is not just a single object. b stands for a whole set of things.

2:51

Maybe it's a set of integers, maybe it's a set of reals.

2:54

All of these variables are individual symbols

2:56

that we can manipulate, so I can

2:57

move this term to the left, and really I'm manipulating a whole family

3:01

of

3:01

relationships between numbers or family of ideas.

3:04

And that is really powerful because you can start to manipulate larger objects,

3:08

potentially entire families of equations or relationships by just moving

3:11

symbols around on the board. The second and third properties

3:15

that make great

3:16

notation

3:17

are suggestiveness

3:23

and what our favorite mathematician, Terry Tao, calls natural transformations.

3:32

And they're related kind of abstract ideas.

3:34

So I'll try to first state it, and then I'll give you an example.

3:38

Suggestiveness is the idea that if two things when you write them down

3:40

look the

3:41

same, then they probably share some property

3:43

that is implied by the similarity

3:45

of the shapes. And natural transformation

3:47

is a related idea that when you could

3:49

apply some operation to the shapes in the space of writing graphical space,

3:54

that they have some corresponding meaning in what you're doing with the

3:57

underlying ideas. So what does that mean?

3:59

A good example of this comes from two different ways that early inventors of

4:04

calculus wrote down the idea of derivatives,

4:07

Leibniz and Newton.

4:09

Leibniz was coming from a more mathematical,

4:11

pure mathematical background.

4:12

He wrote down derivatives as such.

4:14

You would do like a dy over dx, and this is a ratio of differentials.

4:19

It's the ratio of a tiny movement in x to a tiny movement in y, right?

4:22

Remember this from high school.

4:23

And you can manipulate this in a few different ways.

4:26

You can maybe multiply it by other ratio of differentials.

4:30

You can take squares of things and although there

4:34

are rigorous ways you can

4:35

manipulate this, you can also just kind of eyeball certain

4:37

things, and you can

4:39

operate mechanically on the symbols in a way that means something in the space

4:43

of ideas. Whereas Newton, coming from a physics background,

4:46

would say something

4:47

like, "Oh, I only really care about derivatives over

4:49

time, about how fast

4:50

things are moving and accelerating." So he'd say something

4:52

like, "There's

4:53

position, and maybe the first derivative

4:54

is like a dot over the value, and the

4:56

second derivative of acceleration is like two dots over the value."

4:59

And this is

4:59

from an abstraction perspective, it's really nice.

5:02

It's a lot simpler, but you can't really operate on these

5:06

things in a way that

5:06

makes any sense. There's not that much mathematical implied by the

5:10

shape of the

5:11

notations. And so in practice, we find, of course, that this is the more common

5:14

way to work with derivatives in mathematic literature because that Leibniz

5:17

notation is both more suggestive and more amenable to natural

5:20

transformations.

5:22

And then finally, the fourth family of properties

5:24

that we really care about for

5:25

good notations is that they take advantage that they

5:28

are graphical.

5:34

To be graphical means to take advantage of the implicit biases that we have in

5:38

our visual perception system to make sure that whenever we operate on things in

5:42

these two dimensions, that they are kind of naturally intuitive because of

5:45

the

5:45

way that our brain works and our perception system works.

5:48

One way that this is the case is that notations that we work with here

5:51

and in

5:52

the venue and in the signages are all flat.

5:54

They could be three-dimensional.

5:55

You could imagine ways of writing things down involve CAD models and videos and

5:59

other kinds of dimensions, but flat notations

6:01

are useful because we can move

6:02

them around. They're mobile. You can scale them up

6:05

and down really easily.

6:06

So you can take notations that describe really complex,

6:09

really large objects

6:10

and really small objects like bacteria, and you can then bring to bear the same

6:14

tools of geometry, like arrows and shapes and measurement and ratios,

6:18

onto all

6:18

of these notations, different kinds of maps and graphs

6:21

and plots, because they

6:22

ultimately boil down to 2D shapes that are flat, that you can move around and

6:26

you can copy and you can print out.

6:27

A really great example of graphical notation is the coordinate plane.

6:31

So we could take something like Y=MX+B, and we could just say, okay, we'll draw

6:36

a Y-axis here, X-axis here, and maybe this is the line that we want to

6:39

represent with this equation. And this is useful for manipulating

6:44

numbers and algebraic terms, but this is really what gets us to understand the

6:48

relationships between numbers, right?

6:49

If I move the value of B up and down, maybe it means that the line moves up and

6:54

down because that's moving the Y-intercept, or maybe the slope is M,

6:57

and so it

6:57

represents something physical.

6:58

And so the coordinate plane is a really, really useful notation for us because

7:03

it allows us to take advantage of the implicit kind of visual understanding

7:06

that we have as humans, kind of just moving around our space.

7:10

Another

7:11

example of really great visual notation that is one of my favorite examples

7:15

because of how surprising it is, is the arrow. We have arrows here.

7:20

I think pretty much every speaker that we've had today has drawn an arrow at

7:23

some point, pointing to something or gesturing at a direction of movement.

7:27

So it's really kind of we're a fish in the water.

7:30

We're used to the idea of arrows.

7:31

How old do you think this arrow symbol is in recorded history?

7:35

What do you think is the earliest example of the arrow symbol we can find?

7:40

It turns out that the earliest example of arrows that we can find is about 300

7:45

years old. It's from 1737.

7:48

That's pretty wild, right? Before that, if you wanted to draw something and you

7:52

want to point at something, you had to draw a hand pointing at the

7:56

thing,

7:57

which is pretty crazy. And so even something as basic as a

8:00

coordinate plane or

8:01

even the arrow symbol, somebody had to think of and invent,

8:04

and that invention

8:05

allows us to work with these other kinds of notations much, much more easily.

8:10

Another great kind of notation that we are all used to as programmers and

8:14

practitioners in computer science is programming languages.

8:17

Programming languages describe a totally new family of ideas.

8:21

The idea of programs didn't really exist when a lot of this work was done.

8:25

And yet, programming languages are a kind of notation, a thing

8:28

that you write

8:28

down on paper, on a computer screen, to talk about a family of concepts that

8:32

are programs, that are these kind of complex dynamic objects

8:35

that allow

8:36

computers to be turned into useful machines.

8:38

And programming languages have all these elements. They use abstraction.

8:42

They allow for abstraction in that you have variables and functions that

8:46

abstract over important concepts.

8:48

They are suggestive and immutable to natural transformation,

8:51

in that if you're

8:51

refactoring programs, you'll often take a term, take a function,

8:55

replace the

8:55

term with a function. You're moving values around,

8:57

and you're performing

8:58

actions in the space of written programs, in the space of symbols that have

9:03

corresponding meaning in the space of actual programs, this abstract object

9:07

that exists inside a computer, right?

9:09

And then finally, they're maybe more graphical than you

9:11

might initially think,

9:12

because when you indent code, when you syntax highlight code, you

9:15

are using the

9:16

implicit biases that we have as humans of how we perceive color and how we

9:20

perceive space to denote useful ideas in the space of programs like variable

9:24

scope. So these are how we get good notations.

9:29

So now

9:31

we're in 2026,

9:32

we don't have to just invent arrows anymore.

9:35

We can bring the more powerful tools that we have at our disposal, like deep

9:39

learning,

9:41

to try to think about how we might kind of

9:43

in more intentional event, new kinds of abstractions for new kinds

9:48

of ideas. When we build these models,

9:52

as some of the folks talked about previously,

9:56

we have models learning these new kinds of weird ideas.

9:59

When you look inside a model, maybe they learn features that

10:01

are not really

10:02

interpretable to us.

10:04

Maybe if you look inside a model and apply interpretability techniques,

10:07

you get

10:08

things like LoRA layers and strange new features

10:10

that we don't really have

10:11

words for. But if we could somehow figure out how to

10:14

mechanize the process of

10:15

getting good notation, maybe we can mass produce new kinds of

10:19

ways to write

10:19

down ideas that we don't yet have words for or yet have notations for.

10:23

And so how do we invent new notations with deep learning?

10:25

There's actually some prior art in this space.

10:26

In the field of cognitive science, a lot of people use deep learning models,

10:30

things like ConvNets and ResNet, as kind of laboratory models for human

10:34

perception to try to imagine what the evolution of written language may

10:39

have looked like. So they use neural networks as a proxy for

10:43

how humans see the

10:44

world, and then try to use that to imagine how we may have evolved

10:47

from things

10:48

like pictograms to modern language.

10:49

That's not exactly what we're going to do here.

10:51

We're going to instead try to use the same tools of deep learning to try to

10:55

imagine what totally new notations, unconstrained from natural

10:59

notations that we've learned about so far, might look like, and try to actually

11:04

produce a toy model of inventing a new kind of notation.

11:08

So when you want to invent a new notation, you need two things.

11:11

First thing you need

11:12

is you need to decide on an input domain.

11:18

What do you want to represent with the notation?

11:21

And the second thing you need

11:23

is the constraints that are inherent to the medium

11:27

that you're working in.

11:31

So the constraints of the medium.

11:35

For example, so if I'm trying to write things on a

11:38

chalkboard, the constraints

11:39

of the medium are not just that it's flat, but also that maybe you need high

11:42

contrast. Maybe I can't do shading super well,

11:44

so it needs to be just pure

11:45

black and white.

11:47

Maybe it needs to be made of strokes instead of pixels because I'm writing with

11:50

something that generates lines.

11:52

But if you're using a model that's generating raster images,

11:55

maybe your medium

11:55

actually can allow for pixel images and raster images, and we'll see exactly

12:00

that. And so just to produce a toy model of

12:02

inventing a new notation with deep

12:04

learning, let's say we're going to pick a really

12:06

simple input domain, which is

12:08

just a vector of size N, where only one of the values can be

12:12

one and the rest has to be zero. So imagine maybe it's zero

12:18

1000,

12:19

and this is a vector of size five where exactly one value is one, right?

12:22

So we call that a one-hot vector.

12:24

And let's say that's vector N, and we have N values.

12:27

Another way to think about this object is it's like an alphabet.

12:30

The English alphabet has 26 letters. Each letter can only be one letter.

12:34

A letter can't be A and B. And so we're going to try to train a model

12:39

to learn

12:39

a new alphabet that is N letters big.

12:43

So we have this input that's just a vector.

12:45

We put this first into an image generator model G, and this is going to

12:50

generate some kind of an image. In our case, this is a ResNet-style

12:53

convolutional model in my little toy model, and so maybe it generates a 32 by

12:57

32 image.

13:00

And then we're going to have a different model try to read what this model

13:04

wrote down to represent this idea. So this is a generator.

13:07

It's writing something down, and then we have a decoder that's trying

13:09

to decode

13:10

this back into what the decoder thinks the generator tried to write down,

13:15

which we call V prime N. And it's the same size, right?

13:18

And so we optimize this whole thing end to end with gradient descent to try to

13:23

make sure, there was another arrow, try to make sure that the generator and

13:26

decoder learn to talk to each other.

13:29

In other words, when we take some vector, put it through the generator to

13:32

produce an image, put it through a decoder to

13:36

try to figure out what it meant. Hopefully, the values V prime N and VN are

13:40

similar. We try to minimize the difference.

13:42

And at this point, we have a handout

13:45

that we'll start handing out. And if we don't

13:48

have enough handouts for everybody, if you don't get a handout,

13:50

you can also go

13:51

to the URL linus.zone/compile,

13:55

and you'll see the same thing. The handout shows you two interesting

13:58

things.

14:00

The first

14:01

page is in the process of training this thing, where we have in

14:06

the case of the experiment that you're looking at there,

14:08

it's a vector of size

14:09

1024. So the model is learning an alphabet that is 1024

14:13

characters or letters big. It's learning 1024 different symbols.

14:17

It has to learn to distinguish 1,000 different concepts.

14:20

It's doing that in a 32 by 32 grayscale pixel space.

14:24

And you'll see that the first page is the outcome.

14:26

So it's what do these letters look like that these models have made up in their

14:30

own space, unconstrained from human perception.

14:34

In one of the pages, it's that.

14:35

In the other page, it's a sequence of learning.

14:36

Well, you see that in the beginning, you're starting with noise, and the model

14:40

learns basic ideas like black and white, and then over time, it learns more and

14:45

more complex shapes, which is pretty interesting to see.

14:47

The last trick that I forgot to mention in building this thing is that if you

14:51

just train this naively, the model's going to learn some really

14:54

basic hacks.

14:56

Maybe the first letter is the first pixel being on,

14:59

second letter is the second

14:59

pixel being on, and so on and so forth.

15:01

We don't really want that because that doesn't really mean anything to

15:02

humans.

15:03

So we have to go back to our graphical implicit biases of the human visual

15:08

perception and say, what kinds of invariance do we want to

15:10

impose in the

15:11

medium? For example, we want symbols to mean the same thing,

15:15

whether they're

15:15

big or small. And so we impose a scale invariant.

15:18

We make sure that when we get a generator image out of the

15:21

generator, maybe we

15:22

make it smaller and make sure that a smaller symbol decodes to the same

15:24

thing.

15:25

We make it rotational invariance.

15:26

If you take a letter that this thing has learned

15:29

and turn it around, it still

15:30

means the same idea. And so we have rotational invariance,

15:33

scale invariance.

15:34

We also vary the color a little bit, so a brighter or a darker version of the

15:37

same symbol roughly means the same idea.

15:38

And so if you train this system, this kind of autoencoder with an image in

15:42

the

15:42

middle, with these invariants, the model, it turns out, learns to encode a

15:46

family of ideas or domain of ideas in its own kind of visual space.

15:49

In essence, it's learned its own graphical

15:52

representation of ideas.

15:54

The reason I've divided this board here in two like this is that on the left

15:58

side is our world. It's the things that we've, by culture, by communities of

16:02

practice, invented to support us, and they've evolved really organically,

16:06

right? And in the right is this totally alien world,

16:09

where we've taken a model

16:10

not just as a way to learn what's already out there, but as a simulator that

16:14

can simulate anything and learn anything.

16:16

They don't have to be learning just the laws of physics that exist in our

16:19

world. They don't have to just learn the

16:20

languages that we already know.

16:22

We can instead use these to try to imagine new kinds of languages, new kinds of

16:26

ways of writing things down, new laws of physics.

16:28

And I think this idea of thinking of a computer and a model as a simulator for

16:31

anything, not just as a simulator for our specific

16:34

world and experience and way

16:35

of being, is really powerful.

16:38

And while we get a lot of value out of simulating things that exist in our

16:43

world, I think it's maybe much more interesting

16:44

to try to use models as a

16:45

laboratory for imagining other kinds of being.

16:48

Other ways of imagining things, other concepts,

16:49

other ways of writing things

16:50

down.

16:51

And this was really what I thought for a long time about, and I would push

16:55

everyone here with all of your new tools and your keyboards to think about it.

16:59

How do we use models that we know and love to not just model the current

17:03

world,

17:03

but to speak of unspeakable ideas and to see what we really can't imagine

17:07

yet

17:07

that might come in the future? Thank you so much, guys,

17:09

and have a wonderful

17:10

conference.

Interactive Summary

The video explores the concept of 'notational intelligence,' arguing that the systems we use to write down and represent information—like symbols, maps, and programming languages—often have a greater impact on our effectiveness than the machines we build. The speaker identifies key properties of effective notation, such as abstraction, suggestiveness, and graphical utility. Furthermore, the video proposes using modern deep learning techniques to move beyond human-invented notations, demonstrating a toy model that generates new, AI-created alphabets constrained by human-like visual perception to help us conceive of entirely new ideas.

Suggested questions

3 ready-made prompts