HomeVideos

I was right again

Now Playing

I was right again

Transcript

372 segments

0:04

You see that face? You know what that

0:05

face is? That that is the smug face of

0:08

somebody who was proven correct. And the

0:10

best part, it's right after caught in

0:14

8K. You're even told by your friends

0:16

you're wrong. So, you're probably

0:18

thinking, "Okay, what what are we

0:19

talking about? What are you even trying

0:21

to say? What were you proven right

0:23

about?" Well, my friends, it turns out

0:26

after many hurt feelings and being told

0:28

I was wrong, images might actually be

0:31

better for models than text. I knew it.

0:35

Don't make it do image recognition on

0:37

that, dude.

0:37

>> No, no, you just do this. What are you

0:38

talking about?

0:41

>> Hey,

0:42

>> now this isn't just me saying this, of

0:44

course. This is actually DeepSeek

0:46

releasing a new paper that suggests that

0:49

if you provide images, you could pull

0:52

out 10 text token from a single image

0:55

token with near 100% accuracy. In other

0:58

words, a model's internal representation

1:00

of an image is 10 times efficient as its

1:03

internal representation of text.

1:05

Meaning, not only do you get the correct

1:08

information out, that you also get it

1:10

way more efficient. And I know I know a

1:13

lot of you right now keep running up

1:14

against limits with Fable. Okay, you get

1:16

like one prompt into it and you're like,

1:18

I'm already out. Maybe, just maybe, this

1:20

is the balm you need to apply to those

1:22

Fable wounds. What a what an analogy.

1:25

There's also been a bit of research into

1:27

this, apparently claiming that yes,

1:29

you're actually going to get 59 to 70%

1:32

lower end to end bill by using images.

1:35

And yes, I myself even did a bit of

1:37

research, spent 2 billion tokens for

1:40

this video. Hey, that's dedication. All

1:42

right, so before we get into the

1:43

nitty-gritty details, I would first like

1:45

to say thank you to my sponsor, Hkins.

1:48

>> Dad, could you teach me how to code

1:51

today? Doesn't look like I have time for

1:53

that, son. My database is down again.

1:56

>> Oh, maybe tomorrow.

2:03

>> Not again.

2:07

>> Go long, son.

2:08

Teach, got to go fast.

2:11

>> Hey teach, how do you have time to play

2:13

with your kid? Isn't your database

2:14

always down?

2:15

>> What do you mean? I chose Planet Scale.

2:17

The database for the consumer who wants

2:20

their downtime to be with their kid.

2:25

>> Nice catch, son. Thanks for choosing

2:27

Planet Scale. You're the best dad ever.

2:30

>> No studies have confirmed Planet Scale

2:31

makes you a better father, but it is a

2:33

really good database.

2:37

Sorry, I was just feeling so vindicated.

2:40

I just had to keep on making that face.

2:42

So, this blog post by Shawn Geodec goes

2:44

in and goes into quite a bit of detail

2:46

about why these things can happen and

2:48

even gives an example by George Mandez

2:50

here about other fun ways you can

2:53

compress data and send it off to Open AI

2:55

or any of these model providers. And

2:57

this one involves just simply running

2:59

your audio through FFmpeg and just

3:01

speeding it up by two to 3x. you will

3:04

save massive amounts of tokens and it

3:06

gets the same results. Okay, so back to

3:08

the images. Why would images work better

3:11

than text? So with text, you say exactly

3:13

what you want to say. It can input and

3:15

process each one of those tokens and

3:17

moves on, right? It that seems rather

3:19

simple. Whereas with images, doesn't it

3:20

have to do a bunch of processing? How is

3:22

this not using a bunch more tokens? So

3:24

here's his basic explanation. Optical

3:26

compression is pretty unintuitive for

3:28

many software engineers. Why on earth

3:30

would an image of text be expressible in

3:32

fewer tokens than text itself? In terms

3:34

of raw information density, an image

3:36

obviously contains more information.

3:38

Goes on to kind of say like, hey, hype

3:39

out the word dog. It's way bigger in an

3:42

image, right? The first explanation is

3:43

that text tokens are discreet while

3:45

image tokens are continuous. Each model

3:47

has a finite number of tokens, say

3:49

around 50,000, and each of those tokens

3:51

correspond to embeddings of say a

3:52

thousand floatingoint numbers. The text

3:54

tokens thus occupy a scattering of

3:56

single points in the space of all

3:58

possible embeddings. What he's trying to

4:00

say is that just imagine a floating

4:02

point somewhere between zero to one and

4:04

you have a thousand of them and you

4:05

could have so many different variations,

4:07

right? An infinite number of different

4:09

variations. But with text tokens, they

4:11

actually have a concrete discrete value.

4:14

One text token to a single, you know,

4:16

combination of a thousand floatingoint

4:18

numbers. Thus, there's only 50,000

4:20

possible representations you could have

4:22

in an infinite space. So very, very

4:24

unoccupied space. By contrast, the

4:26

embeddings of an image tokens is any

4:28

sequence of those thousand numbers. In

4:30

other words, it's continuously through

4:31

that. They can have all sorts of

4:33

different variations. Me much much more

4:35

data can fit into those thousand

4:37

numbers. So, an image token can be far

4:39

more expressive than a series of text

4:41

tokens. Another way of looking at the

4:43

same intuition is that text tokens are

4:45

really inefficient way of expressing

4:46

information. This is often obscured by

4:48

the fact that text tokens are reasonably

4:50

efficient in the way of sharing

4:52

information. meaning me and you can send

4:54

messages to each other and that makes

4:55

perfect sense. But for an LLM that's a

4:57

little bit different. So long as the

4:59

sender and receiver both know the list

5:00

of all possible tokens. When you send an

5:02

LLM a stream of tokens, it outputs the

5:05

next one. You're not passing around

5:07

slices of a thousand numbers for each

5:08

token. You're passing a single integer

5:10

that represents a token ID. But inside

5:13

the model, this is expanded to a much

5:15

more inefficient representation.

5:16

Inefficient because it encodes some

5:18

amount of information about the meaning

5:19

and use of the token. So it's not

5:22

surprising that you could do better than

5:23

text tokens. He also makes an argument

5:26

about the brain how you know when you

5:27

look at a word even if the letters are

5:29

slightly mixed up you can still just

5:30

read the word and so that humans

5:33

actually operate more on images is the

5:36

general argument and sets LLMs are

5:38

supposed to be some sort of

5:39

representation of our brain therefore it

5:41

works. Uh I'm not really into brain

5:44

comparisons. Uh, I'm not quite sure if

5:46

we even know how exactly the brain

5:48

works. And so not knowing how something

5:50

works and then comparing it to something

5:51

we also don't really know exactly how it

5:54

works is not necessarily the greatest

5:57

way to uh to make an analogy for me. But

6:00

anyways, I DIGRESS. I DIGRESS THERE for

6:02

a second. To me, this actually just

6:04

makes sense and I think it makes much

6:05

more sense looking at this. If you just

6:07

look at this image for one second, this

6:09

is an error from the compiler in Rust.

6:11

you immediately know what went wrong and

6:14

how to fix the error. You don't have to

6:16

read all of those text tokens to be able

6:19

to understand what's happening. There's

6:21

something kind of magical about images

6:23

because they encode so much more

6:25

information than you realize. And this

6:28

is obviously like pointing a gigantic

6:30

arrow and saying, "Hey, here's the thing

6:32

you should look at." And then this right

6:34

here is also another gigantic arrow

6:36

being like, "This is the solution." So,

6:38

one can imagine that the image itself is

6:39

actually encoding significantly more

6:41

information, thus giving these models

6:43

more context to operate on, thus having

6:46

better results. So, I thought I would

6:48

give it a little bit of a try. I took my

6:50

entire context file for my game that

6:52

I've currently am using. It's not the

6:54

best context file, honestly. I think I

6:56

could do a lot better, but it is a

6:58

context file nonetheless. And I shoved

7:00

it all into a singular image. I'm using

7:02

something called PX Pipe, Pixelpipe in

7:05

which apparently they're using it with

7:07

Fable and this is the one that I was

7:09

saying does 59 to 70% lower endto-end

7:11

bill. They have a whole bunch of

7:13

information about this saying, "Hey,

7:14

it's actually much much better that

7:16

Fable 5 is actually operating just the

7:18

same that they're able to produce it

7:20

with a lot less money going in and out."

7:22

I mean, this obviously sounds pretty

7:24

good, right? Well, I wanted to see does

7:26

it actually make a difference. And what

7:28

did I test it on? Well, I tested it on

7:30

my video game that I'm currently making.

7:32

So, as of right now, we're actually

7:33

having pretty decent uh progress into

7:35

the game. I can't believe I got Lonely

7:37

Tower enhancement. Off the rip. Let's

7:39

go. Oh my gosh, my tower is going to do

7:41

so much uh damage right now. Don't worry

7:43

about this, okay? The polish will come

7:44

later, but this is just a simple tower

7:47

defense game with ammo, with levels,

7:50

with enhancements, with stat tables, all

7:52

the fun stuff that you can throw on

7:54

these and energy and all that. And I

7:56

have a file that effectively just

7:57

basically describes what's happening and

8:00

how you play the game. More than just

8:01

that, I can also play the game in JSON

8:04

mode. So right now it's running. It's

8:06

letting you know it's ready. It's

8:07

actually running the full game

8:09

experience. The only difference is is

8:11

that it's running it in JSON mode. It's

8:13

not running it in ray rendering to the

8:16

screen mode. So if you want to move the

8:18

mouse, you have to send in a mouse

8:20

operation that says, "Okay, this is how

8:21

I'm going to move." And if you do that,

8:23

it just works. It actually moves the

8:25

mouse and flows through the whole

8:26

program so the AI can control it through

8:29

what appears to be something like MCP.

8:31

So I ran two experiments. I effectively

8:33

let the game be played by these models

8:37

five concurrently on my system just

8:40

running non-stop playing the game over

8:42

and over again for a couple hours. Now

8:44

these models are very slow at playing

8:46

the game. So on the context as string we

8:48

had 50 successful games. Of those here's

8:51

basically the stats breakdown. Now, the

8:53

thing you really want to look at is

8:54

time. I think time is really, really

8:56

important here. On average, it took

8:58

about 9 minutes to play the game. Okay,

9:01

that's pretty good, right? Second, you

9:03

can see how much overall damage it did

9:05

starting from the 10 percentile all the

9:07

way up to the 90th percentile. And of

9:09

course, max. And you can see that it did

9:11

a decent amount of damage. It ended up

9:13

letting out uh so many enemies through

9:15

before dying. It also had this many

9:17

shots taken and hit, this much criticals

9:20

being done. So overall, it did pretty

9:22

good, right? I I'm happy that I could

9:24

see this. This right here means that it

9:26

played the game really, really smart,

9:28

meaning that it actually took a tower

9:30

and placed certain cards in a very

9:32

specific order, thus maximizing the

9:34

damage. So a good percentage of the

9:36

time, it actually made it to the point

9:37

of maximizing the damage really, really

9:39

well. Now, if we look at the context as

9:41

an image side of things, I only got 23

9:44

successful games. And the reason being

9:46

is if you look at the time, the time was

9:48

outrageous. Minimally, it was like 20

9:51

minutes on a game. The average time was

9:53

about 20 minutes on a game, and some of

9:55

the upper games were taking 30 minutes

9:57

to play through. It was just

9:59

significantly slower. But if you look at

10:01

the damage, the damage isn't really that

10:04

much different. It's lower on the max

10:06

side and it's pretty much higher all the

10:08

way through. So, it took a lot longer,

10:10

but it did a lot more damage. But the

10:11

weird part is these things that I'm

10:13

trying to take out, which is the amount

10:15

of enemies that reach the end before you

10:17

died. meaning that a lot of enemies

10:19

reached the end and had just higher

10:21

amounts of health, thus doing more

10:23

damage, thus killing the uh person much

10:25

much faster. And so it's like, okay,

10:27

what can we really glean from this?

10:28

Well, they have about the same amount of

10:30

hits, so that's not really much

10:32

different. The damage is near similar.

10:34

Everything was about near similar except

10:36

for the time. The time seemed to be

10:38

horrible. And so my personal guess is

10:41

that through the discovery of the server

10:43

and maybe some light processing of the

10:46

image, it was able to piece together,

10:47

yes, I'm playing a tower defense game.

10:49

Yes, I kind of know what I'm doing.

10:51

Therefore, it just went and played and

10:54

it kind of could guess its way through.

10:55

And I don't think it did nearly as well.

10:57

And I think the time indicator is really

10:59

what is kind of telling the truth here.

11:01

every single game just took massively

11:03

longer and I just didn't complete nearly

11:05

as many games because it ended up being

11:07

like four hours of running and a lot of

11:10

tokens. And here are the statistics that

11:12

I actually got out from cursor itself.

11:14

As you can see, the full context as a

11:17

string had uh 324 events, 286,

11:22

a bunch of tokens used, 306 million. The

11:25

total tokens for the image context was

11:27

only 213. Now remember, I played well

11:30

less than half of the game. So again,

11:32

what can I draw from this data? What I

11:34

can draw from the data is that it's

11:35

actually probably much harder and much

11:37

more confusing to actually use this

11:39

strategy. And you would probably want to

11:42

do a bit of experimenting. For me to

11:44

even get to this experiment, I had to

11:46

spend about 1.5 billion tokens in just

11:48

experiments to at least verifiably prove

11:51

that things were going the way that I

11:52

wanted them, that uh actual games were

11:55

being played, that you know, the

11:56

objectives were being had. So, at the

11:58

end of the day, should you just go rush

11:59

out and listen to Twitter and convert

12:01

everything to to images now? Probably

12:04

not. But there actually could be a world

12:06

where this actually makes a lot of

12:07

sense, where a lot of good things could

12:09

happen if you do it. And so, like all

12:11

solutions that you hear on the internet

12:13

and all of the Twitter excitement, it

12:15

always comes down to the classic senior

12:18

engineer response. It depends. And you

12:20

probably need to do the homework

12:22

yourself. So, please don't just listen

12:24

to every last cool thing that's thrown

12:26

across the internet. Instead, try

12:29

things, experiment, look for

12:30

optimizations, because you never know,

12:32

maybe there's something super duper

12:34

cool. And also, just to be real, like a

12:37

lot of these technologies, they're

12:38

designed around text. Their primary

12:39

input is text and all that. And so,

12:41

maybe there's an entire world where

12:42

caching and all this other stuff just

12:44

isn't optimized yet for images. So, even

12:46

if you are getting a win, maybe a much

12:48

larger win could be coming down the

12:50

pipe. If only we just wait a little bit

12:53

and this technology is used more. So,

12:55

you know, it's really really hard to

12:56

compare these things because it is a

12:58

very kind of almost apple to oranges

13:00

comparison. All right. Hey, the name is

13:03

the primogen.

Interactive Summary

The video discusses the emerging theory that feeding images to AI models might be more efficient and performant than using raw text tokens. The author conducts an experiment comparing performance and costs in a tower defense game controlled by an LLM, finding that while image-based context can be more token-efficient in theory, it may significantly increase processing time and latency in practical applications. Ultimately, the author concludes that such strategies are highly use-case dependent and require individual experimentation rather than blind adoption based on internet hype.

Suggested questions

3 ready-made prompts