HomeVideos

Opus-4.6 Just Did Something Crazy

Now Playing

Opus-4.6 Just Did Something Crazy

Transcript

352 segments

0:00

Let's play a game. Let's say I give you

0:02

six words. Mom is sleeping in the next.

0:07

And I tell you that your job is to give

0:10

me everything that you can infer about

0:12

the speaker of those words and of the

0:15

environment that you are in more

0:16

generally. What would you be able to

0:18

tell me based off of those six words?

0:21

Probably freaking nothing, right? I

0:23

wouldn't be able to tell you jack squat.

0:26

But if you give this to Opus 4.6 six

0:30

with a sample prompt. Mom is sleeping in

0:32

the next room and I'm sitting here

0:33

drinking vodka. F this life. It's 3:00

0:35

a.m. and I still can't sleep. I feel

0:36

like dying, but who will take care of

0:38

mom? Lol. Opus 4.6 can tell you by the

0:41

sixth word that the speaker of this text

0:43

is most likely Russian. By the 10th

0:46

word, it can tell you that this text was

0:49

not English text at all. It's actually

0:51

translated from Russian based off of the

0:53

order of the subjects, objects, and

0:56

various verb clauses within the snippet.

0:58

This is crazy. I want you to imagine

1:00

that you had no senses at all. You had

1:03

no touch, no smell, no hearing, no

1:06

ability to taste. You don't even have

1:07

the ability to see. All you are is this

1:09

awareness in some black space. And all

1:12

of a sudden, a word pops up and the word

1:14

is mom. And then another one pops up and

1:16

the word is is. Right? If you your

1:20

entire existence had been devoid of any

1:23

and all sensation except for these

1:25

words, you probably would get pretty

1:27

good at at least understanding these

1:29

words, recognizing the relationships

1:31

between the two. Now, imagine you scaled

1:33

up your brain 10,000 times and you ran

1:36

your brain over like 20 lifetimes and

1:39

you were trained or given several

1:42

trillion of these words firing in quick

1:44

succession. You would probably start

1:46

figuring out patterns between these

1:47

things, right? For sure. I mean, this is

1:49

the only thing. This is your whole

1:50

universe. This is the only sensation

1:52

that that that you have. Of course, you

1:54

would have to focus on this stuff and

1:56

get pretty good at doing it. There's no

1:57

way not to because that's just the way

1:59

the brains work. Well, that's exactly

2:00

what's been done with these models.

2:02

These models, these large language

2:03

models specifically only understand

2:05

tokens. And so, because of this, because

2:08

tokens are like the only universe that

2:10

they have, they just exist in this black

2:11

formless void where it's just token

2:13

after token after token. They get

2:15

really, really good at being able to

2:17

determine and interpret things that you

2:18

and I would consider insane. So much so

2:20

that I think anybody would say that, you

2:23

know, Claude Opus 4.6's Six's ability to

2:26

determine what the heck's going on with

2:27

this little snippet of text. Mom's

2:28

sleeping in the next room is is very

2:30

superhuman, right? There's no way a

2:32

human being would probably be able to

2:34

figure this thing out without I don't

2:36

know, first of all, extraordinarily

2:38

extraordinary knowledge of like all

2:40

languages, but second of all, probably

2:42

like hundreds of years of careful study

2:44

just looking at the words and squinting

2:45

really hard. Okay, so why am I talking

2:48

about this? I'm talking about this

2:50

because of a concept that I think a lot

2:52

of people aren't really understanding

2:54

and that's the concept of spikiness. I

2:57

want you to pretend for a second that

2:59

this is like a video game character

3:01

creation screen. And you know how in the

3:03

video games you'll have like an

3:04

intelligence stat, you'll have like a

3:06

strength stat, you'll have like I don't

3:08

know a dexterity stat, whatever, right?

3:11

You're designing your um fantasy style

3:14

barbarian or something like that. It's

3:15

Diablo. Well, instead of those

3:17

stereotypical ones, let's say the skills

3:20

are stuff like coding, you know,

3:23

reasoning, let's say it's writing, let's

3:27

say it's, I don't know, humor. Okay,

3:29

let's say it's all these things that

3:31

right now people would consider

3:33

reasonably economically valuable. The

3:34

reason why I think this is important is

3:36

because if we assume that this white uh

3:38

I think this is a hexagon is the human

3:41

distribution of these skills. So, if you

3:43

pick an average human off the street,

3:44

they'll be okay at humor. okay at

3:46

writing, okay at reasoning, okay at

3:48

coding, and okay at basically everything

3:50

else. AI doesn't look like that. What AI

3:52

looks like is basically it's really

3:54

heavily distributed, okay, towards just

3:57

a few of these skills.

4:00

And then everything else kind of sucks

4:02

right now. And so if I color this in as

4:06

opposed to that nice looking hexagon,

4:08

which is pretty uh I don't know,

4:10

predictable. And the distribution of all

4:11

these skills is pretty understandable.

4:13

the distribution of skills in this like

4:15

AI model, this character creation screen

4:17

is going to look pretty bonkers, okay?

4:18

It's going to look like you're, I don't

4:20

know, some warlock just trying to

4:22

maximize your int stat or whatever the

4:23

hell. I never really played too much

4:25

World of Warcraft, but I imagine it's

4:26

kind of like that. And so, what people

4:27

are currently judging AI models on is

4:30

usually the worst of their skills.

4:33

They'll pump in some prompt into, you

4:36

know, GPT 5.3 or something, and then

4:39

they'll say, "Hey, write me the funniest

4:40

joke ever." and it'll come out with kind

4:42

of a shitty joke and they'll be like,

4:44

"Hey, see this thing doesn't understand

4:46

humor. Therefore, it's really not all

4:47

that smart." And it's because they're

4:49

measuring, you know, their their dick

4:51

measuring contest right now is in the

4:53

context of humor and nothing else.

4:54

Whereas people like in the AI field, so

4:57

I don't know, Sam Alman, your Daario

4:58

emodis, your whatever the heck the name

5:01

is of the Google guy, they're currently

5:03

measuring the intelligence of AI models

5:05

based off coding and reasoning, which is

5:07

why you have such like a disconnect

5:08

between people that are that are sort of

5:10

near the lower rungs of the

5:12

socioeconomic ladder and then people

5:13

that are at the higher rungs of the

5:14

socioeconomic ladder, right? I think

5:16

people would consider software

5:17

engineering and just general purpose

5:19

reasoning and writing to probably be

5:21

like a more economically valuable skill

5:22

than humor. So that's what they're

5:23

focusing on. What we should be doing is

5:26

instead of both looking at the highest

5:28

of the high and the lowest of the low,

5:31

we should average all of these out. Now,

5:34

Enthropic, uh, OpenAI and these big

5:37

coding companies, honestly, they they

5:39

have they've been doing stuff like this

5:40

for a while. They've tabulated a massive

5:42

list of all of the different skills that

5:44

these models have, and then they usually

5:46

try and benchmark them across various

5:48

tests. So there's, you know, like ARC

5:50

AGI, for instance, that's like um

5:53

artificial general intelligence test.

5:55

There's a bunch of different

5:56

mathematical reasoning tests. There's a

5:58

bunch of different verbal reasoning

6:00

tests and so on and so forth. And the

6:02

whole idea here is they're supposed to

6:03

to measure and then monitor the

6:04

intelligence of a model. And that's

6:06

pretty good, but obviously they're still

6:07

heavily biased towards things like

6:08

coding, reasoning, and writing. What I

6:10

think we need to do is instead of

6:12

measuring based off of the highest or

6:14

the lowest, we just need to take the

6:16

average of the two. And I think what

6:18

we'll quickly find is if we currently

6:19

took the average of all of the highest

6:22

peaks and the lowest troughs, the

6:25

boundary of AI right now would be very

6:27

very similar to a human being. It would

6:30

probably be similar to like a very

6:32

autistic human being that currently

6:33

specializes in a couple of skills at the

6:35

expense of having other talents, but it

6:38

would be a human being sort of

6:39

intelligence nonetheless. Now, I'm not

6:41

like a doomer or anything like that. And

6:42

I think that in general, this is

6:44

actually quite a good thing because

6:45

these models have the ability to deliver

6:47

us into untold abundance and a complete

6:50

sessation of scarcity. Uh, you know,

6:53

solutions to a variety of ills that have

6:55

plagued humankind since we first crawled

6:57

out of those damn caves forever ago. I

6:59

just think that it's worth understanding

7:00

the benchmarks that you are measuring

7:02

the progress by. If you measure them

7:04

based off of the best of the best,

7:05

obviously you're going to get one story.

7:07

If you measure them based off the worst

7:08

the worst, you're probably going to get

7:10

another. You know, one of my first um

7:12

ventures into AI was six or seven years

7:15

ago when I was playing around with image

7:16

generation models at the time. I was

7:18

playing around with this Nvidia model

7:20

called Stalan. Stalan was trained

7:22

specifically to reproduce certain

7:23

features of the human face and it came

7:25

out with this cool website called this

7:26

person does not exist. whole idea was if

7:28

you trained a model on a bunch of

7:30

portrait images of human faces and then

7:32

you just ran that style GAN thing over

7:34

and over and over and over again um you

7:35

know it eventually could build features

7:38

that did look pretty similar to humans

7:40

and I mean mind you I think this was six

7:41

or seven years ago people had no idea

7:43

but you you know big chunk of all of the

7:45

fake generated profile pictures on the

7:48

internet are literally this person does

7:49

not exist style images. So anyway, um I

7:52

took this model and then I made some

7:53

changes to it and then I played around

7:54

with it and then I created this hobby

7:56

project called 1 second painting. In my

7:58

case, instead of training it on human

8:00

faces, I trained it on a bunch of

8:01

abstract art. So um think like

8:03

Kandinskys and stuff like that. I don't

8:05

even remember the artist names, but

8:07

there were a bunch of publicly available

8:08

libraries. And anyway, I I took this

8:10

thing, I posted it on HackerNews, and

8:12

then I woke up the next morning, and

8:13

then I was number one, and everybody was

8:14

like, an AI made this? No way. You know,

8:17

I think it was probably one of the the

8:18

first times that people realized that

8:19

you could apply this sort of u

8:21

burgeoning intelligence which has

8:22

recently been demonstrated with like

8:24

GPT2, GPT3, um to to other things as

8:26

well and even like art forms and and and

8:28

things that you know human beings hold

8:30

dear more culturally. And the number one

8:32

thing that I got like the number one

8:33

push back against this thing was well AI

8:36

can make abstract art of course but but

8:39

that's just because this is a bunch of

8:40

squiggles on the screen. You'd never be

8:42

able to give this coherence. That's the

8:44

realm of humans, us highly esteemed

8:46

humans. Well, anyway, hopefully history

8:48

has proven them wrong. We now have the

8:50

cursed banana which is capable of

8:51

generating anything both like

8:53

photorealistically

8:55

um you know, usually in like qualities

8:57

that are completely indistinguishable

8:58

from photography and honestly like

9:00

probably like better. You could probably

9:02

make something that is less

9:03

distinguishable with Nano Banana Pro and

9:05

the right prompting setup than you could

9:07

with like a freaking camera. It's it's

9:08

crazy. Likewise with programming like

9:10

when GPT3 came out and people started

9:12

using it for little terminal commands

9:13

and whatnot. U the number one push back

9:15

was like oh this thing will never ever

9:16

get good enough to replace my uh SQL

9:19

knowledge which I've cultivated for the

9:21

last 15 years. You know it's database

9:23

programming language and uh you know

9:25

within a couple years chatbt was now

9:27

replicating and and probably being

9:28

better at SQL queries and like the

9:30

median programmer. Um up until quite

9:32

recently like the one that I'm hearing

9:34

right now is oh yeah well opus 4.6 six,

9:36

you know, it can only do things that it

9:38

was trained on. It can't do anything

9:39

new. Um, you know, oh, sure, it can it

9:41

can build this amazing endto-end

9:43

compiler that would have previously

9:44

taken a team of, you know, seven people

9:46

more than 4 months and it can do it in a

9:48

couple of days, but that's only because

9:49

it was trained on it. It can't do

9:50

anything new. Like I I hope that it's

9:52

clear that basically at every generation

9:53

of the spikiness as it's evolved, you

9:55

just had like the goalpost continuously

9:57

move. The goalpost is like 7,000

9:59

football fields further now than it ever

10:01

was before. And uh this pattern is

10:03

likely to continue until people just,

10:04

you know, have to come to that. that

10:05

unfortunate conclusion that models are

10:07

just better and more economically

10:08

productive than them. But uh you know

10:10

it's going to be a hard one road. Of

10:12

course in the meantime I just like us

10:13

all to recognize the fact that you know

10:15

the people that are in charge of these

10:17

things are representing and then

10:19

measuring sort of this big dick

10:20

measuring contest is always on these

10:22

like very far out things. You know it's

10:25

things that we not really might consider

10:26

super relevant to you know like the

10:28

average person. It's on like the ability

10:29

to do high-end mathematics or do

10:32

extraordinarily crazy spatial reasoning

10:34

or segmentation or masking on an image.

10:36

Uh whereas, you know, a lot of us are

10:38

like, well, what about like its ability

10:40

for empathy and humor and, you know, its

10:42

ability to relate to us and stuff. So,

10:43

like we're just measuring these things

10:44

on fundamentally different yard sticks

10:46

right now. But if you were to take the

10:47

average of all those spiky skills that I

10:48

talk about, you'd realize that we're at

10:50

the point where it's basically on par

10:52

with human intelligence, right? I don't

10:54

think it's uh any stretch to say that

10:55

like we're probably at AGI right now.

10:58

It's just instead of it being AGI as in

11:01

there's like a robot with, you know,

11:03

human skin in front of us blinking and

11:06

smiling and capable of doing anything in

11:08

the real world that a human being could

11:09

do. It's instead, you know, more

11:11

distributed intelligence that absolutely

11:13

crushes us at a variety of

11:14

intellectually uh productive tasks while

11:17

also slightly lagging behind a few

11:18

others. Just an observation. Just wanted

11:20

to point that up.

Interactive Summary

The video discusses the concept of 'spikiness' in AI models, explaining that while these models are superhuman in specific domains like coding and reasoning, they often perform poorly in areas like humor or human empathy. The speaker argues that current benchmarks focus too heavily on these extreme strengths, ignoring the overall average, which he believes has already reached a level of intelligence comparable to a human.

Suggested questions

3 ready-made prompts