HomeVideos

OpenAI BROKE the Industry Overnight....

Now Playing

OpenAI BROKE the Industry Overnight....

Transcript

420 segments

0:00

OpenAI just announced their AI chip,

0:02

Jalapeno, and I don't think anyone was

0:04

quite ready for this. Here's the big

0:07

headline that you need to know. This

0:09

chip beats Nvidia on the metrics that

0:13

matter when we're talking about using

0:15

and serving AI. So, this is their first

0:18

generation chip. It's supposed to be

0:19

bad. It's not supposed to be competitive

0:22

with all the other chips on the market.

0:23

That's not how things work. Instead,

0:25

this is a first generation chip. This is

0:27

the first chip that OpenAI has produced

0:29

and it beats Nvidia. It beats Google's

0:31

TPUs and it's not only for OpenAI's

0:34

models. This is tested across many

0:36

different open- source models. This is a

0:38

general chip. This is an industryleading

0:41

chip. How did OpenAI come out with this

0:45

killer industry destroying chip on their

0:48

first attempt? If you can't guess, then

0:51

you haven't been paying attention. This

0:53

hopefully will not come as a surprise

0:55

for anyone. They're saying right here,

0:57

they say it. We used AI to design the

0:59

chip and design the chip so AI could

1:02

program it. AI played a direct role in

1:05

Jalapenño's development, enabling the

1:07

team to move from initial design to tape

1:10

out in 9 months. By the way, this isn't

1:13

just OpenAI saying whatever they feel

1:15

like saying. Semi analysis actually went

1:18

there and tested this chip. We'll take a

1:20

look at the report in a bit. Some of the

1:22

pieces in there are kind of wild, but

1:25

really fast. Couple of things that are

1:26

important to understand. First and

1:28

foremost, this isn't like a GPU that can

1:30

be good at many different things. This

1:32

chip that they're building, the

1:33

Jalapeno, it's an ASIC application

1:35

specific integrated circuit. So, it's

1:36

basically it just does one thing. It

1:38

woke up one day and asked, "What is my

1:40

purpose?" And they told it to you pass.

1:43

But instead of passing butter, all it

1:44

does is it works with large language

1:47

models to do the inference. So the

1:49

outputs of your favorite chatbot, it's

1:50

designed to just do that as fast and as

1:54

efficient as possible. And the metric

1:55

that they're going by is performance per

1:58

watt. So basically, how much useful work

2:00

does a chip get done for every watt of

2:02

electricity that we give it. So really

2:04

what we're measuring is how efficiently

2:06

are we converting energy into AI output.

2:10

And energy is the big bottleneck right

2:12

now. That's the thing that's the hardest

2:13

to get. So how well does the jalapeno

2:16

line up against all the other giants?

2:18

all the other competitors that have been

2:20

working and perfecting their chips for

2:22

years. Well, here's the semi analysis

2:24

blog post. Shout out to those guys and

2:26

gals. Those have just amazing coverage

2:29

of the space. Okay, so remember we're

2:30

looking at the performance per watt.

2:32

That's kind of like what we're trying to

2:34

maximize. That's the most important

2:36

metric. How well does the Jalapeno

2:38

compared to every other chip? Jalapeno

2:40

smokes every other chip. That's semi

2:43

analysis who had a chance to run their

2:45

own test on this chip. So this is not

2:48

hearsay. They're saying jalapeno beats

2:49

blackwell on that metric performance per

2:51

watt across almost all scenarios without

2:54

being tuned for any specific point in

2:55

the curve. You want to do it fast. It

2:58

does it fast and it smokes every other

3:00

chip. You want to have a high

3:01

throughput. You want to have a lot of

3:02

stuff going through it. It still just

3:04

knocks everybody out. At low concurrency

3:06

scenarios, Jalapeno demonstrates

3:07

remarkable interactivity hitting over

3:09

700 tokens per second per user at

3:12

concurrency one on the DeepSseek R1

3:14

model. Now, keep this in mind, the Deep

3:17

Seek R1 model, because this chip did not

3:21

have the code needed to run this model

3:23

efficiently. Just put a pin in that

3:25

because you're going to want to see the

3:27

other side of the coin here in a second.

3:29

Now, I'll leave a link to semi analysis.

3:31

I encourage everybody to check it out

3:32

because there's tons of context here.

3:34

So, I'm not just cherry-picking the good

3:36

or the bad. There's a lot of stuff here

3:37

that we are not going to go into. They

3:40

have some caveats in terms of what tests

3:42

were were run. So they verified their

3:45

tests that were run in person in the

3:47

lab. They did not run the full suite,

3:49

nor have they seen their preferred suite

3:51

for comparing chip performance. So

3:53

there's still more tests to be done. But

3:55

this is where the other shoe drops. This

3:57

is kind of the other side of the coin.

3:59

And this is one of the more wildest

4:01

little nuggets of wisdom or things that

4:04

Simeon House has found at OpenAI. So

4:06

really fast, let's say you have this AI

4:08

chip. So this is our AI chip and here we

4:12

have our AI model. Does the AI model

4:14

just kind of go and run on the chip? No.

4:17

The chip can't do anything without

4:19

something telling it what exactly to do.

4:22

And that thing personified in this case

4:24

by the kernel. Kernel is basically just

4:27

code. It's it's written code.

4:29

Specifically, it's a small hyper

4:31

optimized piece of code that tells that

4:33

chip exactly how to do math. how to do

4:36

specific math to run this AI model like

4:38

matrix multiplication or various

4:40

attention mechanisms and good kernels

4:43

are notoriously hard to write. It's a

4:46

specialist work. There's a lot of

4:47

optimizations. So even if you have an A+

4:50

like a really good AI model, you have an

4:51

A+ chip. You have the best hardware. You

4:54

have like a D or an F level kernel C

4:57

minus whatever this could easily you

4:59

know cut the performance in half. So

5:01

when it was time to run the deepseek

5:04

probably the R1 model on this Jalapeno

5:07

chip, OpenAI had no internal

5:09

implementations of MLA kernels. MLA is

5:12

the multi- head latent attention. If

5:15

you're not familiar with it, it doesn't

5:16

really matter. It just DeepS had this

5:18

funky way of doing things that most

5:21

other labs they they just don't do it

5:23

that way. So they just have an unusual

5:25

architecture. Most other labs don't use

5:27

that. OpenAI models don't use that. So

5:29

open never had to have this code this

5:32

this kernel for MLA. So they're trying

5:34

to run a test. They're trying to use

5:36

this DeepSeek model. So this is

5:38

Deepseek. This is Jalapeno. And their

5:41

kernel just doesn't have the correct

5:43

code to make this run as good as it

5:46

possibly can on this. We'll come back

5:49

and dig a little bit deeper because this

5:50

is I think incredibly important to

5:52

understand because what happened in that

5:55

moment is codeex basically wrote

5:58

functional and efficient kernel very

6:00

very quickly to be able to run this

6:02

unusual sort of architecture on this

6:05

chip. Now why is that such a big big

6:08

deal? Because right now, one of the most

6:10

well-known, most used chips for AI is of

6:13

course the GPU that's by Nvidia. And

6:16

Nvidia has CUDA. CUDA is this

6:19

programming platform that has tons of

6:22

stuff, libraries, tools, etc., etc.,

6:24

that allows people to write good kernel

6:27

to run or train their AI models on the

6:30

GPU, the AI chips that Nvidia produces.

6:33

It's been around since like 2007. So

6:35

it's like two decades worth of progress

6:38

and additions and improvements plus very

6:40

simply a lot of engineer expertise,

6:42

right? So there's a lot of smart people

6:44

that have worked on this and spent time

6:46

on this. No one has anything close in

6:49

terms of an ecosystem. This is referred

6:51

to as the CUDA moat because if you're

6:54

writing on CUDA, then it's easy to hire

6:56

people that know the language. There's

6:57

tons of tools. There's tons of

6:59

knowledge. Like this is the path of

7:00

least resistance. It's super easy. But

7:02

if you're using CUDA, then you're using

7:04

Nvidia's chips. If somebody comes out

7:06

with a brand new chip that is awesome

7:08

and amazing and, you know, theoretically

7:10

better than the GPU in order to start

7:12

using it, you have to kind of rebuild

7:14

that kernel and everything everything,

7:16

you know, on whatever new platform or

7:19

you have to do it from scratch. So, I'll

7:21

write like the the skull and bones here,

7:22

right? This is difficult. Nobody wants

7:24

to do it. So, CUDA has 20 years of work

7:27

behind it. thousands, tens of thousands

7:30

of 100 thousand plus engineers, tools,

7:32

libraries, etc. The Jalapeno, you know,

7:35

has like however long it's been around,

7:37

like a month and probably like five

7:39

people that know how it works. But none

7:42

of that matters because OpenAI writes

7:44

Jalapeno kernels like assembly each

7:46

kernel. So that that code that tells the

7:48

chip how to run that particular AI

7:51

model. It gets handtuned code, some

7:53

running to 3,000 lines backed by

7:55

correctness checks and a custom

7:57

sanitizer. Early kernel work was human

7:59

loop but now it's getting fully

8:01

automated but now it's shifting to a

8:03

more scaled up internal version of

8:05

codeex one which openai plans to pitch

8:07

to enterprise customers. So, OpenAI is

8:10

not making this human friendly, right?

8:12

So, they're writing it like assembly.

8:15

They're writing at the most foundational

8:17

level. So, software development, you

8:18

might have these abstraction layers in

8:20

the top. It's something that's readable

8:22

and human friendly and assembly is at

8:24

the bottom rung of that ladder where

8:27

just you're writing everything raw by

8:29

hand. You control everything. The

8:31

machine does exactly what you say.

8:34

Nothing is handled for you. So Codeex

8:37

writes this kernel right at like the

8:41

metal right at that level. So Nvidia

8:44

spent 20 years developing their moat,

8:47

their CUDA moat. This is a big part of

8:49

why they're winning. AMD tried very

8:52

hard, but just their software kind of

8:54

kept disappointing people. It wasn't

8:56

quite as good. They failed to create the

8:58

same robust kind of ecosystem that

9:01

Nvidia has. Google has been spending a

9:03

lot of resources and time on this to try

9:05

to develop their ecosystem to try to

9:07

create something like CUDA. So for their

9:10

TPUs, they were trying to create

9:12

something similar that says TPU. And

9:14

next comes Open AI and they just go lol.

9:18

We're not going to do any of that. They

9:20

just didn't do any of that. And the

9:22

reason they can do that is because let's

9:24

say there are a few thousand people on

9:26

the planet that can write truly great

9:29

kernel code for these applications.

9:31

They're very scarce. There's not a lot

9:33

of them. You have to pay them a lot of

9:34

money and they mainly work with CUDA. If

9:38

AI can write competitive kernel for any

9:41

chip, then that scarcity. It just

9:44

disappears. This post goes into which

9:46

programming language they use for these

9:48

specific kernel tasks. They call it

9:50

Gluon. Gluon is OpenAI is a kernel

9:52

programming language. We're not going to

9:53

go into the details. Again, read the

9:55

post it. It's excellent. But the point

9:57

is, if I'm reading this right, and by

9:59

the way, keep in mind, I'm not an

10:01

expert, but if I'm reading this

10:03

correctly, then here's what happened.

10:05

For the last 20 years, the conventional

10:07

wisdom was basically that you needed to

10:09

build a very developer friendly, human

10:12

friendly ecosystem like CUDA. It needed

10:14

to be very easy, very friendly, and as

10:16

big as as possible, massive, so that

10:18

thousands and thousands of humans can

10:20

program it and and and use it. OpenAI

10:23

kind of flipped the script. They're

10:24

saying actually what probably you need

10:26

is a very small, sharp, tedious, very

10:30

very manual language plus an AI model

10:34

that's smart and more than happy to code

10:36

in that language. So, Gluon is that

10:38

language. Codeex is that AI model and

10:41

this linear layout, that's a type of

10:43

layout algebra that OpenAI invented. As

10:46

a semi- analysis puts it here, in a

10:47

weird twist of fate, OpenAI models like

10:50

GBT 5.6 ICS soul which currently run on

10:52

the Nvidia GPUs have been used to design

10:55

a chip that poses a real threat to the

10:58

CUDA moat. Nvidia's own GPUs are helping

11:00

usher in their potential successor in

11:02

real time. This is the RSI thing that

11:05

we've been talking about the recursive

11:07

self-improvement. It's like a flywheel.

11:09

It's slow at first, but then when it

11:10

gets moving, it it really gets going. We

11:13

train up this model. Then this model is

11:15

able to optimize the chips, design new

11:18

chips, write all the code for those

11:19

chips. It can do it in a way that humans

11:22

can't compete with. Keep in mind, it

11:24

optimized the hardware and also writes

11:26

the software. It handles both sides. Not

11:29

entirely. There's, like they said,

11:31

humans in the loop. There's a lot of

11:32

very smart engineers working on it. So,

11:34

it's not like any of this is 100% fully

11:36

automated yet, as far as we know. By the

11:38

way, the rumors that everyone's hearing

11:40

are just insane. Open recently finished

11:43

the next pre-train code named Bell, the

11:46

successor to Doug, which is expected to

11:48

be the base for Astra and GPT6. It's a

11:51

giant pre-train with over 10 trillion

11:53

total parameters and it could

11:55

potentially be the base for an AGI

11:57

threshold model. Very quick note going

12:00

back to the semi analysis post. Notice

12:02

they say here that that flexibility

12:04

again you can read about what they're

12:05

talking about but that flexibility buys

12:07

valuable optionality for future 10 to 20

12:10

trillion parameter models. So you you

12:13

get that the chip that they build can

12:16

run those models. The the big chunker of

12:18

a model that chip runs it or a 2 to four

12:21

million token context windows. Rumors on

12:24

the rate of progress inside Anthropic

12:26

and OpenAI are truly bonkers. I think

12:28

we'll see a jump the size of one from 03

12:31

to Fable. We'll see that sort of jump

12:35

again in the next 8 months. So, this is

12:37

a chubby on X Kimmanismus. How do you

12:40

say the guy's name? I finally know what

12:42

he looks like. That's him, Kim

12:43

Eisenberg. And he's like the the the AI

12:46

news on Twitter/X. I don't know when and

12:50

if he sleeps because I feel like he's

12:52

just on top of like everything. But

12:54

here's what he posted about three hours

12:56

ago and it really struck a bell with me

12:58

because I'm seeing and hearing more or

13:01

less exactly the same thing. He's

13:03

saying, quote, "I'm currently hearing

13:05

countless rumors about just how

13:06

significant the leaps in capabilities

13:08

will be with the models that have yet to

13:10

be released. The gains are said to be

13:11

substantial across the board at both

13:13

OpenAI and Anthropic. There are also

13:14

persistent rumors that context and

13:16

memory have been solved along with

13:18

reports of self-improving models capable

13:20

of continual learning. Of course, these

13:22

are all just rumors, but given how

13:24

optimistic and bullish these companies

13:25

have been and how cautiously firms such

13:28

as OpenAI have recently acted, delaying

13:30

releases in response to tectonic shifts

13:32

and capabilities, there may be more to

13:33

this than mere speculation. So, as Sam

13:36

Alman puts it here, we made a chip and

13:38

it is fast. So, this program was first

13:40

announced October 2025. Opening eyes

13:42

saying the 9month development cycle

13:45

which is extremely fast for something of

13:49

this nature formally unveiled as

13:51

Jalapeno June 2026. And now we have our

13:54

first benchmarks and third-p partyy

13:57

reports. They're kind of confirming it

13:59

as semian also says you know we still

14:00

have to run their full suite of tests

14:03

but at this point I mean I mean this is

14:06

real. So very small volumes of these

14:08

chips will be used in OpenAI data

14:10

centers this year. They're going to

14:12

start, you know, ramping up using them

14:14

this year and then a ramp up in 2027. So

14:17

notice how Amazon had their training

14:18

chips, Google had their TPUs, Nvidia had

14:21

GPUs, Meta and Microsoft, they had their

14:23

attempts but struggled. And now who's at

14:26

the top of the leaderboard? Who's

14:28

leading the race? It's OpenAI because

14:30

they didn't build the chip first. They

14:33

used other chips, GPUs to train the best

14:36

model, the best coding model. And now

14:38

OpenAI is very independent. They're a

14:42

full stack company now. They have their

14:43

own chips, their own models, they have

14:45

their deployment, product experiences,

14:47

user base. So in conclusion, AI will eat

14:50

everything and we're watching it happen

14:53

live. Let me know what you think about

14:54

this whole thing. I'm still kind of

14:56

processing everything. I'm still trying

14:57

to kind of understand what the heck

15:00

happened. I mean, I've said things like

15:02

this would happen. I've predicted some

15:05

things like this. Not this specific, but

15:07

we've talked about AI models, developing

15:09

hardware and chips. Like, we we've

15:10

talked about all this. And still, when

15:12

it's here and in your face, it it's

15:14

still kind of shocking. You still don't

15:16

expect it to be this blatantly good and

15:19

for it to get developed this blatantly

15:20

fast. And I knew it was going to be

15:22

fast, but now that you're seeing it,

15:23

you're feeling it, you're like, whoa,

15:25

dang. Like, this is getting nuts. So,

15:28

with that eloquent speech at an end,

15:31

I'll bid you a due let me know what you

15:33

think about this whole thing and what do

15:35

you think about Nvidia? How is this

15:36

going to affect them? Let me know in the

15:37

comments. Thank you so much for

15:38

watching. My name is Wes Roth. will see

15:40

you in the next

Interactive Summary

OpenAI has officially unveiled 'Jalapeno', their first custom-designed AI inference chip. Remarkably, this chip outperforms industry standards like Nvidia's Blackwell in performance per watt. The development of this chip, which took only nine months, was accelerated by using AI to design the chip and create its software kernels. By bypassing the need for a human-centric ecosystem like Nvidia's CUDA, OpenAI leverages specialized 'Gluon' code and AI-driven automation to achieve high efficiency, marking a significant shift toward a fully independent, full-stack AI company.

Suggested questions

3 ready-made prompts