HomeVideos

GPT-5.6 Sol Ultrafast.... Just Got UltraFaster (Cerebras 4th Gen)

Now Playing

GPT-5.6 Sol Ultrafast.... Just Got UltraFaster (Cerebras 4th Gen)

Transcript

566 segments

0:00

One of the topics on this channel is

0:01

just how big can your chips get? Now, if

0:04

you've followed any of the videos that

0:06

I've done over the last few years, some

0:08

of the bangers have been about Cerebrus,

0:10

this massive wafer scale engine. It's as

0:13

big as your face or a dinner plate. And

0:16

the idea is that instead of having just

0:18

one chip being your compute engine or

0:21

maybe two chiplets or maybe you know a

0:24

dozen smaller chiplets,

0:26

why not build them all onto the same

0:28

piece of silicon and stitch them

0:30

together? And I don't mean just simply

0:32

3D bond. I mean literally build it into

0:36

the silicon. That's what Cerebras has

0:38

been doing for the last five, six years

0:41

with their wafer scale engine versions

0:43

one, two, and three. We've covered them

0:44

all on this channel. And for me

0:46

personally as an engineer, it's a marvel

0:49

what they can do. They've partnered with

0:51

TSMC to be able to do this this what

0:54

they call cross retical stitching to be

0:55

able to make one uh chip bigger than

0:58

just simply the size that you can print

1:00

chips by doing, you know, little

1:02

connections in between. Not only that,

1:04

this chip built for machine learning is

1:07

able to arguably self-rep and deal with

1:11

however many defects you have from uh

1:13

from the manufacturing. [snorts] What

1:16

this video is about is the next

1:18

generation for that, especially on based

1:21

on the fact that Cerebrus has been in

1:23

partnership with OpenAI to launch 5.6

1:27

Soul, the ultraast edition, now

1:29

available at 750 tokens per second.

1:32

That's on Wave Scale Engine 3. On 3.5,

1:35

they're promising a lot more, but this

1:37

time they've built it. Instead of just a

1:39

single unit, you can now get a rack. And

1:42

the rack comes with a new design for the

1:44

chip called a backpack. Let's talk about

1:46

it.

1:48

>> This is CS4.

1:55

>> Introducing CS4.

1:58

Faster, simpler, and more powerful than

2:01

any previous system. [music]

2:05

One power rack.

2:08

Three compact backpacks [music]

2:12

with a wafer scale engine in each one.

2:16

[music]

2:19

Adaptable to diverse data center

2:21

environments.

2:26

Architected to support future

2:28

generations of wafers scale compute.

2:30

[music]

2:34

Optimized for automated manufacturing

2:36

and deployment at scale. A revolutionary

2:40

system built to deploy the fastest AI

2:42

across hyperscaled data centers.

2:45

Powering the next generation of

2:47

ultraast. [music]

2:51

This is CS4.

2:55

>> [music]

3:08

>> This

3:10

is the S4.

3:12

This is our next generation system.

3:16

>> What's your minimum specification?

3:23

>> So, let's set the scene. If you were to

3:24

buy one of the Cerebrus' chips, the

3:27

previous ones, webcale engine one, two,

3:29

three, you built it in a 16U form factor

3:33

design. It was all self-contained.

3:35

Within that, you were liquid cooled. Uh,

3:38

but in terms of the data center, you

3:40

were air cooled. It would be around 24

3:42

kW built on TSMC N5 with 900,000 cores,

3:48

44 GB of SRAMM and 21 pabytes pers of uh

3:53

SRAMM uh bandwidth in order to enable

3:56

that fast speed. Putting it all together

3:59

meant that if you applied it to machine

4:01

learning, now the company's been through

4:03

a few cycles. Initially it was

4:04

convolutional neural networks, then it

4:07

was uh endto-end training and now we're

4:09

talking more about inference. You pull

4:11

that on board and even with the latest

4:14

models, you can get super fast

4:16

performance because all your data, all

4:18

your activations and all your weights

4:20

stay on the chip. With 44 GB of SRAMM,

4:25

you can't fit big models on a single

4:27

chip. But the way they designed the

4:29

system is that within 2 and a half

4:31

megawws initially you can put 64 chips

4:34

though that we have increased efficiency

4:37

over time. So you can go from 64 to 88

4:40

to 92. The point being is that you

4:43

aren't limited just to small models here

4:45

with enough chips. You can do large

4:47

models lightning fast speed. And that's

4:50

what kind of went through the OpenAI

4:52

deal that was announced earlier this

4:54

year. uh 750 megawatts, gigawatts, uh I

4:58

think I did the math once. We're talking

4:59

about thousands and thousands of chips,

5:01

about 20 billion dollars of revenue for

5:04

Cerebras. And that's kind of why where

5:06

they IPOed earlier this year with a big

5:08

fanfare. Uh I think the initial IPO

5:11

launch was about 55 um billion dollars.

5:13

They reached a peak of about 90 billion

5:16

uh within that first day. It's around

5:17

about more like 44 billion 40 billion

5:20

now um market cap. but they still offer

5:22

a unique product that nobody else in the

5:24

market does. Now, recently when I've

5:28

spoken with our good friend Sally Wood

5:30

Foxton over at E Times, uh, and she

5:32

follows Cerebras just as closely as I

5:33

do, we've been wondering about next

5:35

generation because Waver Scale Engine 3

5:38

was announced two years ago. Uh, I've

5:41

got a big video on this and you can

5:42

catch that. The we've always been asking

5:45

them, what's next? What's next? What's

5:46

next? And I've had conversations with

5:48

CEO uh Andrew Feldman and some of the

5:51

people there about what's coming next.

5:53

And today as I'm filming this actually

5:56

uh it's going to be announced I guess in

5:58

about an hour or so. We were lucky to

5:59

have a early look at media day. They're

6:02

going to be launching the next

6:04

generation of the wave scale engine. Uh

6:06

but it comes in two parts and I kind of

6:08

want to go through those two parts with

6:09

you.

6:11

First up is the chip, right? The chip is

6:14

the wafer scale engine and that goes

6:16

into the cerebrra system. So this is why

6:18

we have WSE for wafer scale and CS for

6:20

cerebrra system. WSSE 3.5 is the new

6:25

chip. And the reason it's called.5 is

6:28

because it's a turbo edition of the wave

6:31

scale engine 3. If we were talking in

6:33

terms of Intel and AMD CPU terms, uh

6:36

this would be effectively binned a lot

6:38

more aggressively. though say that

6:40

they've done a lot more than that in

6:42

able to get multiple times more

6:44

performance out of this chip compared to

6:46

the standard wave scale engine 3. We're

6:49

talking about pretty much double the

6:51

frequency here. Uh which does mean

6:53

double the power. Uh but on top of that

6:55

it also means a lot more in terms of

6:57

memory bandwidth and the ability to make

7:00

larger models inside large clusters of

7:03

these chips work together. So previously

7:07

we were talking uh you know around about

7:10

125 petlops of sparse FP16. We're now

7:13

going up to 250

7:16

um of that using the same TSMC N5 using

7:19

the same 900,000

7:21

uh cores and with the same principle

7:23

that in terms of physical yield out of

7:26

TSMC they get roughly 100%. You know

7:29

it's like this t-shirt yield at what

7:31

dice size? You can find it at

7:32

shop.teato.com.

7:34

Now the way they work, the way this chip

7:37

works as all the other previous ones

7:38

have done is that if a core of the

7:41

900,000 has a defect, they can smart

7:44

route around it. Uh you you know

7:46

basically firmware they can disable it

7:48

and they have the routing to enable

7:50

that. Um that enables 100% physical

7:54

yield. Now parametric yield based on uh

7:56

frequency and voltage has always been um

7:59

a talking point externally uh based on

8:01

Cerebris's uh margins that they now have

8:04

to provide in their financials. But this

8:06

chip they're saying is something that

8:07

they've managed to majorly increase uh

8:10

the frequency and as a result you'd have

8:12

a voltage increase but also design to

8:15

that effect. So it's not just simply

8:17

having the same design and bin, it's

8:19

optimized what they do have with TSMC

8:21

with their foundry partner on the same

8:23

node to get this massive increase in

8:26

performance. Now this wafer scale chip,

8:29

same as generation 1, two, and three,

8:30

you don't have it flat, you have it

8:33

vertical. This allows for cooling and

8:36

power to come in and then IO from the

8:39

sides. What's different this time around

8:41

is that instead of having it in this big

8:44

16U design where arguably it was built

8:48

for a unit, Cirus has optimized it so

8:51

that you can now put three of these in a

8:53

rack. And it's the rack design that

8:57

makes this a step above just simply a

9:00

skewing difference.

9:02

The rack design, the CS4,

9:05

uh, has three of these systems, of which

9:06

at the front you have a set of fans and

9:09

power supplies. And at the back, you

9:12

have what's affectionately called the

9:14

backpack. And it looks a little like

9:16

this. As you can see, there are massive

9:19

tubes. There's a space here for uh the

9:22

back for the chip to sit vertically. And

9:24

the idea is that in front of this and

9:26

behind this, you have the power delivery

9:28

and the cooling. Uh, I believe one of

9:30

these chips is probably pushing based on

9:32

how many power supplies there are, 40

9:34

kW. That means three in a rack is about

9:36

120, which kind of sounds like Nvidia's

9:39

NVL72 of course. Um, but what's

9:42

interesting here is that ne built onto

9:44

this backpack is how the IO is managed.

9:47

So, one change for this chip based on

9:49

previous generation is that they have

9:51

twice the networking 2.4 terabits per

9:53

second networking coming off of each

9:55

chip. um those networking modules are

9:58

actually modular and the way it's

10:01

cerebrus and CEO Andrew Falman and uh

10:04

also Sean Lee uh the CTO have explained

10:06

it to me and to basically to everybody

10:09

else is that in the future if there are

10:11

other networking standards or other

10:12

networking connectivity they need these

10:15

modules can be replaced with that new

10:18

feature with that new function. So if

10:20

you need more chip to chip or more

10:22

switch or more external switch uh or you

10:25

need 800 gig connections or more 200 gig

10:28

connections or 100 gig connections then

10:31

this modularity will allow that and make

10:33

it a lot easier with this essentially

10:37

they're building a platform to be honest

10:40

where you've got the power in the front

10:42

with a few fans and then in the back

10:44

you've got liquid cooling which comes

10:46

out which uh connects through the top of

10:48

the rack.

10:49

into the wafer. Um, assuming that future

10:52

wafers uh are going to be of a similar

10:54

size because they're literally at the

10:56

max of what a wafer can be, it should be

10:58

very easy to put in. You can have, you

11:00

know, different power input. You roughly

11:02

the same amount of cooling or perhaps

11:03

you'll be able to use chilled water and

11:05

do more cooling. Uh and then you have

11:07

this modular um networking coming above

11:10

and below uh that enables you know

11:13

future chips uh to then use different

11:16

sorts of uh scale up and scale out

11:18

networking and ultimately yeah we get

11:21

what looks like a backpack or

11:22

realistically looks like something that

11:24

should be on the set of Ghostbusters. Um

11:27

I was actually speaking with a couple of

11:28

other press on site and they were saying

11:30

that the the yellow box they should have

11:32

you know put a orange box in the middle

11:33

they should have put a flux capacitor

11:35

inside. Um but the the the idea here is

11:38

that altogether they've got a chip

11:40

that's um twice the performance

11:43

supporting 10x larger models by simply

11:46

enabling large scale support. uh they've

11:49

also uh due to the modularity of the

11:52

networking improved the chip to chip

11:55

latency uh such that what used to be

11:57

five microsconds is now two microsconds

12:00

uh for chip to chip in fast mode I think

12:02

they have a regular mode which is more

12:03

like three microsconds but the idea here

12:06

is that for the customers who want this

12:08

chip and they want them you know at

12:11

scale to supply the one trillion

12:14

parameter models or the three trillion

12:16

parameters models or even the 10

12:17

trillion parameter models going forward.

12:20

This system is designed to be scalable

12:22

three in a rack, many racks in a data

12:24

center um and provide that performance

12:28

uh that they can. So, wave scale engine

12:31

3 previous generation did about 2300

12:35

tokens per second um on a standard model

12:37

that fit all within memory I think in

12:39

one chip. Um with this new chip, you can

12:41

now get 4,400 tokens per second. Uh what

12:45

this means is for like a big model for

12:47

like a GBT 5.6 Soul uh what can do 750

12:51

tokens today can do almost 1500 tokens

12:54

per second uh when this gets installed.

12:56

I did ask the CEO because of this

12:59

massive agreement with OpenAI. At the

13:01

time we thought, okay, that's a lot of

13:03

CS3s are going to be deploying. Um, does

13:05

this actually mean that part of that

13:07

agreement is, you know, uh, CS4s with

13:09

this new chip or maybe even CS5s and

13:12

CS6s? And he said, look, when these, uh,

13:14

agreements are made, there's obviously

13:16

some supply that's of the current

13:18

generation, but you go into it knowing

13:19

that a lot of that uh, revenue is going

13:22

to come from future generation chips. So

13:24

using the CS4 uh you know is going to be

13:27

a base to provide some of that

13:28

performance going forward is going to be

13:30

a key part of that arrangement. One

13:33

other thing here on the CS4 um Cerebras

13:37

were fairly open on this cuz when they

13:39

built the first generation um system

13:41

with the first generation wafer scale

13:43

they were still a relatively young

13:45

company. So they were building it, you

13:47

know, knowing that it was going to be

13:48

small volume, but it meant that they

13:50

didn't necessarily do things in the most

13:52

efficient way in terms of, you know,

13:53

supply chain and building out the

13:56

systems. And we're talking like

13:57

machining custom parts and stuff. And so

13:59

in reality, even though you've got the

14:01

chip and the chip costs what the chip

14:03

costs, the rest of the infrastructure

14:06

costs a lot more than what it probably

14:07

should have done if it was done at

14:09

scale. What they've done with the new

14:10

generation with the CS4 with this

14:12

platform is reduce components by about

14:16

half um or I believe or maybe even 60%.

14:20

But they've also managed to automate it

14:21

a lot more. So one of the issues that

14:24

keeps coming up in cerebrous calls is

14:26

that can cerebras meet demand by with

14:29

how many systems they can bring to

14:31

market. And it turns out if you make the

14:32

systems easier to build and more

14:34

automatable, you can build more systems.

14:37

Uh let's just put it this way. they

14:40

don't need to be fighting over capacity

14:42

at TSMC uh because they're making you

14:45

know one big chip and okay they might

14:47

sell a few thousand at this point a year

14:50

going from a few thousand to 10,000

14:51

isn't that big a jump for TSMC compared

14:54

to say another vendor that might want to

14:56

go from a 100,000 wafers a month to

14:59

200,000 wafers a month is still you know

15:02

smaller quantities than most but it

15:04

costs a lot more uh per design

15:07

did say that this generation ation. This

15:09

new generation will cost more than last

15:11

generation. By how much they wouldn't

15:13

say um but they do have early systems

15:17

shipping now and more ramping to volume

15:20

by the end of the year. And also on top

15:22

of that that because they've built a

15:24

platform, this is a platform that will

15:26

have the next generation wafer scale

15:29

engine in uh and then CS uh 5, CS6 and I

15:33

think they said CS5 is coming second

15:35

half of 2027. So, they're wanting to

15:37

build more of a regular cadence on these

15:40

sorts of things. So, I'll definitely be

15:42

there at that event making sure u or at

15:44

least seeing what they're bringing to

15:46

market.

15:47

So, yeah, in terms of benchmarks,

15:49

they're talking here where you use um

15:52

GPUs for the prefill and then Cerebras

15:55

here for the decode. We're talking GPTS

15:57

12 billion. If you use GPU only, you

16:00

might be getting 250 tokens per second.

16:03

uh with Cerebrus with their new CS4 you

16:06

can get over 4,500 tokens per second. Uh

16:10

and you can see a bunch of other lines

16:11

on this graph. Um Kimmy K2.7 1 trillion

16:15

parameter uh you're going up to 3,000

16:18

tokens per second. And then here right

16:20

at the end you've got the GPT 5.6 SUL

16:23

going up all the way to 1500 tokens per

16:25

second.

16:26

Gen on genen. These are sizable

16:28

increases as you can imagine. Um but

16:31

ultimately it's going to be what they

16:32

end up uh being used for. One of the big

16:35

things at the event actually was this

16:37

talk about disagregation, about prefill

16:40

decode, about splitting your workload up

16:42

into different segments and then running

16:44

uh these segments on the most ideal

16:46

hardware for the type of compute you

16:48

need. It's uh something that we used to

16:51

call partitioning back in the day. And I

16:53

think disagregation is the wrong word to

16:55

be talking about it. We should be

16:56

talking about partitioning workloads. Um

16:58

and the idea is that if you partition

17:01

the thing that needs compute to compute

17:02

hardware and partition things that need

17:04

memory to memory hardware in this case

17:06

because cerebrus is so has such better

17:09

memory performance then this is where

17:12

you get the extra four. Now you might

17:15

think that this is why Nvidia bought

17:18

Grock, right? Grock is an SRAM based

17:20

design and they wanted all the extra

17:22

performance for the decode so they can

17:25

put up some you know somewhat similar

17:27

numbers to Cerebrus. Um and obviously

17:30

Nvidia uh talks a big game when it comes

17:33

to their Grock. I'm more interested in

17:35

what happens when Grock integrates

17:36

NVLink. That's where I think it'll

17:38

actually be competitive because I'm not

17:41

actually hearing much interest for the

17:42

Grock solution with Nvidia right now.

17:45

Um, I've seen a few slides that say that

17:47

for every rack of NVL72, you need nine

17:50

racks of Grock. Um, and you know, the

17:52

power consumption therein, that's not to

17:54

say that Cerebrus, you know, is is is

17:56

power consumption heavy. You need a fair

17:58

amount of racks, but the performance

17:59

you're getting out Cerebrus, uh, looks

18:01

like it blows away anything that Grock

18:03

can provide right now. Uh, and the fact

18:05

that Cerebrris is still an independent

18:07

company means that they can go and work

18:09

with companies like AMD. They announced

18:11

at AMD's advancing AI that they'll be

18:14

selling Helios plus Cerebrus systems. So

18:16

that's AMD's MI455X plus Cerebra systems

18:20

for those who need ultra fast tokens per

18:23

second. They've also done an agreement

18:24

with AWS so that if you want tranium 3

18:28

with fast language model inference then

18:31

you can do that as well. Uh I asked on

18:33

stage whether they play nice with Nvidia

18:36

and it turns out yes they do. If anybody

18:38

wants an Nvidia installation, they do

18:40

well with that uh also. Uh makes me

18:42

wonder exactly who OpenAI are using um

18:45

instead of you know Cerebras for the GPU

18:48

side of things. But for them they want

18:51

the integration of the hardware to be

18:54

relatively simple when you need high

18:56

performance um or high speed tokens of

19:00

large models that are more accurate than

19:03

say the smaller models. I recently did a

19:05

video about Talis and the fact that they

19:08

can build dedicated uh single model

19:11

chips that only run a single small model

19:13

at like 17,000 tokens per second. And

19:15

based on all the comments on that video,

19:17

you guys really liked the high speed

19:19

that that provides. The thing is that

19:21

hardware is obviously very model

19:23

specific. With Cerebras, you can run

19:25

essentially any transformer model or any

19:28

series of models um at a high speed, you

19:31

know, but the trade-off is obviously

19:32

configurability with power and

19:34

deployment and the fact that Cerebras

19:36

has got a lot of business already,

19:38

whereas Talis was just acquired by AMD.

19:42

The idea here is that in any workflow,

19:44

if you're writing code, you want your

19:47

code to be done now, not still working

19:50

when you're have to go off and get a

19:52

coffee every time you have an input. You

19:54

know, I have this problem when I'm doing

19:56

my script kitty stuff um with with some

19:58

of the large language models. I have to

20:00

wait for it to be done before I can

20:02

decide what to do next. With speed, the

20:04

ability is that you can run ideas

20:06

concurrently a lot faster uh and

20:09

actually get to those results sooner. So

20:11

the idea is more tokens, more tokens,

20:13

more tokens at high performance with

20:16

these large models in mind. That's why

20:18

Cerebrris is saying that by the end of

20:20

the decade they want to massively

20:22

increase performance and massively

20:24

increase um support of large models up

20:27

to say uh 10 trillion parameters which

20:30

is kind of what the rest of the market

20:32

was talking about anyway. Now obviously

20:35

as a user you can't interact with

20:38

thousands of tokens per second but if

20:40

you're working on a large database and

20:42

things need to be edited here there and

20:43

everywhere um or you've got models

20:45

talking to each other in an agentic

20:47

workflow then the faster they talk to

20:49

each other the faster you can get your

20:51

result or whatever backend processing

20:53

needs to happen. And right now there is

20:56

a market for this. people are prepared

20:58

to pay extra for those tokens um in

21:01

terms of you know million dollars per

21:03

million tokens and this is where

21:05

Cerebras sees their value as a business

21:08

going forward. Demand for speed will

21:10

always be there. It's even got to the

21:12

point where some companies are saying

21:14

CUDA is not the moat. Speed is the moat

21:17

and Cerebrus is one of those chips that

21:19

gets you there with that speed. Now,

21:22

you'll know on this channel I cover a

21:24

bunch of AI hardware. Um, and some of it

21:27

is really cool. Some of it's doing it in

21:29

in such impressive ways. Um, in speed,

21:32

Cerebrus is king. There are others who

21:35

deal with larger context lengths that

21:37

deal with, you know, special mathematics

21:38

to make it more energy efficient, but in

21:40

speed, Cerebrus is still king. And it's

21:44

really good that they've announced, you

21:45

know, even though it's say a 3.5 and not

21:47

a four, that we have new silicon. Um, I

21:50

really wanted actually them to come out

21:52

with, you know, like a backpack that

21:54

people could wear. Um, no doubt the

21:56

thing weighs, you know, a couple of

21:57

hundred kilos. Anyway, but the idea is

21:59

that now they're building a rack, 120 kW

22:03

roughly rack with three systems inside

22:06

and not just one 16U box. Um,

22:11

next thing to ask is whether they'll

22:12

have a buy it now on their website. So,

22:15

let me know your thoughts. Is Cerebras a

22:17

company you're interested in? I'm not

22:18

invested in them in any way. Uh they're

22:21

not a client of mine at this point.

22:23

Fingers crossed. Um but let me know what

22:26

you think about these guys, what you

22:27

want to see next, because they

22:28

definitely want to hear feedback from

22:29

the launch, from the event, from the

22:31

product. Uh cuz, uh they've got a team

22:34

that's, you know, going through all of

22:35

these, making sure the next generation

22:37

works better. So, please do leave

22:39

comments and questions for the CEO.

22:40

Maybe we'll get him on uh camera later

22:43

about what you want to hear about Cirrus

22:45

and what makes you think that they're

22:47

either good or bad um or what's coming

22:50

next.

Interactive Summary

This video details the launch of the Cerebras CS4 system, a high-performance AI platform built around the new WSE-3.5 wafer-scale engine. The system utilizes a novel 'backpack' design within a rack-mount configuration, significantly increasing performance, memory bandwidth, and scalability for large language model inference. The video also discusses Cerebras's market position, their partnership with companies like OpenAI, AMD, and AWS, and their strategy to differentiate themselves through raw speed rather than relying on traditional GPU ecosystems.

Suggested questions

4 ready-made prompts