GPT-5.6 Sol Ultrafast.... Just Got UltraFaster (Cerebras 4th Gen)
566 segments
One of the topics on this channel is
just how big can your chips get? Now, if
you've followed any of the videos that
I've done over the last few years, some
of the bangers have been about Cerebrus,
this massive wafer scale engine. It's as
big as your face or a dinner plate. And
the idea is that instead of having just
one chip being your compute engine or
maybe two chiplets or maybe you know a
dozen smaller chiplets,
why not build them all onto the same
piece of silicon and stitch them
together? And I don't mean just simply
3D bond. I mean literally build it into
the silicon. That's what Cerebras has
been doing for the last five, six years
with their wafer scale engine versions
one, two, and three. We've covered them
all on this channel. And for me
personally as an engineer, it's a marvel
what they can do. They've partnered with
TSMC to be able to do this this what
they call cross retical stitching to be
able to make one uh chip bigger than
just simply the size that you can print
chips by doing, you know, little
connections in between. Not only that,
this chip built for machine learning is
able to arguably self-rep and deal with
however many defects you have from uh
from the manufacturing. [snorts] What
this video is about is the next
generation for that, especially on based
on the fact that Cerebrus has been in
partnership with OpenAI to launch 5.6
Soul, the ultraast edition, now
available at 750 tokens per second.
That's on Wave Scale Engine 3. On 3.5,
they're promising a lot more, but this
time they've built it. Instead of just a
single unit, you can now get a rack. And
the rack comes with a new design for the
chip called a backpack. Let's talk about
it.
>> This is CS4.
>> Introducing CS4.
Faster, simpler, and more powerful than
any previous system. [music]
One power rack.
Three compact backpacks [music]
with a wafer scale engine in each one.
[music]
Adaptable to diverse data center
environments.
Architected to support future
generations of wafers scale compute.
[music]
Optimized for automated manufacturing
and deployment at scale. A revolutionary
system built to deploy the fastest AI
across hyperscaled data centers.
Powering the next generation of
ultraast. [music]
This is CS4.
>> [music]
>> This
is the S4.
This is our next generation system.
>> What's your minimum specification?
>> So, let's set the scene. If you were to
buy one of the Cerebrus' chips, the
previous ones, webcale engine one, two,
three, you built it in a 16U form factor
design. It was all self-contained.
Within that, you were liquid cooled. Uh,
but in terms of the data center, you
were air cooled. It would be around 24
kW built on TSMC N5 with 900,000 cores,
44 GB of SRAMM and 21 pabytes pers of uh
SRAMM uh bandwidth in order to enable
that fast speed. Putting it all together
meant that if you applied it to machine
learning, now the company's been through
a few cycles. Initially it was
convolutional neural networks, then it
was uh endto-end training and now we're
talking more about inference. You pull
that on board and even with the latest
models, you can get super fast
performance because all your data, all
your activations and all your weights
stay on the chip. With 44 GB of SRAMM,
you can't fit big models on a single
chip. But the way they designed the
system is that within 2 and a half
megawws initially you can put 64 chips
though that we have increased efficiency
over time. So you can go from 64 to 88
to 92. The point being is that you
aren't limited just to small models here
with enough chips. You can do large
models lightning fast speed. And that's
what kind of went through the OpenAI
deal that was announced earlier this
year. uh 750 megawatts, gigawatts, uh I
think I did the math once. We're talking
about thousands and thousands of chips,
about 20 billion dollars of revenue for
Cerebras. And that's kind of why where
they IPOed earlier this year with a big
fanfare. Uh I think the initial IPO
launch was about 55 um billion dollars.
They reached a peak of about 90 billion
uh within that first day. It's around
about more like 44 billion 40 billion
now um market cap. but they still offer
a unique product that nobody else in the
market does. Now, recently when I've
spoken with our good friend Sally Wood
Foxton over at E Times, uh, and she
follows Cerebras just as closely as I
do, we've been wondering about next
generation because Waver Scale Engine 3
was announced two years ago. Uh, I've
got a big video on this and you can
catch that. The we've always been asking
them, what's next? What's next? What's
next? And I've had conversations with
CEO uh Andrew Feldman and some of the
people there about what's coming next.
And today as I'm filming this actually
uh it's going to be announced I guess in
about an hour or so. We were lucky to
have a early look at media day. They're
going to be launching the next
generation of the wave scale engine. Uh
but it comes in two parts and I kind of
want to go through those two parts with
you.
First up is the chip, right? The chip is
the wafer scale engine and that goes
into the cerebrra system. So this is why
we have WSE for wafer scale and CS for
cerebrra system. WSSE 3.5 is the new
chip. And the reason it's called.5 is
because it's a turbo edition of the wave
scale engine 3. If we were talking in
terms of Intel and AMD CPU terms, uh
this would be effectively binned a lot
more aggressively. though say that
they've done a lot more than that in
able to get multiple times more
performance out of this chip compared to
the standard wave scale engine 3. We're
talking about pretty much double the
frequency here. Uh which does mean
double the power. Uh but on top of that
it also means a lot more in terms of
memory bandwidth and the ability to make
larger models inside large clusters of
these chips work together. So previously
we were talking uh you know around about
125 petlops of sparse FP16. We're now
going up to 250
um of that using the same TSMC N5 using
the same 900,000
uh cores and with the same principle
that in terms of physical yield out of
TSMC they get roughly 100%. You know
it's like this t-shirt yield at what
dice size? You can find it at
shop.teato.com.
Now the way they work, the way this chip
works as all the other previous ones
have done is that if a core of the
900,000 has a defect, they can smart
route around it. Uh you you know
basically firmware they can disable it
and they have the routing to enable
that. Um that enables 100% physical
yield. Now parametric yield based on uh
frequency and voltage has always been um
a talking point externally uh based on
Cerebris's uh margins that they now have
to provide in their financials. But this
chip they're saying is something that
they've managed to majorly increase uh
the frequency and as a result you'd have
a voltage increase but also design to
that effect. So it's not just simply
having the same design and bin, it's
optimized what they do have with TSMC
with their foundry partner on the same
node to get this massive increase in
performance. Now this wafer scale chip,
same as generation 1, two, and three,
you don't have it flat, you have it
vertical. This allows for cooling and
power to come in and then IO from the
sides. What's different this time around
is that instead of having it in this big
16U design where arguably it was built
for a unit, Cirus has optimized it so
that you can now put three of these in a
rack. And it's the rack design that
makes this a step above just simply a
skewing difference.
The rack design, the CS4,
uh, has three of these systems, of which
at the front you have a set of fans and
power supplies. And at the back, you
have what's affectionately called the
backpack. And it looks a little like
this. As you can see, there are massive
tubes. There's a space here for uh the
back for the chip to sit vertically. And
the idea is that in front of this and
behind this, you have the power delivery
and the cooling. Uh, I believe one of
these chips is probably pushing based on
how many power supplies there are, 40
kW. That means three in a rack is about
120, which kind of sounds like Nvidia's
NVL72 of course. Um, but what's
interesting here is that ne built onto
this backpack is how the IO is managed.
So, one change for this chip based on
previous generation is that they have
twice the networking 2.4 terabits per
second networking coming off of each
chip. um those networking modules are
actually modular and the way it's
cerebrus and CEO Andrew Falman and uh
also Sean Lee uh the CTO have explained
it to me and to basically to everybody
else is that in the future if there are
other networking standards or other
networking connectivity they need these
modules can be replaced with that new
feature with that new function. So if
you need more chip to chip or more
switch or more external switch uh or you
need 800 gig connections or more 200 gig
connections or 100 gig connections then
this modularity will allow that and make
it a lot easier with this essentially
they're building a platform to be honest
where you've got the power in the front
with a few fans and then in the back
you've got liquid cooling which comes
out which uh connects through the top of
the rack.
into the wafer. Um, assuming that future
wafers uh are going to be of a similar
size because they're literally at the
max of what a wafer can be, it should be
very easy to put in. You can have, you
know, different power input. You roughly
the same amount of cooling or perhaps
you'll be able to use chilled water and
do more cooling. Uh and then you have
this modular um networking coming above
and below uh that enables you know
future chips uh to then use different
sorts of uh scale up and scale out
networking and ultimately yeah we get
what looks like a backpack or
realistically looks like something that
should be on the set of Ghostbusters. Um
I was actually speaking with a couple of
other press on site and they were saying
that the the yellow box they should have
you know put a orange box in the middle
they should have put a flux capacitor
inside. Um but the the the idea here is
that altogether they've got a chip
that's um twice the performance
supporting 10x larger models by simply
enabling large scale support. uh they've
also uh due to the modularity of the
networking improved the chip to chip
latency uh such that what used to be
five microsconds is now two microsconds
uh for chip to chip in fast mode I think
they have a regular mode which is more
like three microsconds but the idea here
is that for the customers who want this
chip and they want them you know at
scale to supply the one trillion
parameter models or the three trillion
parameters models or even the 10
trillion parameter models going forward.
This system is designed to be scalable
three in a rack, many racks in a data
center um and provide that performance
uh that they can. So, wave scale engine
3 previous generation did about 2300
tokens per second um on a standard model
that fit all within memory I think in
one chip. Um with this new chip, you can
now get 4,400 tokens per second. Uh what
this means is for like a big model for
like a GBT 5.6 Soul uh what can do 750
tokens today can do almost 1500 tokens
per second uh when this gets installed.
I did ask the CEO because of this
massive agreement with OpenAI. At the
time we thought, okay, that's a lot of
CS3s are going to be deploying. Um, does
this actually mean that part of that
agreement is, you know, uh, CS4s with
this new chip or maybe even CS5s and
CS6s? And he said, look, when these, uh,
agreements are made, there's obviously
some supply that's of the current
generation, but you go into it knowing
that a lot of that uh, revenue is going
to come from future generation chips. So
using the CS4 uh you know is going to be
a base to provide some of that
performance going forward is going to be
a key part of that arrangement. One
other thing here on the CS4 um Cerebras
were fairly open on this cuz when they
built the first generation um system
with the first generation wafer scale
they were still a relatively young
company. So they were building it, you
know, knowing that it was going to be
small volume, but it meant that they
didn't necessarily do things in the most
efficient way in terms of, you know,
supply chain and building out the
systems. And we're talking like
machining custom parts and stuff. And so
in reality, even though you've got the
chip and the chip costs what the chip
costs, the rest of the infrastructure
costs a lot more than what it probably
should have done if it was done at
scale. What they've done with the new
generation with the CS4 with this
platform is reduce components by about
half um or I believe or maybe even 60%.
But they've also managed to automate it
a lot more. So one of the issues that
keeps coming up in cerebrous calls is
that can cerebras meet demand by with
how many systems they can bring to
market. And it turns out if you make the
systems easier to build and more
automatable, you can build more systems.
Uh let's just put it this way. they
don't need to be fighting over capacity
at TSMC uh because they're making you
know one big chip and okay they might
sell a few thousand at this point a year
going from a few thousand to 10,000
isn't that big a jump for TSMC compared
to say another vendor that might want to
go from a 100,000 wafers a month to
200,000 wafers a month is still you know
smaller quantities than most but it
costs a lot more uh per design
did say that this generation ation. This
new generation will cost more than last
generation. By how much they wouldn't
say um but they do have early systems
shipping now and more ramping to volume
by the end of the year. And also on top
of that that because they've built a
platform, this is a platform that will
have the next generation wafer scale
engine in uh and then CS uh 5, CS6 and I
think they said CS5 is coming second
half of 2027. So, they're wanting to
build more of a regular cadence on these
sorts of things. So, I'll definitely be
there at that event making sure u or at
least seeing what they're bringing to
market.
So, yeah, in terms of benchmarks,
they're talking here where you use um
GPUs for the prefill and then Cerebras
here for the decode. We're talking GPTS
12 billion. If you use GPU only, you
might be getting 250 tokens per second.
uh with Cerebrus with their new CS4 you
can get over 4,500 tokens per second. Uh
and you can see a bunch of other lines
on this graph. Um Kimmy K2.7 1 trillion
parameter uh you're going up to 3,000
tokens per second. And then here right
at the end you've got the GPT 5.6 SUL
going up all the way to 1500 tokens per
second.
Gen on genen. These are sizable
increases as you can imagine. Um but
ultimately it's going to be what they
end up uh being used for. One of the big
things at the event actually was this
talk about disagregation, about prefill
decode, about splitting your workload up
into different segments and then running
uh these segments on the most ideal
hardware for the type of compute you
need. It's uh something that we used to
call partitioning back in the day. And I
think disagregation is the wrong word to
be talking about it. We should be
talking about partitioning workloads. Um
and the idea is that if you partition
the thing that needs compute to compute
hardware and partition things that need
memory to memory hardware in this case
because cerebrus is so has such better
memory performance then this is where
you get the extra four. Now you might
think that this is why Nvidia bought
Grock, right? Grock is an SRAM based
design and they wanted all the extra
performance for the decode so they can
put up some you know somewhat similar
numbers to Cerebrus. Um and obviously
Nvidia uh talks a big game when it comes
to their Grock. I'm more interested in
what happens when Grock integrates
NVLink. That's where I think it'll
actually be competitive because I'm not
actually hearing much interest for the
Grock solution with Nvidia right now.
Um, I've seen a few slides that say that
for every rack of NVL72, you need nine
racks of Grock. Um, and you know, the
power consumption therein, that's not to
say that Cerebrus, you know, is is is
power consumption heavy. You need a fair
amount of racks, but the performance
you're getting out Cerebrus, uh, looks
like it blows away anything that Grock
can provide right now. Uh, and the fact
that Cerebrris is still an independent
company means that they can go and work
with companies like AMD. They announced
at AMD's advancing AI that they'll be
selling Helios plus Cerebrus systems. So
that's AMD's MI455X plus Cerebra systems
for those who need ultra fast tokens per
second. They've also done an agreement
with AWS so that if you want tranium 3
with fast language model inference then
you can do that as well. Uh I asked on
stage whether they play nice with Nvidia
and it turns out yes they do. If anybody
wants an Nvidia installation, they do
well with that uh also. Uh makes me
wonder exactly who OpenAI are using um
instead of you know Cerebras for the GPU
side of things. But for them they want
the integration of the hardware to be
relatively simple when you need high
performance um or high speed tokens of
large models that are more accurate than
say the smaller models. I recently did a
video about Talis and the fact that they
can build dedicated uh single model
chips that only run a single small model
at like 17,000 tokens per second. And
based on all the comments on that video,
you guys really liked the high speed
that that provides. The thing is that
hardware is obviously very model
specific. With Cerebras, you can run
essentially any transformer model or any
series of models um at a high speed, you
know, but the trade-off is obviously
configurability with power and
deployment and the fact that Cerebras
has got a lot of business already,
whereas Talis was just acquired by AMD.
The idea here is that in any workflow,
if you're writing code, you want your
code to be done now, not still working
when you're have to go off and get a
coffee every time you have an input. You
know, I have this problem when I'm doing
my script kitty stuff um with with some
of the large language models. I have to
wait for it to be done before I can
decide what to do next. With speed, the
ability is that you can run ideas
concurrently a lot faster uh and
actually get to those results sooner. So
the idea is more tokens, more tokens,
more tokens at high performance with
these large models in mind. That's why
Cerebrris is saying that by the end of
the decade they want to massively
increase performance and massively
increase um support of large models up
to say uh 10 trillion parameters which
is kind of what the rest of the market
was talking about anyway. Now obviously
as a user you can't interact with
thousands of tokens per second but if
you're working on a large database and
things need to be edited here there and
everywhere um or you've got models
talking to each other in an agentic
workflow then the faster they talk to
each other the faster you can get your
result or whatever backend processing
needs to happen. And right now there is
a market for this. people are prepared
to pay extra for those tokens um in
terms of you know million dollars per
million tokens and this is where
Cerebras sees their value as a business
going forward. Demand for speed will
always be there. It's even got to the
point where some companies are saying
CUDA is not the moat. Speed is the moat
and Cerebrus is one of those chips that
gets you there with that speed. Now,
you'll know on this channel I cover a
bunch of AI hardware. Um, and some of it
is really cool. Some of it's doing it in
in such impressive ways. Um, in speed,
Cerebrus is king. There are others who
deal with larger context lengths that
deal with, you know, special mathematics
to make it more energy efficient, but in
speed, Cerebrus is still king. And it's
really good that they've announced, you
know, even though it's say a 3.5 and not
a four, that we have new silicon. Um, I
really wanted actually them to come out
with, you know, like a backpack that
people could wear. Um, no doubt the
thing weighs, you know, a couple of
hundred kilos. Anyway, but the idea is
that now they're building a rack, 120 kW
roughly rack with three systems inside
and not just one 16U box. Um,
next thing to ask is whether they'll
have a buy it now on their website. So,
let me know your thoughts. Is Cerebras a
company you're interested in? I'm not
invested in them in any way. Uh they're
not a client of mine at this point.
Fingers crossed. Um but let me know what
you think about these guys, what you
want to see next, because they
definitely want to hear feedback from
the launch, from the event, from the
product. Uh cuz, uh they've got a team
that's, you know, going through all of
these, making sure the next generation
works better. So, please do leave
comments and questions for the CEO.
Maybe we'll get him on uh camera later
about what you want to hear about Cirrus
and what makes you think that they're
either good or bad um or what's coming
next.
Ask follow-up questions or revisit key timestamps.
This video details the launch of the Cerebras CS4 system, a high-performance AI platform built around the new WSE-3.5 wafer-scale engine. The system utilizes a novel 'backpack' design within a rack-mount configuration, significantly increasing performance, memory bandwidth, and scalability for large language model inference. The video also discusses Cerebras's market position, their partnership with companies like OpenAI, AMD, and AWS, and their strategy to differentiate themselves through raw speed rather than relying on traditional GPU ecosystems.
Videos recently processed by our community