The Reason My Phone is Ringing Daily: Etched
333 segments
What do you do when a company coming out
of stealth says that they're ready to
take on Nvidia? They've got a chip that
only does Transformers, but it does
Transformers really well. It's headed up
by two Harvard computer science students
and it's got some betal backed money.
This company is etched and we finally
know a little bit about the chip they're
designing, so stick around to find out
more.
What's your minimum specification?
Now, Etched is one of those companies
that for the people who talk to me
regularly, they keep asking me about who
are they? What do they do? Why is there
so much interest in this startup based
on the people behind it and the amount
of money they've managed to raise? What
makes them so special against the
Nvidias, the AMDs, and everybody else?
Well, a couple of years ago, this
company etched, their CEO came out and
said, "We're building a chip for
transformers." Now, by that time,
transformers had kind of been embedded
in language models and is essentially
the mainstay going forward, but it's a
very specific workload. And the idea was
as they presented it is that if you
build a chip just for transformers then
it will be able to do transformers
really fast and really efficiently. The
only problem is if something else comes
along that chip might be useless. And
the CEO said as much. He said that if
the market pivots then what they've done
might be for nothing. That was a couple
of years ago and since then they raised
$800 million uh dollars in a series B
and then $300 million 3 months later or
3 weeks later in a series C. So they've
raised quite a bit of money and they
finally lifted the lid on what the
architecture or at least part of what
the architecture is around as it relates
to this new generation chip. Now first I
want to talk about the people behind the
company. Now the CEO of Etched is a lad
called Gavin Uberti who at the time the
company was formed was a 23-year-old
dropout of Harvard in the computer
science division. His co-founder is
Robert Wacken who of a similar age also
from Harvard but worked more on the
spinout side of uh the university
helping companies go from you know
nothing to something with money and
their CTO is a professor called Mark
Ross whose history actually includes
Cisco and Cypress Semi. It lists on his
LinkedIn that he was he built the team
and eventually built the first uh
multiport Ethernet switch in the market.
So, there's some history here. Also,
turns out he's worked a little bit on
the technology that's coming up in Soho.
Now, when again the company was
announced, there was a bit of a hoo-ha
coming out of nowhere saying, "We're
going to challenge Nvidia uh with an
eight-way box claiming 500,000 tokens
per second," which is uh many, many more
times what Nvidia was promoting at the
time. And it was around then that our
good friend Sally from EU Times actually
got to meet with a company and went
around the headquarters. And I remember
having discussions with her saying they
said that they had chips back and they
were um beeping and worring and she saw
a number of servers with flashing lights
but didn't actually see any silicon.
Fast forward a year, a year and a bit
and we now have the essentially the
company coming out of stealth properly.
First of all, with that fund raise and
now with another subsequent fund raise,
but also more details about this SOHU
chip that's apparently now taped out,
which kind of means that what Sally saw
was perhaps a test chip. But this tape
out, and we actually have some images of
the chip, involves two key technologies.
They've let us in uh gave us some small
amount of information about by which I
mean they've given us the names and
roughly how it works, but not quite
exact. And this is what people who have
been calling me up have been asking
about is this a legitimate workload. So
the two main technologies, one, the
first of which is low voltage compute,
which on the face of it sounds pretty
simple. When you have your compute
units, when you have your accelerators
in your AI silicon, you run them at a
frequency which allows them to switch
and switch fast. But if you have a
design that is embarrassingly parallel
and streamlined, kind of like a GPU,
then you can perhaps lower the voltage,
multiply how many units you have out.
And because the effect of lowering the
voltage voltage has a compound effect on
the frequency and power, you can be in a
more efficient regime.
The way that etched has described it
though in terms of low voltage inference
is that we're talking super low voltage.
Now, your processor in your PC, if you
got a high-end desktop at home, may be
running around 1.2, 1.3 volts, more if
you've overclocked it, even more if
you're still running a Pentium 4. The
most efficient desktop and laptop
processors, like the one I've got just
off to the side here, run more around
sort of 8.91 volts. And your GPU uh in
your system probably runs about that
similar level of voltage.
um these systems will go down to8 volts
when they're in idle. But there's a
class of processor that goes even below
that. And what you have to understand is
when you build a transistor, the minimum
voltage you need for that transistor to
switch on and off cuz that's what these
are switches is called your threshold
voltage.
Now, as you rise up through frequency,
the voltage you need to enable that
switching increases. That's why when you
need 5 GHz on your CPU, you need a lot
of voltage to do that. But these ultra
low voltage transistors run at around
450 molts,
0.45
volts, uh, which is really low. And the
reason why not a lot of people do it is
one, you have to be able to design it
and design them well. Two, you have to
also deal with the fact that they're not
running at super high multi- gigahertz
frequencies. And there's one industry,
one high compute industry that has
embraced this technology. It's called
Bitcoin mining. Turns out the Bitcoin
mining algorithm, the SHA 256 algorithm
or whatever it is, is ideally suited for
low voltage transistors that you can
make a trillion of. That's why if you go
to the main Bitcoin silicon providers,
uh I know Bitmain's a big one, they will
run the latest generation process node
technology TSMC 3 nanometer with ultra
low voltage transistors at around 400
450 molts, but then make a,000 chips
that run in a 3000 W system. The problem
there is Bitcoin mining is a very low
communication workload. You might have a
few bits of data go in, do your
calculation, and a few bits of data go
out. Machine learning isn't like that.
Machine learning is all about the data
flow, being able to transfer data to and
from everywhere. If you've got low
voltage compute that just deals with
small in, small out, that's great. What
about when you start have to transmit
that data across a chip or across a data
center? That's where the second
technology gets a bit more interesting.
The second technology etched introduced
is called cluster scale memory. So,
their chip, as we've seen, has HBM on
top um or HBM inside the package, but a
lot of what else is inside is SRAMM or
what we believe to be a large SRAMM
scratch pad that all the compute units
inside the chip can use. And the idea of
cluster scale memory is that that
scratch pad is available for other chips
to use through high-speed interfaces.
Now, high-speed interfaces require
voltage, so they're not part of the low
voltage inference we just mentioned. Um,
but it might mean that if you're using
less power on your compute, you can use
more power in your communications. So if
you have a scaleup fabric of a thousand
chips, more of the power going to the
communication pathway to this cluster
scale memory to this what looks like
it's um it's when you do RDMA uh where
you do a direct memory access to another
chip, but in this case it's an SRAMM
RDMA. You're actively accessing the
cache and the scratchpad of another chip
directly. That sounds like a really
interesting technology.
That being said, questions about
bandwidth, power consumption, which IP
is being used, uh, all of which
unanswered by etched. Uh, in this
instance, they're simply saying that
saying that in order to scale their
compute, in order to scale their low
voltage inference, to scale this
scratchpad SRAM, they're using this
cluster scale memory. Um they did show
off a couple of images of uh their
serverbased solutions. And the one thing
that most people saw when they uh
realized when they saw it the first time
is that's a lot of cabling. It's a lot
of black cabling going here, there, and
everywhere. A lot more than what you
might see with an Nvidia or an AMD
system, for example. And our first
thought is, well, that must be due to
all the connectivity needed for this
cluster scale memory. Um, some of it's
going to be for cooling. These are going
to be liquid cool chips when they're
deployed,
but a lot of them could be for this high
bandwidth interface. In reality, maybe
what etched have solved here isn't so
much the compute problem, but maybe a
bandwidth problem. The thing you have to
realize is when they came out and said,
"Hey, we do low voltage inference and
this cluster scale memory."
nothing specifically in those two mean
transformers. So that must mean that the
the architecture or the way that these
compute units are laid out through low
voltage inference is directly mapped to
transistor workload uh to transformer
workloads. In that regard, it means that
what we end up with with is a hardwired
transformer.
Still programmable because you still
have to be able to run different sorts
of transformers. So, it's not as rigid
as the Talis chip we saw a few weeks
ago, but it still has to be transformer
based. Um, I wanted to catch up with the
team and I haven't had a chance to have
a really good on thereord chat with them
because things like transcendental
functions still need to be solved. how
the memory addressing system works
because if there's no memory maps unit,
then that might make things complicated.
Or the fact that these are scratch pads,
how do you keep track of what's in
memory and what's out? Or maybe that all
has to be done by the compiler. And in
that instance, we really have to pray to
the compiler gods that their compilers
work. Um, but Sally has been back to the
headquarters and apparently they do have
lots of a lot more lights flashing these
days um with their new uh silicon chip
that's been taped out. I expect that we
will hear more from them later this
year. Um, since George and I recorded
our Hot Chips preview video, which you
can find link here, um, they have now
came out as one of the top sponsors for
that event. yet they're not listed as a
talk just yet. So, we might see
something at the event or hear something
at the event that may not be part of the
official. Who knows? All I know is we
want to hear more and it sounds like the
team is ready to at least talk more. The
other angle to this is their customer
base. They're already saying that
they've got a lot of interest from a
number of hyperscalers and players in
the market. Um, though I'll refer to the
old Jim Keller adage of, well, if you
want to sell a million chips, you got to
be able to sell 100,000. If you want to
sell 100,000, you got to be able to sell
10,000. If you want to sell 10,000, you
got to be able to sell a,000. And if you
want to sell a,000, you got to be able
to sell a 100. So, the question with
always with these startups is where are
your hundreds? Where are your thousands?
Where are your 10,000s? Um Robert Wacken
the co-founder he did say at raise on
stage that the key ethos at etched is
deliver deliver
as in their internal metrics. It was all
about delivering a chip a year earlier
than the startups who basically formed
around about the same time. and they're
going to keep on this aggressive cadence
of maybe not quite year on year but at
least as soon as they can get their next
generation chips out that they will do
that. I'm hoping that if that is the
case then we might see a road map
hopefully by the end of the year talking
about Gen 2, Gen 3 and how exactly this
low voltage inference architecture
evolves over time. So even though they
are a couple of guys from Harvard, even
though they have what sounds like Peter
Teal money, they are making a lot of
waves very quickly and a lot of noise
that everybody's asking me about. So the
question is when I do go see them, what
do you want to know?
Now what I do know is you really should
head to shop.teato.com
to get your TechT techato merch. It's
kind of cool.
Ask follow-up questions or revisit key timestamps.
The video discusses Etched, a startup founded by former Harvard students, aiming to compete with Nvidia by creating chips specifically optimized for transformer workloads. The company has raised significant capital and is moving out of stealth mode with a new chip, Sohu, which utilizes 'low voltage compute' and 'cluster scale memory' technologies. Despite the hype, questions remain regarding their ability to scale production and the long-term flexibility of their hardwired architecture.
Videos recently processed by our community