HomeVideos

The Reason My Phone is Ringing Daily: Etched

Now Playing

The Reason My Phone is Ringing Daily: Etched

Transcript

333 segments

0:00

What do you do when a company coming out

0:04

of stealth says that they're ready to

0:06

take on Nvidia? They've got a chip that

0:09

only does Transformers, but it does

0:11

Transformers really well. It's headed up

0:15

by two Harvard computer science students

0:18

and it's got some betal backed money.

0:21

This company is etched and we finally

0:23

know a little bit about the chip they're

0:25

designing, so stick around to find out

0:28

more.

0:32

What's your minimum specification?

0:38

Now, Etched is one of those companies

0:40

that for the people who talk to me

0:42

regularly, they keep asking me about who

0:45

are they? What do they do? Why is there

0:48

so much interest in this startup based

0:51

on the people behind it and the amount

0:54

of money they've managed to raise? What

0:56

makes them so special against the

0:58

Nvidias, the AMDs, and everybody else?

1:02

Well, a couple of years ago, this

1:04

company etched, their CEO came out and

1:07

said, "We're building a chip for

1:09

transformers." Now, by that time,

1:11

transformers had kind of been embedded

1:13

in language models and is essentially

1:16

the mainstay going forward, but it's a

1:18

very specific workload. And the idea was

1:21

as they presented it is that if you

1:23

build a chip just for transformers then

1:26

it will be able to do transformers

1:27

really fast and really efficiently. The

1:31

only problem is if something else comes

1:33

along that chip might be useless. And

1:36

the CEO said as much. He said that if

1:38

the market pivots then what they've done

1:41

might be for nothing. That was a couple

1:44

of years ago and since then they raised

1:47

$800 million uh dollars in a series B

1:50

and then $300 million 3 months later or

1:53

3 weeks later in a series C. So they've

1:56

raised quite a bit of money and they

1:58

finally lifted the lid on what the

2:01

architecture or at least part of what

2:03

the architecture is around as it relates

2:05

to this new generation chip. Now first I

2:09

want to talk about the people behind the

2:11

company. Now the CEO of Etched is a lad

2:14

called Gavin Uberti who at the time the

2:17

company was formed was a 23-year-old

2:21

dropout of Harvard in the computer

2:23

science division. His co-founder is

2:26

Robert Wacken who of a similar age also

2:29

from Harvard but worked more on the

2:31

spinout side of uh the university

2:34

helping companies go from you know

2:37

nothing to something with money and

2:39

their CTO is a professor called Mark

2:42

Ross whose history actually includes

2:44

Cisco and Cypress Semi. It lists on his

2:47

LinkedIn that he was he built the team

2:50

and eventually built the first uh

2:52

multiport Ethernet switch in the market.

2:55

So, there's some history here. Also,

2:57

turns out he's worked a little bit on

2:58

the technology that's coming up in Soho.

3:02

Now, when again the company was

3:03

announced, there was a bit of a hoo-ha

3:05

coming out of nowhere saying, "We're

3:07

going to challenge Nvidia uh with an

3:09

eight-way box claiming 500,000 tokens

3:12

per second," which is uh many, many more

3:15

times what Nvidia was promoting at the

3:17

time. And it was around then that our

3:20

good friend Sally from EU Times actually

3:22

got to meet with a company and went

3:24

around the headquarters. And I remember

3:27

having discussions with her saying they

3:29

said that they had chips back and they

3:32

were um beeping and worring and she saw

3:35

a number of servers with flashing lights

3:37

but didn't actually see any silicon.

3:39

Fast forward a year, a year and a bit

3:42

and we now have the essentially the

3:44

company coming out of stealth properly.

3:46

First of all, with that fund raise and

3:48

now with another subsequent fund raise,

3:50

but also more details about this SOHU

3:52

chip that's apparently now taped out,

3:55

which kind of means that what Sally saw

3:56

was perhaps a test chip. But this tape

3:59

out, and we actually have some images of

4:01

the chip, involves two key technologies.

4:05

They've let us in uh gave us some small

4:07

amount of information about by which I

4:09

mean they've given us the names and

4:11

roughly how it works, but not quite

4:13

exact. And this is what people who have

4:15

been calling me up have been asking

4:17

about is this a legitimate workload. So

4:20

the two main technologies, one, the

4:22

first of which is low voltage compute,

4:25

which on the face of it sounds pretty

4:27

simple. When you have your compute

4:29

units, when you have your accelerators

4:31

in your AI silicon, you run them at a

4:35

frequency which allows them to switch

4:37

and switch fast. But if you have a

4:40

design that is embarrassingly parallel

4:43

and streamlined, kind of like a GPU,

4:46

then you can perhaps lower the voltage,

4:48

multiply how many units you have out.

4:51

And because the effect of lowering the

4:52

voltage voltage has a compound effect on

4:55

the frequency and power, you can be in a

4:58

more efficient regime.

5:00

The way that etched has described it

5:02

though in terms of low voltage inference

5:04

is that we're talking super low voltage.

5:07

Now, your processor in your PC, if you

5:09

got a high-end desktop at home, may be

5:11

running around 1.2, 1.3 volts, more if

5:15

you've overclocked it, even more if

5:17

you're still running a Pentium 4. The

5:20

most efficient desktop and laptop

5:22

processors, like the one I've got just

5:23

off to the side here, run more around

5:26

sort of 8.91 volts. And your GPU uh in

5:31

your system probably runs about that

5:33

similar level of voltage.

5:35

um these systems will go down to8 volts

5:38

when they're in idle. But there's a

5:40

class of processor that goes even below

5:43

that. And what you have to understand is

5:46

when you build a transistor, the minimum

5:48

voltage you need for that transistor to

5:50

switch on and off cuz that's what these

5:53

are switches is called your threshold

5:55

voltage.

5:57

Now, as you rise up through frequency,

6:00

the voltage you need to enable that

6:02

switching increases. That's why when you

6:05

need 5 GHz on your CPU, you need a lot

6:08

of voltage to do that. But these ultra

6:11

low voltage transistors run at around

6:14

450 molts,

6:17

0.45

6:18

volts, uh, which is really low. And the

6:21

reason why not a lot of people do it is

6:23

one, you have to be able to design it

6:25

and design them well. Two, you have to

6:27

also deal with the fact that they're not

6:29

running at super high multi- gigahertz

6:31

frequencies. And there's one industry,

6:35

one high compute industry that has

6:37

embraced this technology. It's called

6:39

Bitcoin mining. Turns out the Bitcoin

6:43

mining algorithm, the SHA 256 algorithm

6:46

or whatever it is, is ideally suited for

6:50

low voltage transistors that you can

6:52

make a trillion of. That's why if you go

6:55

to the main Bitcoin silicon providers,

6:58

uh I know Bitmain's a big one, they will

7:01

run the latest generation process node

7:03

technology TSMC 3 nanometer with ultra

7:07

low voltage transistors at around 400

7:10

450 molts, but then make a,000 chips

7:14

that run in a 3000 W system. The problem

7:17

there is Bitcoin mining is a very low

7:22

communication workload. You might have a

7:25

few bits of data go in, do your

7:27

calculation, and a few bits of data go

7:30

out. Machine learning isn't like that.

7:32

Machine learning is all about the data

7:34

flow, being able to transfer data to and

7:37

from everywhere. If you've got low

7:40

voltage compute that just deals with

7:42

small in, small out, that's great. What

7:45

about when you start have to transmit

7:47

that data across a chip or across a data

7:50

center? That's where the second

7:52

technology gets a bit more interesting.

7:54

The second technology etched introduced

7:57

is called cluster scale memory. So,

8:00

their chip, as we've seen, has HBM on

8:03

top um or HBM inside the package, but a

8:06

lot of what else is inside is SRAMM or

8:09

what we believe to be a large SRAMM

8:13

scratch pad that all the compute units

8:16

inside the chip can use. And the idea of

8:18

cluster scale memory is that that

8:21

scratch pad is available for other chips

8:23

to use through high-speed interfaces.

8:26

Now, high-speed interfaces require

8:29

voltage, so they're not part of the low

8:31

voltage inference we just mentioned. Um,

8:33

but it might mean that if you're using

8:35

less power on your compute, you can use

8:38

more power in your communications. So if

8:42

you have a scaleup fabric of a thousand

8:44

chips, more of the power going to the

8:47

communication pathway to this cluster

8:50

scale memory to this what looks like

8:52

it's um it's when you do RDMA uh where

8:57

you do a direct memory access to another

8:59

chip, but in this case it's an SRAMM

9:02

RDMA. You're actively accessing the

9:05

cache and the scratchpad of another chip

9:08

directly. That sounds like a really

9:10

interesting technology.

9:12

That being said, questions about

9:15

bandwidth, power consumption, which IP

9:18

is being used, uh, all of which

9:20

unanswered by etched. Uh, in this

9:22

instance, they're simply saying that

9:24

saying that in order to scale their

9:26

compute, in order to scale their low

9:28

voltage inference, to scale this

9:30

scratchpad SRAM, they're using this

9:32

cluster scale memory. Um they did show

9:36

off a couple of images of uh their

9:40

serverbased solutions. And the one thing

9:43

that most people saw when they uh

9:45

realized when they saw it the first time

9:46

is that's a lot of cabling. It's a lot

9:49

of black cabling going here, there, and

9:50

everywhere. A lot more than what you

9:52

might see with an Nvidia or an AMD

9:55

system, for example. And our first

9:57

thought is, well, that must be due to

9:59

all the connectivity needed for this

10:01

cluster scale memory. Um, some of it's

10:04

going to be for cooling. These are going

10:05

to be liquid cool chips when they're

10:07

deployed,

10:08

but a lot of them could be for this high

10:11

bandwidth interface. In reality, maybe

10:14

what etched have solved here isn't so

10:18

much the compute problem, but maybe a

10:20

bandwidth problem. The thing you have to

10:22

realize is when they came out and said,

10:24

"Hey, we do low voltage inference and

10:26

this cluster scale memory."

10:29

nothing specifically in those two mean

10:32

transformers. So that must mean that the

10:35

the architecture or the way that these

10:38

compute units are laid out through low

10:40

voltage inference is directly mapped to

10:44

transistor workload uh to transformer

10:46

workloads. In that regard, it means that

10:51

what we end up with with is a hardwired

10:53

transformer.

10:55

Still programmable because you still

10:56

have to be able to run different sorts

10:58

of transformers. So, it's not as rigid

11:00

as the Talis chip we saw a few weeks

11:02

ago, but it still has to be transformer

11:05

based. Um, I wanted to catch up with the

11:08

team and I haven't had a chance to have

11:09

a really good on thereord chat with them

11:12

because things like transcendental

11:13

functions still need to be solved. how

11:16

the memory addressing system works

11:18

because if there's no memory maps unit,

11:21

then that might make things complicated.

11:23

Or the fact that these are scratch pads,

11:25

how do you keep track of what's in

11:26

memory and what's out? Or maybe that all

11:29

has to be done by the compiler. And in

11:31

that instance, we really have to pray to

11:33

the compiler gods that their compilers

11:36

work. Um, but Sally has been back to the

11:39

headquarters and apparently they do have

11:41

lots of a lot more lights flashing these

11:44

days um with their new uh silicon chip

11:46

that's been taped out. I expect that we

11:49

will hear more from them later this

11:51

year. Um, since George and I recorded

11:54

our Hot Chips preview video, which you

11:58

can find link here, um, they have now

12:01

came out as one of the top sponsors for

12:04

that event. yet they're not listed as a

12:06

talk just yet. So, we might see

12:10

something at the event or hear something

12:12

at the event that may not be part of the

12:14

official. Who knows? All I know is we

12:17

want to hear more and it sounds like the

12:19

team is ready to at least talk more. The

12:22

other angle to this is their customer

12:26

base. They're already saying that

12:27

they've got a lot of interest from a

12:29

number of hyperscalers and players in

12:31

the market. Um, though I'll refer to the

12:34

old Jim Keller adage of, well, if you

12:36

want to sell a million chips, you got to

12:37

be able to sell 100,000. If you want to

12:39

sell 100,000, you got to be able to sell

12:41

10,000. If you want to sell 10,000, you

12:43

got to be able to sell a,000. And if you

12:44

want to sell a,000, you got to be able

12:46

to sell a 100. So, the question with

12:48

always with these startups is where are

12:50

your hundreds? Where are your thousands?

12:52

Where are your 10,000s? Um Robert Wacken

12:55

the co-founder he did say at raise on

12:58

stage that the key ethos at etched is

13:01

deliver deliver

13:04

as in their internal metrics. It was all

13:07

about delivering a chip a year earlier

13:10

than the startups who basically formed

13:13

around about the same time. and they're

13:15

going to keep on this aggressive cadence

13:17

of maybe not quite year on year but at

13:20

least as soon as they can get their next

13:22

generation chips out that they will do

13:24

that. I'm hoping that if that is the

13:26

case then we might see a road map

13:28

hopefully by the end of the year talking

13:30

about Gen 2, Gen 3 and how exactly this

13:34

low voltage inference architecture

13:37

evolves over time. So even though they

13:42

are a couple of guys from Harvard, even

13:44

though they have what sounds like Peter

13:47

Teal money, they are making a lot of

13:49

waves very quickly and a lot of noise

13:52

that everybody's asking me about. So the

13:54

question is when I do go see them, what

13:56

do you want to know?

13:59

Now what I do know is you really should

14:01

head to shop.teato.com

14:03

to get your TechT techato merch. It's

14:06

kind of cool.

Interactive Summary

The video discusses Etched, a startup founded by former Harvard students, aiming to compete with Nvidia by creating chips specifically optimized for transformer workloads. The company has raised significant capital and is moving out of stealth mode with a new chip, Sohu, which utilizes 'low voltage compute' and 'cluster scale memory' technologies. Despite the hype, questions remain regarding their ability to scale production and the long-term flexibility of their hardwired architecture.

Suggested questions

3 ready-made prompts