HomeVideos

Jane Street on GPUs, Trading, and Hiring: A Conversation with Dwarkesh

Now Playing

Jane Street on GPUs, Trading, and Hiring: A Conversation with Dwarkesh

Transcript

905 segments

0:00

Jane Street are partners of my podcast

0:02

and one of the fun ideas we had is why

0:04

don't I come visit a data center uh for

0:08

training that you guys run. So I just

0:10

got a tour of this uh Texas data center

0:13

from Ron Minsky who co-heads the

0:15

technology group and Dan Ponttovo who

0:18

heads the physical engineering team. So

0:20

thank you guys for showing me around

0:23

it's worth I've never been here before

0:24

so I also got I was also getting a tour

0:26

which was great. Previously I was

0:27

confused well how can you be doing GPU

0:29

things if you need to be trading on

0:30

nanconds and maybe you can talk through

0:33

what is the actual time horizon of the

0:37

trading you guys do can you afford to

0:39

have be running big models to in in the

0:42

middle of making trading decisions

0:44

>> I think the thing to understand here is

0:45

there isn't one time horizon there are

0:47

many time horizons uh there are trading

0:49

systems we build and trades that we do

0:52

where in order to be competitive you

0:54

actually have to turn around a packet in

0:56

in under 100 nanconds and like that's a

0:59

very different regime, right? You know,

1:01

people sometimes talk about like, oh,

1:03

can you guys write high performance

1:05

stuff in Okamel? It's like we can, but

1:07

like for this kind of speed, it's like

1:09

it doesn't matter if you write in Okamel

1:10

or Rust or C++. You can't use a CPU,

1:13

right? you are going to be on an FPGA

1:15

that's like direct wire attached to the

1:16

network and you're going to be turning

1:18

around the packet so fast that if you

1:19

like attached an oscilloscope to the

1:21

wire on the way in and the wire on the

1:22

way out you would see the packet start

1:24

to leave before it's done being

1:26

consumed. So it's like a very different

1:28

very specialized regime but like when

1:30

you're in that time regime you really

1:32

can't do very much computation. the

1:34

decisions you're making are going to be

1:35

very simple. And in fact, there's this

1:37

kind of whole curve of trade-offs

1:39

between how smart is the decision that

1:42

you're making, be it a model or some

1:44

other kind of maybe even like

1:45

handwritten decision-m process, and how

1:48

fast the turnaround is. And like the the

1:52

right way to build uh an optimal trading

1:54

strategy is really to have a kind of

1:56

ensemble approach where for some kinds

1:57

of decisions you're making very s simple

1:59

decisions very quickly. For some kind of

2:01

decisions, you're operating at the scale

2:03

of, you know, instead of thinking of a

2:05

hundred nanos, maybe like a handful of

2:08

mics or tens of microsconds or hundreds

2:10

of microscs or milliseconds. And in some

2:13

cases, there are processes where if you

2:14

can get that decision turned around, you

2:16

know, in an hour or that day, that's

2:19

totally fine. And you're you're kind of

2:22

competitive on a time basis at each of

2:25

these horizons. Uh but you're making

2:27

very different kinds of decisions at all

2:29

of them. Maybe you can't say but what

2:31

what is it exactly these models are

2:33

predicting like surely it's just not the

2:35

next thing in the order book or maybe it

2:36

is

2:36

>> right so we're definitely like dancing

2:38

towards stuff that's hard to talk about

2:39

but I think the simplest and most

2:41

important one that we've been thinking

2:43

about like we think about it now but

2:45

like also 25 years ago when I started at

2:47

James Street when I was building like

2:48

models out of linear regression you know

2:51

and stuff like that like a very useful

2:53

kind of thing is to predict a fair value

2:55

for a thing like what do we think this

2:56

thing is worth and that fits in in a

2:58

very kind of composable a into lots of

3:00

different trading processes. That's not

3:02

the only kind of thing that we use as a

3:04

prediction target, but it's an important

3:05

one.

3:06

>> It seemed like a meme I was getting for

3:07

a while about like what trading firms do

3:09

is like you you got to get the colo and

3:10

the where the NASDAQ exchanges and it's

3:13

very important that your machines are

3:14

right there without

3:15

>> thinking without thinking too much about

3:16

the exact details of what we put where.

3:18

like your inference processes might be

3:20

on CPU, might be on FPGA, might be on

3:23

GPU

3:24

>> depending on the kind of constraints of

3:27

how much compute you need, how big the

3:28

model is, what kind of latency

3:29

turnaround do you need. And yeah, like

3:31

bigger, slower things you can put

3:33

farther away. It's annoying to have to

3:35

put all the compute right by the

3:36

exchange. And for the stuff that's like

3:38

really really fast, like being in the

3:40

colo isn't enough, you care about like

3:42

how long is the spool of wire that gets

3:45

you there. like you're literally like

3:46

measuring out the length of the fiber

3:48

runs when you're when again when you're

3:50

like at this very very low nanocond

3:52

scale. Um but in general like the like

3:55

bigger models give you a lot more

3:57

flexibility in terms of where they

3:58

physically go. If we're putting GPUs in

4:00

some of these colloccated uh facilities

4:02

that are next to the exchanges right

4:03

now, you have to work with with their

4:05

rules. You know, who is who is that

4:06

provider? Who is who is giving you that

4:08

space? Yeah. And your power, your

4:10

cooling, all those constraints now are

4:12

maybe slightly tighter than if you have

4:13

a facility facility that you're

4:15

designing and operating. Um so you're

4:17

now having to kind of come up with ways

4:19

to, hey, maybe I could only get one GPU

4:21

in a rack because it consumes so much

4:22

power. So now I have to spread it all

4:23

out rather than being able to do liquid

4:25

cooled in one rack. So these are all uh

4:27

things we need to keep in mind as our

4:29

comput you know comput.

4:32

>> You guys recently signed a $6 billion

4:34

comput deal with core reef.

4:35

>> Mhm.

4:36

>> What are you going to use that for?

4:37

>> The rest of the AI world has scaling

4:38

laws. We have scaling laws too and there

4:41

are lots of models that we want to

4:42

train. I think the thing that's

4:44

interesting and maybe different between

4:46

us and the kind of more traditional AI

4:49

labs is the amount of diversity in model

4:54

architecture and the amount of

4:55

experimentation that we're doing. So a

4:57

lot of the value you get from all of

4:58

this is just people are like trying lots

5:00

of very different new things in the

5:02

model designs and giving researchers

5:05

just like faster iteration time so they

5:07

can discover more ideas and drive more

5:09

innovation. It just turns out to be

5:10

incredibly important.

5:11

>> In the case of these foundation labs,

5:14

there's some gain from have training

5:15

just one model that does everything that

5:16

is fully general rather than building a

5:19

bunch of custom different models. Can

5:21

you give me a sense of why there's a

5:23

different trade-off at Jane Street? For

5:25

us, some of the specialization is about

5:27

being adapted to consume the right kind

5:29

of data, right? And there are just like

5:31

many possible data sources that we might

5:32

be feeding in. There are like a bunch of

5:35

just differences in the data rates that

5:38

we need to achieve. Like just like

5:40

another thing that like just makes us

5:41

need to kind of specialize some of what

5:43

we're doing is just like the overall

5:44

kind of both inference and trading

5:46

dynamics are made different by just like

5:48

the the bytes to flop ratio being

5:50

different. We have like way more data

5:52

that we are using to train the models,

5:54

but the data is kind of bite for bite

5:56

less informative just because financial

5:58

data is very noisy.

6:00

>> Yeah. Um, and so the models tend to be

6:03

smaller and the data tends to be noisier

6:05

noisier and there tends to be a lot more

6:06

of it. And it's also different between

6:08

different models that we build for

6:09

different applications, right? As we try

6:11

and figure out like how can we leverage

6:13

more of the information that we get.

6:15

It's like oh now there's like all of the

6:17

kind of decisions from like how do we

6:19

store and load data efficiently to how

6:21

do we shape the model to how do we make

6:23

the inference process, you know, have

6:25

both the throughput and latency that it

6:27

needs. those there's going to be a whole

6:28

different set of trade-offs there. And

6:30

so there's just like a lot of value in

6:33

kind of working that out and picking the

6:35

the best thing that you can do for for

6:38

different applications.

6:39

>> What is the um inference workload

6:41

actually or how does it compare to what

6:44

your traditional big chatbot LLM company

6:47

is doing

6:48

>> in broadstrokes? Latency matters more as

6:50

you might expect. Um, batching is still

6:53

an issue. Like depending on the model

6:54

you're doing, you might have models or a

6:56

part of models that are kind of

6:57

disagregated for different symbols that

6:59

you're looking at. And so the same kind

7:00

of like pulling in data from multiple

7:02

sources and batching them together makes

7:04

a difference. I think another thing

7:05

that's interesting is just like the data

7:07

rates are really high. Like the amount

7:10

the the aggregate data rate that you get

7:12

in a large LLM LLM lab from like all of

7:15

the different users is also very high.

7:17

But the amount of sequential data that

7:19

you're going to get from any one user is

7:21

not that high. Whereas when you the the

7:24

data that you're pulling is the bytes

7:25

that are coming out of like the NASDAQ

7:27

feed, it's like oh man the data rate

7:29

that you want that kind of sequentially

7:31

consumed in one domain kind of causally

7:34

one after the other is really high. And

7:36

so again like the dynamics change and

7:39

like but I think a lot of the same kind

7:41

of basic engineering questions are not

7:44

so dissimilar but like all the constants

7:45

are twiddled to different places and so

7:47

you end up making different choices.

7:49

>> What what does that mean in terms of the

7:51

how you how you had to design these

7:52

systems where whether it's in terms of

7:54

storage or whatever else.

7:56

>> Yeah, there's I think more emphasis on

7:57

the performance of the data loading than

7:59

you might otherwise see. I think we're

8:02

doing a lot of work to build out our own

8:04

kind of largecale uh data storage

8:07

system, our own kind of internal object

8:08

store um where we've like used various

8:10

kind of vendor products but um over time

8:13

I think for some of these research

8:14

focused use cases we kind of need to

8:16

operate at a much larger scale and need

8:19

also to deal with a diversity of data

8:21

centers right and this is like less a

8:23

training time less a inference time and

8:24

more of a training time question of like

8:27

we just can't get all the compute we

8:28

want all in the same place And I don't

8:31

know, I feel like in general like a an

8:34

important trick in like effectively

8:37

running a technical organization is

8:38

feeling is figuring out what shortcuts

8:40

you can take. One shortcut that we were

8:42

like privileged to be able to take for

8:43

many years is we got to pretend like

8:45

there was only one CPU architecture on

8:47

the planet. Like everything was for like

8:48

x8664.

8:50

We pretended like none of these other

8:51

things existed. Uh and that simplified a

8:54

bunch of things. Uh and we also had like

8:56

one big research data center and one big

8:58

storage cluster and that also simplified

9:00

a ton of things. And actually both of

9:01

those have now been unwound like you

9:04

just can't get the amount of power like

9:06

you cannot wire in enough thunderbolts

9:08

into like the same data center to power

9:09

all the things you need. You need to get

9:11

the data centers built all over the

9:12

place. So there's a big disagregation

9:14

problem and that gives you a problem

9:15

like oh now you have to think about like

9:17

your compute scheduling and your storage

9:19

scheduling being intertwined one with

9:21

the other and there's a ton of data. So

9:23

moving around is like non-trivial. Um,

9:25

and also we had to give up on this x86

9:27

only thing because Nvidia has a bunch of

9:30

cool new products that mean that you

9:32

need to support ARM. Now zooming out, I

9:34

want to ask a very naive question.

9:35

There's maybe a naive view that uh, you

9:39

know, if you have AGI, it can like

9:41

immediately do what Jane Street does.

9:44

Give me a sense of like why that naive

9:49

view is naive.

9:50

>> Yeah. And I don't want to totally

9:51

discount it like you know there's a

9:54

world that we should take seriously

9:55

where like you know we're going to build

9:58

large language models or some other AI

10:00

systems that are like strictly smarter

10:02

than all humans on the planet and more

10:04

capable at all cognitive tasks and like

10:08

>> yeah that's going to be weird and that's

10:09

like a different that's a that's a

10:10

different a different state of things.

10:12

Um, and in that case, yeah, you know,

10:15

maybe large amounts of things that Jane

10:16

Street does will be automated away and,

10:19

you know, maybe we'll all just like, you

10:20

know, sit back and, you know, drink more

10:22

margaritas or something. I don't know

10:23

what that world looks like, but it

10:25

doesn't feel like we're particularly

10:27

close to that now. I think that like in

10:29

general I think it's like easy to

10:31

underestimate the richness and

10:33

complexity of the work both that like a

10:36

company like Jane Street does but really

10:37

that is done in kind of any really like

10:39

ambitious high difficulty like company

10:42

scale task. I think trading in

10:45

particular feels to me as like kind of

10:47

AGI complete sort of like NPcomplete.

10:50

It's like meaning like that like all of

10:52

the different problems of the world end

10:54

up influencing what you're doing in a

10:56

trading context because at the end of

10:58

the day trading involves figuring out

11:02

what things are worth which means making

11:04

predictions about the future and lots of

11:06

different things flow into that and as

11:08

various pieces of that get automated you

11:10

know you have the usual thing of like

11:12

the other hard parts that we don't yet

11:13

know how to automate well that ends up

11:15

being where the competitive edge lies I

11:17

feel like humans and like human

11:19

cognition are like more valuable than

11:21

ever. Like I have never been more

11:22

desperate to hire more engineers and

11:24

more traders than I am today because

11:26

everything people are doing is more

11:28

valuable than it was. I mean some of

11:30

this is just me being somewhat skeptical

11:33

that we are quite as close to the models

11:36

that are like smarter than humans at all

11:38

the things as some people seem to think.

11:40

Maybe it's like physical infrastructure

11:42

like actually getting the colo. Maybe

11:43

it's actually like the software

11:44

infrastructure that you build. Like give

11:46

me a sense of what it is that would

11:48

>> Yeah, we build like a huge variety of

11:50

complicated pieces of software, have

11:52

people thinking about lots of different

11:54

trading problems, some of which are not

11:55

very electronic at all. Like the

11:57

business is just like way more diverse

11:59

than I think people give it credit for.

12:01

And there's an idea of like, oh yeah,

12:02

it's like it must be that like simple

12:04

thing where you just like you just have

12:06

smart people who like make smart

12:08

decisions and write good software. And

12:09

like if we could just automate the

12:10

smartness part, that would be the whole

12:13

thing. And I think it's just way more

12:14

complicated than that. What what do you

12:15

mean by the non electronic parts of

12:17

trading?

12:18

>> I mean, there's still trading that

12:19

happens

12:21

via chat between people talking to each

12:23

other and making decisions and like

12:25

someone like sizing up how much adverse

12:27

selection they think the person on the

12:29

other side of the phone represents.

12:30

That's like still a real part of the

12:32

business. Um there's just like, you

12:35

know, there's just different kinds of

12:37

securities that have taken longer to get

12:40

more automated. The bonds business for

12:42

example is just like not nearly at the

12:45

level of automation that you see in

12:48

equities. Indeed, we I think we were

12:50

kind of confused about this of like I

12:52

think those of us who have been like in

12:53

the business for a while. We kind of I

12:55

mean I I started a little too late to

12:56

really see the kind of transition of

12:58

equities becoming electronic. But I

13:00

think people who are you know paying

13:01

attention a little earlier than me were

13:02

like yeah and I guess everything else

13:03

comes next. And like you know what it's

13:05

been like you know 25 30 years and like

13:07

not everything has gone that way. the

13:09

systems are still, you know, we don't

13:11

have a lot of people like standing on

13:12

the floor of exchanges anymore, but

13:14

there's still lots of trading that is

13:16

deeply intermediated by humans and human

13:18

judgment.

13:18

>> King of which, how much are humans in

13:20

the loop on between the model and the

13:22

and the trading decision? Many of your

13:24

most profitable days happen when like

13:27

weird stuff happens and there are events

13:29

and the world kind of goes crazy and

13:30

like nobody knows what's going on and

13:32

like that's when it's like very hard to

13:35

provide liquidity in those contexts and

13:37

so you get paid more for doing it and

13:39

there's often a lot of volume on days

13:40

like that

13:41

>> and doing that well often involves human

13:44

judgment of like thinking about like how

13:47

is today different from all of the other

13:49

days and you know to the degree that we

13:51

can we want to build the models that

13:53

work well through phase transitions but

13:56

also we think humans work better than

13:57

models do through phase transitions and

13:59

sometimes you need this kind of meta

14:01

judgment to decide what to do and so

14:04

there's a even for the systems that are

14:06

largely automated there are decisions to

14:08

be made by the people who are watching

14:10

and we always have people who are

14:12

watching right I think an important part

14:13

of trading is paying attention to and

14:15

thinking about what's happening during

14:17

the trading day even if the individual

14:19

transactions are going by far too fast

14:21

for a human to kind of weigh in on a

14:23

kind of transaction bytransaction basis.

14:26

>> Dan, what what have been the more

14:27

notable changes over the last 20 years

14:30

that you've been doing in buildings like

14:32

these?

14:33

>> Yeah, people are actually care about

14:34

data centers and want to talk about it.

14:36

You know, been working on cooling for a

14:38

while and now all a sudden people people

14:40

talk about it and and and think it's

14:41

interesting. So that's like that's fun

14:43

and exciting. And for for folks on my

14:45

team, I think they feel that way as

14:46

well. There's people who have been in

14:48

the data center industry for 20 years

14:50

that kind of still want to do it the way

14:52

they used to. And I think that's kind of

14:53

falling by the wayside now. Uh you're

14:55

finding ways where people are um

14:58

challenging previous thoughts. Hey,

15:00

these my entire data center is backed up

15:01

by generators. But generators are some

15:03

of the longest lead time items you can

15:05

buy. So maybe we take those away and

15:07

only put it for a core part of the

15:09

system that needs that resiliency. Um

15:11

that gets our GPUs on six months faster.

15:13

Let's do it. So those are things that uh

15:15

you know maybe Maybe it's not the best

15:17

engineering decision, but it's truly the

15:19

best business decision. And I think it's

15:22

stuff like that that has been coming up

15:23

more and more often.

15:24

>> It feels like every year people change

15:26

the what what is bottlenecking scaling

15:28

AI compute right now as you're doing

15:32

more negotiations and trying to acquire

15:33

more comput. What what is the current

15:34

bottleneck and what do you expect it to

15:36

be for

15:36

>> putting aside comput and memory and all

15:38

that fun stuff. So generators, uh

15:40

transformers, um some of the cooling

15:41

equipment that's used now for the liquid

15:43

cooling is is is in in a lot of demand.

15:45

So um and it changes rapidly. What I

15:47

tell you today is is going definitely

15:49

going to be different 2 weeks from now.

15:51

Um we do this thing we work very closely

15:54

with internal teams on the procurement

15:56

side uh to to stock up on some of this

15:58

stuff. Stuff that we know is fungeible

15:59

across all our data centers. We will

16:01

warehouse and have it ready to go. Um,

16:03

there's components like generators where

16:05

you're not going to put a giant

16:06

generator in a in a warehouse or or you

16:09

know, for instance, if you're doing

16:10

something behind the meter like a

16:11

turbine, you're gonna you're going to

16:13

have to think about those markets a

16:15

little bit more. Um, where you're

16:17

getting them, where you're staging them,

16:18

you can't just leave them off to the

16:20

side. Um, so I think the components

16:23

definitely change. Those are some of the

16:24

big ones. And uh you know as we get to

16:27

more and more density you know I think

16:30

one hope is that the buildings get a

16:32

little bit smaller and maybe and and and

16:33

we're able to like you know build the

16:35

buildings faster get all that compute

16:37

kind of in a nice tight bundle and then

16:39

all the infrastructure around it's got

16:40

to got to be maybe pre-built and

16:42

delivered to site right modular data

16:44

centers or modular infrastructure is

16:46

becoming more and more of a thing where

16:48

these components especially the long

16:49

lead components are being um designed

16:52

and built offsite and shipped to site.

16:54

So almost as close to plug-and-play as

16:57

you can get.

16:57

>> Well, one of the points you made earlier

16:59

is that as uh as the racks themselves

17:03

get more uh dense uh you know more and

17:06

more of the data center is like the

17:07

infra around

17:09

>> the actual racks which actually is kind

17:10

of similar to um like a a a package on

17:16

like a a chip right or like a chip on a

17:19

package. It's like the the compute is a

17:20

very small part of the total area of

17:22

package.

17:23

>> Yeah, it's it's interesting. thing. I

17:24

mean I I don't you know um it it doesn't

17:27

solve any problems per se. I mean maybe

17:29

it creates it creates others. Sure. Like

17:31

you know you get to a one megawatt rack,

17:32

right? People are like what does that

17:34

even mean? One megawatt in a rack and

17:35

and you know you know the the cooling

17:37

kind of the pipe is just going to get

17:39

larger that you're bringing there and uh

17:40

the amount of power whether it's kind of

17:42

the AC power that we're using now or 800

17:44

volt DC where it's where it's going in

17:46

the future. You still have to bring all

17:47

that those components to a spot. And the

17:49

thing that's like interesting from our

17:51

point of view is like you know we could

17:53

design these these engineering things

17:55

but at the end of the day whether it's

17:57

Nvidia or an ASIC or who they have to

18:00

sell a component that can work in a data

18:02

center and they're they're thinking very

18:04

hard about what they sell um because you

18:07

need people to use it right if you're if

18:09

you build a one megawatt data center um

18:11

one megawatt rack but there's no way to

18:13

power and cool it kind of useless. So,

18:15

you know, we're working very closely

18:17

with with kind of almost everyone in in

18:19

that space to think about what are the

18:21

components you need to be able to

18:22

support these next generations

18:24

>> because the lead times you're talking

18:25

about, you know, over a year sometimes

18:27

and you're just you're deciding on the

18:29

infrastructure before you're placing an

18:30

order for the chips.

18:32

>> So, you know, you're trying, for

18:34

instance, the you know, TPUs, they use

18:36

lower temperature water and they're

18:37

they're half as dense as as you know, an

18:39

NBL72 GP300, right? So that requires a

18:43

different strategy and and you want to

18:45

make sure you can handle those in the

18:46

future.

18:47

>> One of the things that allows

18:48

hyperscalers to commit to large amounts

18:50

of compute is that they have some

18:53

reserve use for excess compute that

18:55

they're not using for training or

18:57

inference of LLMs at a particular time.

18:59

For example, like Meta, if they're not

19:01

using some of the GPUs they bought, they

19:03

can just say we'll just make our Insta

19:06

ad uh serving uh model slightly better

19:09

for today. What is the equivalent sort

19:11

of reserve use of compute for Jane

19:13

Street that's just a lower bound on how

19:15

much that's worth for you?

19:16

>> Part of what's going on is like in many

19:17

ways we're just like very compute

19:19

constrained. There's lots of innovation

19:22

and experimentation and new ideas that

19:24

people have that is bounded by the

19:26

amount of compute that we have. And so

19:28

like in some ways like if we just think

19:30

about like like we do a we try and and

19:32

do a kind of moderately rigorous job of

19:34

thinking about the value of the new

19:38

different runs that we can do and the

19:41

value of the runs that we're turning

19:42

away is really quite high, right? So

19:44

like we're doing what we think are the

19:46

most valuable things but you know if you

19:48

know if it turns out we have more

19:49

compute than we need for those there's

19:50

just like a ton of other

19:52

>> research and experimentation that we can

19:54

do in that space. So like we're we're

19:56

we're nowhere near to like being like oh

19:58

too much compute like we sort of have

20:00

have the opposite problem. I think

20:01

there's also really lowhanging fruit in

20:04

that direction. Just like it's valuable

20:06

to retrain the models more often.

20:08

There's some decay in the quality of

20:09

models over time and being able to rerun

20:11

them like that's that's kind of has

20:14

immediate and clear value to the firm.

20:17

Uh there's also some amount of bulk

20:20

inference tasks that that we can do that

20:22

can like fill in the gaps in the systems

20:24

where there's nothing else to schedule.

20:26

Um so we don't quite have the thing that

20:29

looks like the analog of like the

20:31

Instagram ad serving thing, but there is

20:33

just like a ton of other like kind of

20:36

dark space of like things that we're not

20:38

doing, but we would if we had more

20:40

compute. So we're like pretty

20:41

unconcerned about getting value out of

20:43

these. Here thing there is like there is

20:46

a bunch of embedded bets like we are

20:48

like investing a lot of money in in this

20:50

stuff and you could imagine that like

20:53

things won't get better at the rate that

20:55

we are thinking they will in terms of

20:56

like the value of the individual models

20:58

and trades that we're doing and like

21:00

it's a competitive environment. maybe

21:02

other people will out compete us. We're

21:03

like I think one part of remaining good

21:05

is like always being nervous about other

21:07

ways that competitors can like figure

21:10

out doing similar things to what you're

21:12

doing and reduce the value of of that.

21:14

So like there are ways there are ways

21:16

that it might not work out but uh

21:19

certainly with anything like the current

21:20

mix of compute jobs that we have we're

21:22

just like very far from having this

21:24

problem. It's it's interesting to this

21:26

doesn't exactly answer it but like you

21:28

know you could disconnect the uh the

21:30

powering the data center from the chips

21:32

and say okay well you know I I I might

21:34

need to use this compute later let me

21:36

commit to the data center and the power

21:37

now but like delay the decision on the

21:39

chips which are very expensive right and

21:41

and just be slightly long power and data

21:44

center for that that that point of time

21:46

where you might need that compute um and

21:48

then we'll build in situations where hey

21:50

maybe we can kind of offload some of

21:51

that capacity to somebody else it's much

21:54

easier for us I is to offload power and

21:56

data center capacity than is the chips

21:57

themselves for obvious reasons. But uh

22:00

you can you can really bifrocate those

22:01

two.

22:02

>> This also changes the considerations

22:04

around hiring. I mean you already have

22:06

like the highest bar for hiring but it

22:08

just increases even more if you hire one

22:11

more person that is one person who will

22:13

need compute to do their experiments and

22:17

that compute is going to be traded off

22:19

against somebody else who's excellent on

22:21

your team who could be doing experiments

22:23

themselves. I I hear what you're saying,

22:25

but we don't think, oh, it would be

22:27

weird to hire more researchers because

22:29

then we'd have to give them more comput.

22:30

It's more like the research is

22:31

incredibly valuable. The researchers are

22:34

incredibly valuable. This is a good

22:36

argument for buying more compute. Um,

22:38

and so we're like very axed to grow the

22:41

amount of compute. Like these days, we

22:42

are in something like the range of like

22:44

tens of thousands of GPUs and we will in

22:46

not too long be in the range of hundreds

22:48

of thousands of of GPUs. And we think

22:52

it's like well justified by the business

22:55

like you know it's not it's it's not

22:57

like it's it's not like you know we're

22:59

worried about like oh you know can we

23:01

justify it based on like the penals of

23:03

the trading strategy. It's like no no no

23:04

it's like these are clearly good

23:06

investments. Um so it doesn't feel like

23:08

it's slowing us down on the hiring

23:10

front. In some ways, the the biggest

23:12

impediment to growth is that it takes

23:15

time to like really train people and

23:17

absorb them into the culture and kind of

23:19

build build them up and build the place

23:21

up. Like we want Jane Street to continue

23:23

to be a great place to work. Like I I

23:25

just don't think of the hardware thing

23:26

as at all being the thing that slows us

23:27

down. And and I think the real limiting

23:30

factors are finding great people and

23:32

having the mentorship capacity for them.

23:34

>> I guess this might be a good opportunity

23:35

for you guys to mention what kinds of

23:37

roles you're currently hiring for. Oh

23:39

man, why don't you start in the in the

23:40

engineering space?

23:41

>> Yeah, I I'll start. I mean, I think so,

23:42

we're generally just looking for really

23:44

smart people, people that that that are

23:46

interested in in in doing this stuff and

23:48

and that's, you know, mechanical

23:50

engineers, electrical engineers, project

23:52

managers, architects, people that help

23:54

design and build some of these spaces.

23:56

And, you know, our our our remit uh in

23:59

my team is is really to to find the

24:01

spaces, to design them, to construct

24:03

them, and then to operate them, right?

24:04

So, it's full life cycle. So in each one

24:06

of those you kind of need people you

24:08

know lots of engineers, lots of what we

24:10

call physical engineering which is a

24:12

madeup term that that we came up with

24:13

but uh you know mechanical engineers and

24:16

structural engineers maybe electrical

24:18

engineers those types of folks

24:20

>> and and machine learning and trading in

24:22

general is really like a whole team

24:23

sport and so we want to hire people from

24:26

lots of different backgrounds and with

24:28

lots of different capabilities. Uh we're

24:30

certainly like very excited to hire

24:32

people with kind of you know specific

24:34

like machine learning backgrounds of

24:36

like you know designing architectures

24:38

and building models in various cases. We

24:40

both I mentioned that we have like a

24:42

bunch of like custom architectures and

24:44

stuff for like our own bespoke kind of

24:46

kinds of data that we need like the data

24:48

kind of characteristic of the markets.

24:50

Um we also build LLMs and people who

24:53

experience in all sorts of part of the

24:55

life cycle of LLM training. we're

24:58

interested in hiring and have been

25:00

growing that area. Um, you know, we we

25:04

hire lots of like people with like

25:07

generally good scientific and technical

25:10

backgrounds from like math and CS and

25:12

physics and engineering and stuff to be

25:14

traders and like there's a kind of mix

25:16

of skills there. Uh, but that's like an

25:19

area we continue to be very excited to

25:20

hire in. On the software engineering

25:22

side, there's like a general software

25:24

engineering role which we're always

25:26

eager to get great people for uh that,

25:29

you know, I think just rewards a little

25:30

bit, you know, it feels a little silly

25:32

to say, but just like, you know, as Dan

25:34

was saying, smart, curious people with

25:36

really good CS backgrounds, uh, you

25:39

know, fit into that generalist role and

25:40

there's lots of different kinds of

25:41

things they can end up doing. There's

25:43

also a bunch of interesting specialized

25:45

areas where we really are excited. Like

25:47

here's a thing that's kind of new. With

25:49

all of this scale, we are much more

25:51

interested in fleetwide optimization

25:53

than we were in the past. Like

25:55

>> we our old view about about performance

25:57

optimization was that it was much more

25:59

about, you know, making the things that

26:01

were most speedritical as fast as

26:03

possible. And more generally, yeah,

26:05

compute's kind of cheap and like people

26:07

are expensive and we're not spending

26:08

that much time optimizing our general

26:09

compute. But like, man, we're doing a

26:12

lot of general compute now. You know,

26:13

you start investing billions of dollars

26:15

in this stuff and it just becomes more

26:16

valuable there. And there are people who

26:18

have experience in doing this at some of

26:19

the hyperscalers and we'd love to hire

26:21

more people with that kind of background

26:22

to think about the optimization problems

26:24

that we're hitting which are like

26:26

related in important ways different but

26:28

like you know so it's like both a

26:30

related challenge and a new one. Um

26:33

we're like we do a lot of fun like

26:36

hardware engineering stuff. We're like

26:37

working on our own AS6 people with that

26:39

kind of experience is super exciting.

26:41

Um,

26:44

one thing that we mentioned a little

26:45

earlier at lunch was like we're starting

26:47

to think about building out a formal

26:48

methods team using basically

26:50

mathematical proof to make software

26:52

engineering more effective.

26:54

>> That's like a new very speculative area

26:56

and we're like very excited to find

26:58

people there. We feel like that's a kind

26:59

of a set of a whole community of people

27:02

who in the past I feel like I've always

27:03

had to disappoint by like yeah we're not

27:05

interested in formal methods but like I

27:06

think the whole AI revolution makes

27:08

formal methods suddenly a much more

27:10

interesting field and so it's a place

27:12

we're excited to invest in. So I don't

27:15

know and like I don't know project

27:16

managers people who do front-end dev

27:18

actually like for most of Jane Street's

27:20

experience we pretended like this whole

27:22

web thing had never happened and like

27:23

almost all all of our tools were just

27:25

like in the terminal but you know it

27:27

turns out it's useful to be able to like

27:28

draw a straight line and you know have a

27:30

tool tip and things like that. So, we've

27:32

actually invested a lot in building

27:33

really good tools for doing front-end

27:35

development and building tools for

27:37

people and having great front-end

27:38

engineers who are both really good

27:40

software engineers and have a good sense

27:42

of what it means to make an application

27:44

that's good for a person is really

27:46

important. I say like as as a general

27:48

meta point about all of this, I think

27:49

that like in all of the like legitimate

27:52

and real excitement around AI tooling, I

27:54

think people sometimes like

27:56

kind of miss out on the importance of

27:58

the human element of all of this. I

27:59

think that we really we really care a

28:01

ton about building tools that are good

28:03

for people and that comes that includes

28:06

the AI tooling itself, right? I think

28:08

trying to drive tooling in a way that

28:09

increases human understanding and agency

28:12

and efficiency is like that's the core

28:14

thing. We are limited more than anything

28:15

else by the amazing people who work

28:18

there and like being able to find more

28:19

of the right people and grow the

28:20

organization so that we can get more

28:22

done. Uh and so we have a very kind of

28:25

humanoriented way that we think about

28:26

the systems that we build.

28:29

Um, it's been really cool to have you

28:31

guys um, make these fun puzzles and

28:35

challenges. I think in general you do

28:37

that, but also you um, you've been uh,

28:40

you guys have made a couple for the

28:41

listeners of the podcast in particular.

28:42

And I think people who are listening to

28:44

this might find it interesting to check

28:45

those out. um uh including one by the

28:48

way which not only was nobody who

28:50

submitted to the competition able to

28:52

solve but Jane Street itself cannot

28:55

solve which um uh which involves finding

28:58

back doors to various LLMs that have a

29:00

trigger phrase baked into them. Anyways,

29:02

I mentioned this because um uh to the

29:04

extent people are interested in learning

29:06

more, I think these are the kinds of fun

29:07

puzzles that might give some indication

29:08

of what work is like and um uh why

29:11

things like fun place.

29:12

>> Yeah, puzzles are a deeply embedded part

29:14

of the culture. So, it's kind of great

29:15

to use them as a way to reach out to

29:17

people as well.

29:17

>> Yeah. Yeah. Um, I guess the plug here in

29:20

this case is janestreet.com/doresh.

29:24

Uh, so that people can learn more about

29:25

the open worlds and about all these

29:27

puzzles. Yep. Awesome. Cool. Thanks for

29:29

doing this, guys. Thank you very much.

29:31

>> Our pleasure.

Interactive Summary

The video features a tour of a Jane Street data center and a discussion with Ron Minsky and Dan Ponttovo about their trading infrastructure. They explore the different time horizons of their trading strategies, ranging from nanosecond-speed FPGA-based decisions to longer-term model-driven trading. The conversation highlights how they approach compute constraints, the necessity for specialized hardware, and why they prioritize human judgment alongside automated systems. They also address their hiring strategy, emphasizing their need for diverse talent in software, engineering, and machine learning to tackle complex problems.

Suggested questions

4 ready-made prompts