HomeVideos

Powering Agentic AI with AI-Ready Data Platforms That Turn Data Into Intelligence

Now Playing

Powering Agentic AI with AI-Ready Data Platforms That Turn Data Into Intelligence

Transcript

578 segments

0:02

Thank you. Dennis.

0:03

Good morning everyone.

0:05

Welcome to day three of GTC Taipei 2026.

0:10

I like that you're all the early birds. You got here first.

0:13

You get all the good content today.

0:16

So I have a bit of a tradition.

0:18

Every time I do this, I have to take a selfie with you all.

0:22

So pose, smile, jump, do whatever you want.

0:26

But on the count of three,

0:30

here we go.

0:31

One, two, three.

0:35

Alright.

0:36

Thank you.

0:38

Well good morning.

0:40

Like Dennis said, my name is Jason Hardy.

0:42

I'm the vice president for storage technology here at NVIDIA.

0:46

It's funny being the storage guy in a company that doesn't sell storage.

0:52

So it's been very exciting though.

0:55

And we're finding that data continues to play

1:01

a monumental role in the success of AI, especially as

1:08

we start to see agentic capabilities

1:11

really start to grab ahold and take place.

1:14

There's no denying that from an investment perspective, we're seeing a

1:19

continuation of investment into AI technologies

1:24

and the necessary infrastructure to be able to get to these outcomes, and this

1:29

will continue to happen for many years to come.

1:32

And we're really excited about this.

1:36

Now, AI though, inferencing on the agentic side

1:41

or just even simple engagement just with a chatbot.

1:47

It relies on a grounding

1:50

in the data that sits behind it.

1:54

The data is what describes the personality

1:57

of an enterprise, describes what we do and how we operate personally, but also

2:04

how companies, businesses, enterprises, how they exist,

2:09

what they do, how they operate.

2:11

All of that is grounded in data.

2:16

The problem that we're starting to see

2:18

is that in order for the success of AI

2:23

to really grab ahold, it's becoming

2:27

increasingly dependent on data.

2:31

And we see the facts and statistics around the AI projects

2:35

that are abandoned or that never see past proof of concept.

2:39

But ultimately, at the end of the day,

2:43

AI is as good as the data that's fed to it, that it's used

2:47

against it, that it's trained against.

2:50

And especially as we look at this agentic frontier that we're starting to

2:54

to embrace, and we're really starting to see some amazing capabilities to

2:58

come from, the requirements from a data perspective, especially on

3:04

the enterprise storage systems, the things that

3:08

have maintained that data for 50 plus years.

3:12

The dependency on that is changing those requirements.

3:17

It's changing how those systems are built.

3:19

It's changing how those systems operate.

3:21

It's changing the requirements and the services that even need

3:25

to come from those platforms and systems.

3:30

You see, when we look at an

3:31

enterprise, it's really made up of three different kinds of data.

3:36

There's structured data, while small in volume, very critical to

3:41

how an enterprise operates.

3:43

This is the ERP systems, the SQL systems,

3:47

the databases, that structured data

3:50

that business decisions are made off of.

3:55

Then there's the unstructured data side—your home directory. SharePoint.

4:00

Some people love it. Some people hate it.

4:03

OneDrive, or Google Docs, or NAS shares,

4:08

or PDFs, or video, or audio, or files.

4:13

That is unstructured data, and that is by far the lion's share,

4:18

the largest volume of data in an enterprise.

4:23

This also is the hardest to process

4:27

because it is made up of

4:29

so many different personalities and so many different data types

4:32

and ways to engage with it, that it is the hardest, but the largest

4:37

volume of data that describes an enterprise.

4:41

And now we have this

4:42

third category, this third classification of data.

4:46

This is new.

4:47

This is what's come from the agentic

4:50

push and and the embracing of this technology.

4:53

And this is context. Agent context.

4:57

Some people know context as KV cache.

5:00

It's actually more than that.

5:02

KV cache is a piece of it.

5:04

But it's also the memory, the

5:07

data that gets generated as agents operate, and as they

5:12

consume information that we feed them and how they

5:15

and the task at hand and everything, that's the data that the agent

5:19

is responsible for consuming, whether it be what we feed it

5:24

and it activates off, of or what it generates itself.

5:28

And even, at times, shares from agent to agent.

5:33

This has to persist over long term.

5:37

Some agents live for hours; others live for years.

5:42

So this data becomes as critical

5:46

as the rest of the enterprise ecosystem— as critical as what the structured data

5:52

that we continue to maintain

5:54

with the tools and applications, or even the unstructured data

5:57

that we generate on a daily basis as we just go about our day, our job.

6:03

So across these three landscapes,

6:06

this data here is the foundation of what enterprises are built on.

6:12

But you see, from an agent demand perspective, there is a

6:17

new paradigm that's happening,

6:19

especially for storage.

6:21

You see, traditionally we had stateless environments.

6:24

We had the structured data, we had a user doing a thing, and

6:28

it generated a workflow.

6:30

And then we also had kind of two states for data to manage that—data at rest.

6:36

What's living there?

6:37

It could be an archive. Or data in motion.

6:40

That active data, that data set that we operate against.

6:44

We're modifying or manipulating.

6:47

Now, as we shift into the agentic frontier, though,

6:52

it changes this whole paradigm.

6:54

You see, now we aren't just stateless where it's short lived.

6:59

Now we have long-running sessions.

7:01

Who here has used agents in one way, shape, or form?

7:06

You guys really need...

7:07

You're at an NVIDIA conference.

7:09

Come on.

7:09

You should you should be using this stuff.

7:12

No, these tools are amazing.

7:14

And in fact, even as I use my own agents, I have the liberty

7:19

to give it a task, to go to bed, to wake up,

7:22

and it's worked all night.

7:24

That task is done,we review, and then we iterate from there.

7:28

So it's no longer just... I end my day, and then data is done.

7:33

I end my day, and it continues.

7:36

I'm also blending unstructured

7:38

and structured data together for these agents to operate.

7:42

So it's as much SQL as it is a PDF.

7:48

And then

7:49

it's not just me engaging with my data; it's

7:54

the one, the tens, the hundreds, even the thousands

7:59

of agents that will continue to iterate over

8:02

data, continue to consume it, and it will consume it, not at the rate

8:06

at which a human interfaces with data, but at the rate

8:10

at which agents can start to interface with

8:13

both the data and each other.

8:17

And trust me, it's a lot faster

8:19

than you reading a document.

8:22

So we have scale.

8:24

We have new data types.

8:25

We have data that's active 24/7.

8:28

And this growing context of needing

8:32

to keep both, yes, KV cache, but also all of the memory, and the scratch,

8:38

and the to-do, and the embeddings, and everything that goes on

8:41

to build context around how agents operate.

8:46

Everything that agentic workflows

8:49

are starting to push into, that we're starting to embrace on a day-to-day

8:53

basis, is changing the landscape of what storage actually means.

8:59

And so storage needs to evolve.

9:02

It needs to move from being what we've always known

9:05

as a system of record, to now

9:08

a system of context.

9:13

Now, what makes this hard for the enterprise?

9:16

Obviously, things change very quickly.

9:18

So we need to be able to do things at scale.

9:22

We need to be able to scale cache.

9:24

We need to be able to scale those embeddings.

9:27

That context.

9:28

Everything about how agents operate brings more scale,

9:33

performance, and density.

9:36

We also need to introduce new technologies into the enterprise.

9:40

Things like RAG, or retrieval-augmented generation, or embeddings,

9:44

and the ability to take in all of that data and embed it into a vector database.

9:50

How do we chunk that up?

9:51

How do we operate with that?

9:54

We need to scale that workload

9:56

so that our agents are fed at the speed at which they operate.

10:01

And we cannot forget about the security, the governance, and

10:07

even how enterprises operate today,

10:09

from a hybrid perspective.

10:11

You see, enterprises aren't just in the data center, and they're not

10:15

just in the cloud. They're on both.

10:17

So in that is, how do we best enforce

10:20

and represent the data where it needs to be on prem, cloud, wherever, but also

10:25

ensure that we do this securely, and we do this in a controlled manner,

10:30

do this in a way that we can understand and audit, if ever necessary.

10:34

So we still have many enterprise requirements that need to be maintained,

10:40

but also need to evolve forward to address the agentic demand.

10:48

This is why we believe that storage is moving

10:51

from that system of record to, really, the new data plane.

10:56

And it's the new data plane for agentic AI.

11:01

You see,

11:03

much like Jensen has talked about the

11:06

AI layer cake, the five layers of AI,

11:10

storage has much the same, but there's three of them.

11:15

So maybe it's a cupcake.

11:17

The cupcake of storage.

11:20

But what it means is it's based on high-performance requirements,

11:25

still having enterprise data,

11:27

but then how do we bring agentic context into that?

11:32

When we look at it from a storage perspective,

11:35

just at the raw capability to introduce and support, or to be able to support

11:40

the high-performance requirements, those continue to scale upward, in both

11:47

how do we train and inference, but also how do we bring more context and volume?

11:53

How do we support inferencing at that scale, from a performance and demand?

11:59

This is why we're extremely excited that Vera is here.

12:05

You see, Vera was designed

12:08

to take stuff off the wire—

12:12

information, bits and bytes, packets, process it through, and then give

12:18

it to the GPU as quickly as possible.

12:22

So the 3x memory bandwidth plays a big role in being able to do that.

12:26

The single core... single threaded performance improvements,

12:30

best-in-class, is great for the GPU.

12:35

Well, you know what else it's great for?

12:38

Storage.

12:40

Vera, through the capabilities that we built to be able to drive

12:45

the GPU, benefits the same to storage.

12:50

So as we look at now introducing

12:53

Vera to the market...

12:58

The NVIDIA Vera

12:59

Bluefield-4 STX platform, our opinionated design

13:03

of what storage looks like, benefits from all of those capabilities.

13:08

It's the foundation for how we'll continue to build the AI factory, and in working

13:13

with our partners. It's how we will enable storage to scale

13:19

to meet the demand from an agentic perspective,

13:22

at the very base of high-performance

13:24

storage and object storage

13:26

and all these storage technologies to be able to drive that forward.

13:30

It's how we bring Vera. It's how we bring our software.

13:32

It's how we bring all the greatness of the NVIDIA

13:36

ecosystem, plus our storage partners' capabilities.

13:39

Bring those together, and release that into the market to be able

13:43

to address this demand from the agentic workload.

13:52

Second, though,

13:54

it's not just an infrastructure thing.

13:57

It's also about how we prepare the data, how we govern the data, how we bring

14:02

lineage and indexing and everything necessary

14:04

for enterprise data to be agentic ready.

14:08

This second layer of the storage cupcake

14:11

is about addressing the challenges to make sure

14:15

that we don't become another headline.

14:19

You see,

14:21

if you don't address the data accurately,

14:23

if you don't bring the right data to the context,

14:27

things can go a little sideways.

14:31

Air Canada got in trouble with the Canadian government

14:35

because they had a chatbot hallucinate new refund policies.

14:39

Didn't exist. The courts didn't care.

14:42

United Healthcare got sued for faulty AI, and in that,

14:45

had to pay out a significant amount of money.

14:50

Google had to play a tremendous amount of money

14:54

because of wrong answers, because of wrong demos.

14:57

The same with Deloitte, and same with on and on and on and on.

15:02

This isn't an AI issue.

15:05

This is a data issue.

15:08

And this is why data needs to be especially

15:11

prepared to be able to address the AI demand.

15:16

Now, it's complicated to do this.

15:18

There's no denying that.

15:20

First, you have to ingest your data.

15:22

And this could be whatever type of data.

15:24

It could be PDFs, it could be tables, it could be structured

15:28

data, it could be audio, it could be video,

15:30

It could be very multimodal.

15:33

The last time you opened up a PDF, when was it just text?

15:37

Never.

15:38

So you need to understand not just that there's text in the document, but you have

15:42

to be able to understand and process through the text

15:45

and the tables and all of the imageries

15:48

and how it relates to each other.

15:51

So you go through this ingestion process, then you have to cleanse your data

15:55

so that it's ready. You have to remove duplicity.

15:58

You have to remove or normalize the data.

16:01

Sometimes you have to strip the personable... the PII, the

16:05

sensitive data out of it, and then you want to enrich it.

16:08

You want to bring metadata to it.

16:10

That way, you have data that describes the data.

16:13

Then chunk it. Break it up.

16:16

Make sure it fits into these semantic units, sized appropriately for

16:21

the model that you're running.

16:23

And then you embed it into the vector database.

16:25

And then you index it. And then you govern it on top of it.

16:31

It's not straightforward.

16:34

We realize that for data to

16:37

be properly processed for AI readiness,

16:42

it takes work, and it's complicated.

16:48

You see, agents need accurate data.

16:50

And unfortunately, data has all of this complexity with it.

16:55

It has pipeline complexity.

16:57

How do we how do we use it,

16:59

even in our business? It has structured complexity.

17:03

How do we associate, and again, in a PDF, this graph and this text

17:08

and this... all of this stuff, how does it correlate with each other?

17:12

And then data changes over time.

17:14

It's not static.

17:16

Just like our businesses change over time.

17:19

And then, data has gravity.

17:22

Data has mass.

17:25

Turns out that the best bandwidth in the world is actually FedEx.

17:29

Go Google it.

17:29

It's kind of a meme.

17:33

But really, when we talk about data,

17:35

data is a complex ecosystem that continuously changes.

17:40

Just like we do. Just like our enterprises do.

17:43

So how do we bring this data to agents

17:47

in a way that's accurate for them to be able to consume this?

17:52

And how do we do this to make

17:54

sure that we de-scope as much risk as possible?

17:58

A radiological diagnosis— if something goes wrong,

18:03

that can impact a patient health.

18:06

Or how do we use this technology now

18:09

to reduce the errors, from a radiology perspective?

18:13

Or how do we use this data to identify where

18:17

we can reduce theft inside of the retail space?

18:22

Or how can we use this data to improve

18:25

and drive on first-call resolution from a call center?

18:29

See, agents, while there's risk with them,

18:33

also bring a tremendous benefit.

18:36

Better health care, better financial state,

18:40

better identifying of theft and loss.

18:43

Improved profit. Everything about it.

18:47

You see, this agentic outcome, when we feed it

18:50

accurate data, can give us significant upside.

18:55

And this is why

18:56

we're super focused on aligning the data

19:00

and storage ecosystem to transform,

19:03

to deliver AI-ready data.

19:07

As the agentic harness that Jensen talked about during this keynote, as it starts

19:14

to process data, and as agents create that context

19:17

and go through this loop, and then act on the end of it, and we feed

19:21

more data in, and tools and all the governance and security, and then

19:24

ultimately create that memory that it lives from—

19:28

All of that requires a new paradigm, a new approach

19:33

for the enterprise data ecosystem

19:36

to be able to address this demand. We need to be able to have better search.

19:41

We need to bring intelligence into it.

19:43

We need to be able to identify whether it's structured or unstructured.

19:46

How do we access it?

19:48

We need to ensure that our agents have access to only

19:51

the data they're allowed to.

19:53

All of that comes from this new data plane that is the storage platform.

20:01

So as we now look at this, it's still about

20:05

having enterprise storage systems that have protocols

20:08

and file systems and orchestrate data and snapshots and cloning and everything

20:14

that's necessary from an AI data infrastructure perspective.

20:18

But it's more than that.

20:20

It's about how we bring now further capabilities

20:24

and how we manage that data. How we govern it.

20:28

How we bring deduplication and compression. How we bring keyword search

20:32

into it, better access control, all the analytics necessary.

20:36

This enrichment starts to add functionality

20:40

into that foundation from an enterprise storage system.

20:44

Then how do we prepare the data for AI?

20:47

By bringing in that continuous ingestion,

20:50

extracting data actively as it exists.

20:53

Not later, not in chunks, not in a couple hours, but in real time.

20:59

How do we index it so that the agent knows about the data when it's created?

21:03

So it's making accurate information.

21:06

Storage needs to prepare data on creation.

21:12

And then manage it.

21:15

How do we bring the best management into that

21:17

so that we can redact data, mask it, classify it,

21:21

embed it, ensure that we're not exposing sensitive data.

21:26

Again, all at time of ingestion.

21:31

So all of this now—it's not just an enterprise

21:34

storage problem, but is an AI data platform capability.

21:39

And this is what enterprise storage

21:42

is becoming. The AI data platform.

21:45

And we at NVIDIA have been working with our partners

21:49

in building exactly that—the AI data platform.

21:55

It's about bringing that readiness

21:57

at time of creation, at the source of it, from a CPU perspective starting,

22:02

but including, now, the ability for GPUs to be added

22:06

in to the enterprise storage ecosystem so that now we can have all of that

22:11

amazing functionality for vectorization, all the embeddings, all the even

22:16

inferencing at that data, directly inside the storage platform.

22:25

And to help this along, it's not just

22:28

saying, here's a vector database, here's a GPU, Mr.

22:32

Storage Partner, go take care of it.

22:35

But we've actually been building blueprints that help simplify this.

22:40

This is our RAG blueprint for

22:43

for helping and streamlining the storage

22:47

system capability to vectorize and embed and do all the work necessary

22:52

to prepare data within the storage platform

22:56

for agentic AI readiness.

22:59

The key here is, is that this is done at the point of creation

23:03

where the data exists, not decoupled somewhere else.

23:09

Or what we're doing with

23:11

VSS or Video Search and Summarization blueprint.

23:14

Where now, agents have the ability to understand video content

23:18

because we can summarize it.

23:20

We can we can capture off of it.

23:22

We can understand the context of the video.

23:24

We can process it, all at the point of creation in the storage system.

23:34

Now, the great thing is

23:35

we've been talking about this for a while, and our partners

23:39

have been working with us very closely.

23:42

So it's not just about an idea that we have, but that they've

23:46

put this into execution—

23:49

that Dell with their storage capability

23:52

or a NetApp, or IBM, or the list goes on and on, have all taken this capability

23:58

and have embedded it into their storage systems

24:00

to create an in-place AI data platform solution

24:05

that's embedded and a part of their storage ecosystem now.

24:10

So we're super excited about this because this now starts to prepare data

24:15

and get it ready for, again, at the time of...

24:17

at that point in time of creation, so that

24:20

our agents can now have better context, our agents can

24:23

make better decisions, all based and grounded on data

24:27

that's been prepared for them in a way that they understand.

24:33

So that idea

24:35

of understanding data

24:37

at that time of creation so that my agent can be more effective—

24:41

it's not just an idea. It's being put into execution.

24:48

So now when we look at it, every enterprise, every AI factory

24:55

needs an AI data platform.

25:00

You see, an AI data platform—

25:03

you could think of it like

25:05

the raw goods for an AI factory.

25:10

Just like if we have a factory that builds stainless steel or whatever,

25:15

and you take raw iron ore, you feed it

25:19

through the factory to create that stainless steel at the end.

25:22

I think that's how stainless steel is made.

25:25

The same goes for the AI factory.

25:28

It's about taking that data,

25:31

preparing it at the edge at the time of creation

25:34

in the storage system, and taking that refined good

25:37

and bringing it to the AI factory

25:39

so that it can manufacture intelligence,

25:43

grounded on how the enterprises operate.

25:47

So now, it's how the AI factory becomes more effective, is by grounding

25:52

it in data that AI data platforms have created or have refined, so

25:57

that its accuracy and its intelligence is as high as possible.

26:09

But every AI data platform—

26:12

it still needs to be real-time and accurate.

26:16

It still needs to continuously index.

26:19

It's not one and done.

26:21

It's about continuing to improve on itself as well.

26:25

Not just saying, hey, I have one indexing model or one RAG model.

26:29

It's I've got different ways of understanding this

26:33

and evolving forward to make it enterprise-grade.

26:37

It's also how, again,

26:39

we introduce data, new data types into the enterprise, all the time.

26:44

How do we make sure we understand that?

26:46

How do we make sure we grasp that?

26:47

How do we make sure that we

26:49

incorporate that into the model so that our agents can understand new data

26:53

with new context, to be able to make, again, a more informed decision?

26:59

How can we discover data quickly? New data.

27:03

And then, how do we make this turnkey?

27:06

How do we make this simple?

27:07

And this is really important because this is also

27:09

where the enterprises have a lot to do.

27:12

And it's very complicated to do this stuff.

27:14

So we need to make it easy.

27:17

And that's what we're doing with our storage partners—

27:19

the Dells, and the IBMs, the HPEs, and the NetApps—making this turnkey

27:24

so that it's fully integrated in; it has all the networking, the compute, the software,

27:29

everything necessary so that a consumer of it, a user of it, whether

27:34

it's an agent or a human, is able to just flip the switch.

27:40

Now, it's a little bit more complicated than that, but it really

27:44

is about making it that easy or as easy as possible

27:47

and continuing to bring in the security and the compliance and the governance

27:52

so that it still maintains that trust.

27:57

Because at the end, it's about trust.

28:00

And by doing this and making it simple, we also don't want to compromise

28:04

on the trust of the data.

28:06

And, again, this is why we're really excited about

28:09

what our partners are creating inside the AI data platform ecosystem—

28:13

consuming our blueprints, driving forward, and bringing a lot more value

28:17

in different data types and different capabilities,

28:20

directly into the storage ecosystem.

28:25

The last layer of the cake,

28:27

the cupcake. The frosting.

28:31

Context.

28:33

Addressing context is not easy.

28:37

In fact, addressing context

28:39

required us to think about the structure of memory inside of a data center,

28:44

inside of an AI factory, inside of a GPU.

28:49

So we created this kind of hierarchy of memory here

28:53

that allows us to describe

28:56

what type of memory lives where.

28:59

G1 is HBM. It's the fastest memory.

29:03

It's directly in the GPU. It's where the active KV live.

29:06

It's where the weights get loaded.

29:08

Everything about the model and inferencing

29:11

happens in G1. But G1 is finite.

29:16

G2 is system memory.

29:20

That tightly coupled memory, where both the HBM, G1, and the system, G2,

29:26

can collab or can work together in

29:30

as fast... as fast as possible.

29:33

This is where KV cache spillover happens.

29:36

This is where staging happens for models, for inferencing,

29:39

and all sorts of different things.

29:43

G3 is scratch, that NVMe disk in each server.

29:47

It's where the OS lives, simple things like that.

29:50

And then G4 is high-performance storage, where models live, where checkpointing

29:57

happens, where inferencing takes place.

30:00

G4 even extends into enterprise storage.

30:05

Now, when we were looking at context,

30:08

what we found is that G4

30:11

wasn't quite fast enough, and G3, it wasn't quite enough of it.

30:17

So we created G3.5.

30:19

And this is where KV cache lives.

30:22

Now KV cache is really important because when we are inferencing,

30:29

we have to calculate

30:31

a lot of information on the prefill and the decode stages.

30:35

This is how inferencing happens.

30:37

And the thing about KV cache is, when you cache

30:42

that information, you don't have to recalculate it, which is great.

30:45

That means the GPUs can spend more time working on new data,

30:49

creating new tokens, instead of having to recalculate, which is awesome.

30:54

That means that we get the best out of the factory.

30:57

So in order to address KV cache, we needed something

31:02

that was fast enough to support the model or the GPU

31:05

without having to recalculate, but also was cost effective and didn't

31:10

didn't create too much burden to be able to do that inside the factory.

31:16

So from our perspective, KV cache,

31:19

this is where we've released CMX,

31:22

or KV cache storage capability.

31:27

This new tier of storage ensures that we can offer

31:32

tokens, offer cache to the GPU as quickly as possible, but without the

31:37

extra burden, without being too inefficient,

31:41

All of the great stuff that's necessary to make this effective.

31:45

That addresses KV cache, though.

31:48

Like I said earlier, there's another type of context—

31:51

this memory context, this how an agent operates and

31:57

the data it generates when

31:59

consuming data, and when working with other agents, and everything about it

32:04

that's not just tokens.

32:08

That cache is...

32:09

Jensen talked about it earlier during the keynote was, it's the ontology.

32:14

It's how agents' information and how we are described as

32:19

as users of this, that... how that permeates through the system.

32:23

That context,

32:26

we're still doing a lot of work on, but we do know that we can make

32:30

the factory as efficiently as possible by bringing in this KV cache concept.

32:36

So this new tier of storage, again, in working with our partners, allows for

32:42

better scale of the AI factory,

32:46

from a token throughput perspective. More tokens. More intelligence.

32:52

It also scales to meet the demand of as the factory grows. Start small, go big.

32:58

Support, again, that workload as we continue to scale the factory size.

33:04

It's also designed to

33:06

complement both GPU memory and network storage by sitting in between.

33:11

It doesn't replace HPS or G4.

33:14

It augments it.

33:16

It brings the right type of storage for the right platform

33:19

or the right requirement.

33:22

So KV cache is really designed to address that

33:26

acceleration and that token-per-second output from the AI factory.

33:33

But again, this is designed not by us, but with our partners.

33:38

So we'll continue to go through this with them.

33:41

We'll continue to address this.

33:42

And you'll start to see this permeate the AI factory, because there is

33:47

value in being able to manage KV cache more appropriately,

33:52

to reuse it as much as possible so that the GPU doesn't

33:56

spend time recomputing when it can create new.

34:04

So when we look at the challenges

34:06

of the data center, now, when we look at the challenges

34:09

of the data ecosystem now, it's not about just bringing storage in.

34:15

It's not about just bringing a GPU in.

34:18

It's about extreme co-design across both the stack horizontally...

34:23

I should say vertically and horizontally.

34:27

It's how we bring network and how we bring GPU

34:30

and how we bring compute, especially with Vera now, it's how we bring Spectrum-X in.

34:34

All of this working together to accelerate

34:38

and provide as much value as possible, and then working with our partners

34:43

on the horizontal perspective to be able to create new capabilities

34:48

inside their portfolios.

34:51

It's about addressing the agentic AI demand

34:54

through the need for the AI data platform.

34:57

That scenario, whether it's or the CV cache scenario,

35:02

or even just high performance, just the parallel

35:06

file system, high-performance storage requirement.

35:09

This here is NVIDIA Bluefield-4 STX.

35:15

NVIDIA Vera Bluefield-4 STX.

35:19

This is our design.

35:21

This is our collaboration with the storage

35:26

ecosystem to address the agentic AI demand.

35:29

And we're really excited about the capabilities, espeically

35:32

like I said earlier with what Vera is bringing into the market, the capabilities

35:36

that it's going to be able to allow

35:38

for our storage partners to accelerate with and drive additional context

35:42

and additional value, and then bringing the GPU in with Rubin

35:46

or with RTX Pro, to be able to drive that

35:49

from an AI data platform perspective.

35:51

All of this together is reinventing the storage ecosystem,

35:56

so that, again, it's agentic AI ready.

Interactive Summary

In this presentation, Jason Hardy, VP of Storage Technology at NVIDIA, discusses the evolving role of storage in the era of agentic AI. He highlights how the rise of AI agents necessitates a shift from traditional 'system of record' storage to a dynamic 'AI data platform' capable of managing structured, unstructured, and context-heavy data like KV cache. Hardy emphasizes the importance of preparing data at the point of creation, integrating GPU capabilities directly into storage systems to handle tasks like vectorization, and collaborating with industry partners through NVIDIA's blueprints to build turnkey AI-ready infrastructure. He also introduces new memory tiers, including G3.5 for KV cache, to improve AI inference efficiency by reducing recomputation.

Suggested questions

3 ready-made prompts