Powering Agentic AI with AI-Ready Data Platforms That Turn Data Into Intelligence
578 segments
Thank you. Dennis.
Good morning everyone.
Welcome to day three of GTC Taipei 2026.
I like that you're all the early birds. You got here first.
You get all the good content today.
So I have a bit of a tradition.
Every time I do this, I have to take a selfie with you all.
So pose, smile, jump, do whatever you want.
But on the count of three,
here we go.
One, two, three.
Alright.
Thank you.
Well good morning.
Like Dennis said, my name is Jason Hardy.
I'm the vice president for storage technology here at NVIDIA.
It's funny being the storage guy in a company that doesn't sell storage.
So it's been very exciting though.
And we're finding that data continues to play
a monumental role in the success of AI, especially as
we start to see agentic capabilities
really start to grab ahold and take place.
There's no denying that from an investment perspective, we're seeing a
continuation of investment into AI technologies
and the necessary infrastructure to be able to get to these outcomes, and this
will continue to happen for many years to come.
And we're really excited about this.
Now, AI though, inferencing on the agentic side
or just even simple engagement just with a chatbot.
It relies on a grounding
in the data that sits behind it.
The data is what describes the personality
of an enterprise, describes what we do and how we operate personally, but also
how companies, businesses, enterprises, how they exist,
what they do, how they operate.
All of that is grounded in data.
The problem that we're starting to see
is that in order for the success of AI
to really grab ahold, it's becoming
increasingly dependent on data.
And we see the facts and statistics around the AI projects
that are abandoned or that never see past proof of concept.
But ultimately, at the end of the day,
AI is as good as the data that's fed to it, that it's used
against it, that it's trained against.
And especially as we look at this agentic frontier that we're starting to
to embrace, and we're really starting to see some amazing capabilities to
come from, the requirements from a data perspective, especially on
the enterprise storage systems, the things that
have maintained that data for 50 plus years.
The dependency on that is changing those requirements.
It's changing how those systems are built.
It's changing how those systems operate.
It's changing the requirements and the services that even need
to come from those platforms and systems.
You see, when we look at an
enterprise, it's really made up of three different kinds of data.
There's structured data, while small in volume, very critical to
how an enterprise operates.
This is the ERP systems, the SQL systems,
the databases, that structured data
that business decisions are made off of.
Then there's the unstructured data side—your home directory. SharePoint.
Some people love it. Some people hate it.
OneDrive, or Google Docs, or NAS shares,
or PDFs, or video, or audio, or files.
That is unstructured data, and that is by far the lion's share,
the largest volume of data in an enterprise.
This also is the hardest to process
because it is made up of
so many different personalities and so many different data types
and ways to engage with it, that it is the hardest, but the largest
volume of data that describes an enterprise.
And now we have this
third category, this third classification of data.
This is new.
This is what's come from the agentic
push and and the embracing of this technology.
And this is context. Agent context.
Some people know context as KV cache.
It's actually more than that.
KV cache is a piece of it.
But it's also the memory, the
data that gets generated as agents operate, and as they
consume information that we feed them and how they
and the task at hand and everything, that's the data that the agent
is responsible for consuming, whether it be what we feed it
and it activates off, of or what it generates itself.
And even, at times, shares from agent to agent.
This has to persist over long term.
Some agents live for hours; others live for years.
So this data becomes as critical
as the rest of the enterprise ecosystem— as critical as what the structured data
that we continue to maintain
with the tools and applications, or even the unstructured data
that we generate on a daily basis as we just go about our day, our job.
So across these three landscapes,
this data here is the foundation of what enterprises are built on.
But you see, from an agent demand perspective, there is a
new paradigm that's happening,
especially for storage.
You see, traditionally we had stateless environments.
We had the structured data, we had a user doing a thing, and
it generated a workflow.
And then we also had kind of two states for data to manage that—data at rest.
What's living there?
It could be an archive. Or data in motion.
That active data, that data set that we operate against.
We're modifying or manipulating.
Now, as we shift into the agentic frontier, though,
it changes this whole paradigm.
You see, now we aren't just stateless where it's short lived.
Now we have long-running sessions.
Who here has used agents in one way, shape, or form?
You guys really need...
You're at an NVIDIA conference.
Come on.
You should you should be using this stuff.
No, these tools are amazing.
And in fact, even as I use my own agents, I have the liberty
to give it a task, to go to bed, to wake up,
and it's worked all night.
That task is done,we review, and then we iterate from there.
So it's no longer just... I end my day, and then data is done.
I end my day, and it continues.
I'm also blending unstructured
and structured data together for these agents to operate.
So it's as much SQL as it is a PDF.
And then
it's not just me engaging with my data; it's
the one, the tens, the hundreds, even the thousands
of agents that will continue to iterate over
data, continue to consume it, and it will consume it, not at the rate
at which a human interfaces with data, but at the rate
at which agents can start to interface with
both the data and each other.
And trust me, it's a lot faster
than you reading a document.
So we have scale.
We have new data types.
We have data that's active 24/7.
And this growing context of needing
to keep both, yes, KV cache, but also all of the memory, and the scratch,
and the to-do, and the embeddings, and everything that goes on
to build context around how agents operate.
Everything that agentic workflows
are starting to push into, that we're starting to embrace on a day-to-day
basis, is changing the landscape of what storage actually means.
And so storage needs to evolve.
It needs to move from being what we've always known
as a system of record, to now
a system of context.
Now, what makes this hard for the enterprise?
Obviously, things change very quickly.
So we need to be able to do things at scale.
We need to be able to scale cache.
We need to be able to scale those embeddings.
That context.
Everything about how agents operate brings more scale,
performance, and density.
We also need to introduce new technologies into the enterprise.
Things like RAG, or retrieval-augmented generation, or embeddings,
and the ability to take in all of that data and embed it into a vector database.
How do we chunk that up?
How do we operate with that?
We need to scale that workload
so that our agents are fed at the speed at which they operate.
And we cannot forget about the security, the governance, and
even how enterprises operate today,
from a hybrid perspective.
You see, enterprises aren't just in the data center, and they're not
just in the cloud. They're on both.
So in that is, how do we best enforce
and represent the data where it needs to be on prem, cloud, wherever, but also
ensure that we do this securely, and we do this in a controlled manner,
do this in a way that we can understand and audit, if ever necessary.
So we still have many enterprise requirements that need to be maintained,
but also need to evolve forward to address the agentic demand.
This is why we believe that storage is moving
from that system of record to, really, the new data plane.
And it's the new data plane for agentic AI.
You see,
much like Jensen has talked about the
AI layer cake, the five layers of AI,
storage has much the same, but there's three of them.
So maybe it's a cupcake.
The cupcake of storage.
But what it means is it's based on high-performance requirements,
still having enterprise data,
but then how do we bring agentic context into that?
When we look at it from a storage perspective,
just at the raw capability to introduce and support, or to be able to support
the high-performance requirements, those continue to scale upward, in both
how do we train and inference, but also how do we bring more context and volume?
How do we support inferencing at that scale, from a performance and demand?
This is why we're extremely excited that Vera is here.
You see, Vera was designed
to take stuff off the wire—
information, bits and bytes, packets, process it through, and then give
it to the GPU as quickly as possible.
So the 3x memory bandwidth plays a big role in being able to do that.
The single core... single threaded performance improvements,
best-in-class, is great for the GPU.
Well, you know what else it's great for?
Storage.
Vera, through the capabilities that we built to be able to drive
the GPU, benefits the same to storage.
So as we look at now introducing
Vera to the market...
The NVIDIA Vera
Bluefield-4 STX platform, our opinionated design
of what storage looks like, benefits from all of those capabilities.
It's the foundation for how we'll continue to build the AI factory, and in working
with our partners. It's how we will enable storage to scale
to meet the demand from an agentic perspective,
at the very base of high-performance
storage and object storage
and all these storage technologies to be able to drive that forward.
It's how we bring Vera. It's how we bring our software.
It's how we bring all the greatness of the NVIDIA
ecosystem, plus our storage partners' capabilities.
Bring those together, and release that into the market to be able
to address this demand from the agentic workload.
Second, though,
it's not just an infrastructure thing.
It's also about how we prepare the data, how we govern the data, how we bring
lineage and indexing and everything necessary
for enterprise data to be agentic ready.
This second layer of the storage cupcake
is about addressing the challenges to make sure
that we don't become another headline.
You see,
if you don't address the data accurately,
if you don't bring the right data to the context,
things can go a little sideways.
Air Canada got in trouble with the Canadian government
because they had a chatbot hallucinate new refund policies.
Didn't exist. The courts didn't care.
United Healthcare got sued for faulty AI, and in that,
had to pay out a significant amount of money.
Google had to play a tremendous amount of money
because of wrong answers, because of wrong demos.
The same with Deloitte, and same with on and on and on and on.
This isn't an AI issue.
This is a data issue.
And this is why data needs to be especially
prepared to be able to address the AI demand.
Now, it's complicated to do this.
There's no denying that.
First, you have to ingest your data.
And this could be whatever type of data.
It could be PDFs, it could be tables, it could be structured
data, it could be audio, it could be video,
It could be very multimodal.
The last time you opened up a PDF, when was it just text?
Never.
So you need to understand not just that there's text in the document, but you have
to be able to understand and process through the text
and the tables and all of the imageries
and how it relates to each other.
So you go through this ingestion process, then you have to cleanse your data
so that it's ready. You have to remove duplicity.
You have to remove or normalize the data.
Sometimes you have to strip the personable... the PII, the
sensitive data out of it, and then you want to enrich it.
You want to bring metadata to it.
That way, you have data that describes the data.
Then chunk it. Break it up.
Make sure it fits into these semantic units, sized appropriately for
the model that you're running.
And then you embed it into the vector database.
And then you index it. And then you govern it on top of it.
It's not straightforward.
We realize that for data to
be properly processed for AI readiness,
it takes work, and it's complicated.
You see, agents need accurate data.
And unfortunately, data has all of this complexity with it.
It has pipeline complexity.
How do we how do we use it,
even in our business? It has structured complexity.
How do we associate, and again, in a PDF, this graph and this text
and this... all of this stuff, how does it correlate with each other?
And then data changes over time.
It's not static.
Just like our businesses change over time.
And then, data has gravity.
Data has mass.
Turns out that the best bandwidth in the world is actually FedEx.
Go Google it.
It's kind of a meme.
But really, when we talk about data,
data is a complex ecosystem that continuously changes.
Just like we do. Just like our enterprises do.
So how do we bring this data to agents
in a way that's accurate for them to be able to consume this?
And how do we do this to make
sure that we de-scope as much risk as possible?
A radiological diagnosis— if something goes wrong,
that can impact a patient health.
Or how do we use this technology now
to reduce the errors, from a radiology perspective?
Or how do we use this data to identify where
we can reduce theft inside of the retail space?
Or how can we use this data to improve
and drive on first-call resolution from a call center?
See, agents, while there's risk with them,
also bring a tremendous benefit.
Better health care, better financial state,
better identifying of theft and loss.
Improved profit. Everything about it.
You see, this agentic outcome, when we feed it
accurate data, can give us significant upside.
And this is why
we're super focused on aligning the data
and storage ecosystem to transform,
to deliver AI-ready data.
As the agentic harness that Jensen talked about during this keynote, as it starts
to process data, and as agents create that context
and go through this loop, and then act on the end of it, and we feed
more data in, and tools and all the governance and security, and then
ultimately create that memory that it lives from—
All of that requires a new paradigm, a new approach
for the enterprise data ecosystem
to be able to address this demand. We need to be able to have better search.
We need to bring intelligence into it.
We need to be able to identify whether it's structured or unstructured.
How do we access it?
We need to ensure that our agents have access to only
the data they're allowed to.
All of that comes from this new data plane that is the storage platform.
So as we now look at this, it's still about
having enterprise storage systems that have protocols
and file systems and orchestrate data and snapshots and cloning and everything
that's necessary from an AI data infrastructure perspective.
But it's more than that.
It's about how we bring now further capabilities
and how we manage that data. How we govern it.
How we bring deduplication and compression. How we bring keyword search
into it, better access control, all the analytics necessary.
This enrichment starts to add functionality
into that foundation from an enterprise storage system.
Then how do we prepare the data for AI?
By bringing in that continuous ingestion,
extracting data actively as it exists.
Not later, not in chunks, not in a couple hours, but in real time.
How do we index it so that the agent knows about the data when it's created?
So it's making accurate information.
Storage needs to prepare data on creation.
And then manage it.
How do we bring the best management into that
so that we can redact data, mask it, classify it,
embed it, ensure that we're not exposing sensitive data.
Again, all at time of ingestion.
So all of this now—it's not just an enterprise
storage problem, but is an AI data platform capability.
And this is what enterprise storage
is becoming. The AI data platform.
And we at NVIDIA have been working with our partners
in building exactly that—the AI data platform.
It's about bringing that readiness
at time of creation, at the source of it, from a CPU perspective starting,
but including, now, the ability for GPUs to be added
in to the enterprise storage ecosystem so that now we can have all of that
amazing functionality for vectorization, all the embeddings, all the even
inferencing at that data, directly inside the storage platform.
And to help this along, it's not just
saying, here's a vector database, here's a GPU, Mr.
Storage Partner, go take care of it.
But we've actually been building blueprints that help simplify this.
This is our RAG blueprint for
for helping and streamlining the storage
system capability to vectorize and embed and do all the work necessary
to prepare data within the storage platform
for agentic AI readiness.
The key here is, is that this is done at the point of creation
where the data exists, not decoupled somewhere else.
Or what we're doing with
VSS or Video Search and Summarization blueprint.
Where now, agents have the ability to understand video content
because we can summarize it.
We can we can capture off of it.
We can understand the context of the video.
We can process it, all at the point of creation in the storage system.
Now, the great thing is
we've been talking about this for a while, and our partners
have been working with us very closely.
So it's not just about an idea that we have, but that they've
put this into execution—
that Dell with their storage capability
or a NetApp, or IBM, or the list goes on and on, have all taken this capability
and have embedded it into their storage systems
to create an in-place AI data platform solution
that's embedded and a part of their storage ecosystem now.
So we're super excited about this because this now starts to prepare data
and get it ready for, again, at the time of...
at that point in time of creation, so that
our agents can now have better context, our agents can
make better decisions, all based and grounded on data
that's been prepared for them in a way that they understand.
So that idea
of understanding data
at that time of creation so that my agent can be more effective—
it's not just an idea. It's being put into execution.
So now when we look at it, every enterprise, every AI factory
needs an AI data platform.
You see, an AI data platform—
you could think of it like
the raw goods for an AI factory.
Just like if we have a factory that builds stainless steel or whatever,
and you take raw iron ore, you feed it
through the factory to create that stainless steel at the end.
I think that's how stainless steel is made.
The same goes for the AI factory.
It's about taking that data,
preparing it at the edge at the time of creation
in the storage system, and taking that refined good
and bringing it to the AI factory
so that it can manufacture intelligence,
grounded on how the enterprises operate.
So now, it's how the AI factory becomes more effective, is by grounding
it in data that AI data platforms have created or have refined, so
that its accuracy and its intelligence is as high as possible.
But every AI data platform—
it still needs to be real-time and accurate.
It still needs to continuously index.
It's not one and done.
It's about continuing to improve on itself as well.
Not just saying, hey, I have one indexing model or one RAG model.
It's I've got different ways of understanding this
and evolving forward to make it enterprise-grade.
It's also how, again,
we introduce data, new data types into the enterprise, all the time.
How do we make sure we understand that?
How do we make sure we grasp that?
How do we make sure that we
incorporate that into the model so that our agents can understand new data
with new context, to be able to make, again, a more informed decision?
How can we discover data quickly? New data.
And then, how do we make this turnkey?
How do we make this simple?
And this is really important because this is also
where the enterprises have a lot to do.
And it's very complicated to do this stuff.
So we need to make it easy.
And that's what we're doing with our storage partners—
the Dells, and the IBMs, the HPEs, and the NetApps—making this turnkey
so that it's fully integrated in; it has all the networking, the compute, the software,
everything necessary so that a consumer of it, a user of it, whether
it's an agent or a human, is able to just flip the switch.
Now, it's a little bit more complicated than that, but it really
is about making it that easy or as easy as possible
and continuing to bring in the security and the compliance and the governance
so that it still maintains that trust.
Because at the end, it's about trust.
And by doing this and making it simple, we also don't want to compromise
on the trust of the data.
And, again, this is why we're really excited about
what our partners are creating inside the AI data platform ecosystem—
consuming our blueprints, driving forward, and bringing a lot more value
in different data types and different capabilities,
directly into the storage ecosystem.
The last layer of the cake,
the cupcake. The frosting.
Context.
Addressing context is not easy.
In fact, addressing context
required us to think about the structure of memory inside of a data center,
inside of an AI factory, inside of a GPU.
So we created this kind of hierarchy of memory here
that allows us to describe
what type of memory lives where.
G1 is HBM. It's the fastest memory.
It's directly in the GPU. It's where the active KV live.
It's where the weights get loaded.
Everything about the model and inferencing
happens in G1. But G1 is finite.
G2 is system memory.
That tightly coupled memory, where both the HBM, G1, and the system, G2,
can collab or can work together in
as fast... as fast as possible.
This is where KV cache spillover happens.
This is where staging happens for models, for inferencing,
and all sorts of different things.
G3 is scratch, that NVMe disk in each server.
It's where the OS lives, simple things like that.
And then G4 is high-performance storage, where models live, where checkpointing
happens, where inferencing takes place.
G4 even extends into enterprise storage.
Now, when we were looking at context,
what we found is that G4
wasn't quite fast enough, and G3, it wasn't quite enough of it.
So we created G3.5.
And this is where KV cache lives.
Now KV cache is really important because when we are inferencing,
we have to calculate
a lot of information on the prefill and the decode stages.
This is how inferencing happens.
And the thing about KV cache is, when you cache
that information, you don't have to recalculate it, which is great.
That means the GPUs can spend more time working on new data,
creating new tokens, instead of having to recalculate, which is awesome.
That means that we get the best out of the factory.
So in order to address KV cache, we needed something
that was fast enough to support the model or the GPU
without having to recalculate, but also was cost effective and didn't
didn't create too much burden to be able to do that inside the factory.
So from our perspective, KV cache,
this is where we've released CMX,
or KV cache storage capability.
This new tier of storage ensures that we can offer
tokens, offer cache to the GPU as quickly as possible, but without the
extra burden, without being too inefficient,
All of the great stuff that's necessary to make this effective.
That addresses KV cache, though.
Like I said earlier, there's another type of context—
this memory context, this how an agent operates and
the data it generates when
consuming data, and when working with other agents, and everything about it
that's not just tokens.
That cache is...
Jensen talked about it earlier during the keynote was, it's the ontology.
It's how agents' information and how we are described as
as users of this, that... how that permeates through the system.
That context,
we're still doing a lot of work on, but we do know that we can make
the factory as efficiently as possible by bringing in this KV cache concept.
So this new tier of storage, again, in working with our partners, allows for
better scale of the AI factory,
from a token throughput perspective. More tokens. More intelligence.
It also scales to meet the demand of as the factory grows. Start small, go big.
Support, again, that workload as we continue to scale the factory size.
It's also designed to
complement both GPU memory and network storage by sitting in between.
It doesn't replace HPS or G4.
It augments it.
It brings the right type of storage for the right platform
or the right requirement.
So KV cache is really designed to address that
acceleration and that token-per-second output from the AI factory.
But again, this is designed not by us, but with our partners.
So we'll continue to go through this with them.
We'll continue to address this.
And you'll start to see this permeate the AI factory, because there is
value in being able to manage KV cache more appropriately,
to reuse it as much as possible so that the GPU doesn't
spend time recomputing when it can create new.
So when we look at the challenges
of the data center, now, when we look at the challenges
of the data ecosystem now, it's not about just bringing storage in.
It's not about just bringing a GPU in.
It's about extreme co-design across both the stack horizontally...
I should say vertically and horizontally.
It's how we bring network and how we bring GPU
and how we bring compute, especially with Vera now, it's how we bring Spectrum-X in.
All of this working together to accelerate
and provide as much value as possible, and then working with our partners
on the horizontal perspective to be able to create new capabilities
inside their portfolios.
It's about addressing the agentic AI demand
through the need for the AI data platform.
That scenario, whether it's or the CV cache scenario,
or even just high performance, just the parallel
file system, high-performance storage requirement.
This here is NVIDIA Bluefield-4 STX.
NVIDIA Vera Bluefield-4 STX.
This is our design.
This is our collaboration with the storage
ecosystem to address the agentic AI demand.
And we're really excited about the capabilities, espeically
like I said earlier with what Vera is bringing into the market, the capabilities
that it's going to be able to allow
for our storage partners to accelerate with and drive additional context
and additional value, and then bringing the GPU in with Rubin
or with RTX Pro, to be able to drive that
from an AI data platform perspective.
All of this together is reinventing the storage ecosystem,
so that, again, it's agentic AI ready.
Ask follow-up questions or revisit key timestamps.
In this presentation, Jason Hardy, VP of Storage Technology at NVIDIA, discusses the evolving role of storage in the era of agentic AI. He highlights how the rise of AI agents necessitates a shift from traditional 'system of record' storage to a dynamic 'AI data platform' capable of managing structured, unstructured, and context-heavy data like KV cache. Hardy emphasizes the importance of preparing data at the point of creation, integrating GPU capabilities directly into storage systems to handle tasks like vectorization, and collaborating with industry partners through NVIDIA's blueprints to build turnkey AI-ready infrastructure. He also introduces new memory tiers, including G3.5 for KV cache, to improve AI inference efficiency by reducing recomputation.
Videos recently processed by our community