• What happened: On September 29, NetApp announced "NetApp Novus," an AI storage architecture that separates metadata processing from the reading and writing of the data itself.
  • Why it matters: Because the ability to handle large volumes of file operations and the bandwidth for moving large amounts of data can be scaled independently, Novus aims to reduce the time GPUs sit idle waiting on storage.
  • What to watch next: The 100TB/s figure is not a value measured on a large-scale production system; it includes a projection based on scaling tests. The key questions are how Novus performs in real AI environments when reads and checkpoint writes overlap, and how far the certified configurations offered at launch can scale.

At "INSIGHT 2026" in Las Vegas on September 29, 2026, NetApp announced "NetApp Novus," a new storage architecture for AI.

Novus separates metadata processing, which manages information such as where files are located, from the processing that reads and writes the data itself, so each can be scaled independently in large GPU environments.

The main targets are "AI factories": neoclouds that operate large numbers of GPUs, GPU-as-a-Service providers, hyperscalers, and similar operators.

NetApp touts scaling to read bandwidth beyond 100TB/s, but top speed alone is not the point. Novus is designed so that, in an AI infrastructure where huge numbers of small file operations and large checkpoint writes occur at the same time, neither workload gets in the way of the other.

AD

Separating metadata processing from data transfer

AI model training generates storage workloads of very different natures at the same time.

Loading training data requires opening large numbers of files and checking their locations and attributes. This generates a flood of fine-grained requests to "metadata," which manages things like file names and placement.

Checkpoints, which save the state of training partway through, can involve multiple compute nodes writing terabytes of data all at once.

In the official blog, NetApp's Arindam Banerjee explains that handling both on the same controller causes them to compete for performance.

Even if you add large numbers of SSDs to increase capacity or peak bandwidth, if the process of looking up file locations gets backed up, delays occur before data can even be delivered to the GPUs.

In other words, storage capacity, data transfer bandwidth, and the number of file operations that can be processed concurrently need to be considered as separate performance dimensions.

In Novus, metadata processing is split out into software called "Novus Data Director."

The data itself is stored on data nodes that use NetApp's storage OS, ONTAP. In the initial configuration, Data Director runs on certified Supermicro servers, and the data nodes use the NetApp AFF A90.

The product brief shows a setup in which a client on the GPU server first asks Data Director for file placement information, then accesses the data nodes directly once it has received that information.

Because the data itself does not pass through Data Director, metadata processing capacity and data transfer bandwidth can each be scaled independently.

Novus also combines multiple storage systems into a single "namespace."

A namespace is the file path structure as seen by users. Rather than mounting each added storage system as a separate destination, the aim is to increase capacity and bandwidth while presenting the same file system.

Applying existing pNFS to large-scale AI infrastructure

The idea of separating metadata access from data access is not new in itself.

IETF RFC 8435, published in 2018, standardized the pNFS (Parallel NFS) "Flexible File Layout."

With pNFS, a client receives file placement information from a metadata server and can then access data servers directly.

NetApp's existing ONTAP administration documentation also describes a mechanism for accessing multiple storage systems in parallel using pNFS.

So what is new about Novus is not that it invented a way to separate metadata and data.

Its novelty lies in offering a product configuration for large-scale AI infrastructure that combines a dedicated metadata layer with multiple ONTAP storage systems, each of which can be scaled independently.

On the GPU server side, it uses NFSv4.2 and pNFS Flex Files, which are built into the Linux kernel.

The Novus product page says it also supports nconnect, which widens bandwidth by using multiple connections, and GPUDirect Storage, which transfers data efficiently between storage and GPU memory.

One benefit NetApp emphasizes is that there is no need to distribute a proprietary Novus-specific client to GPU servers.

In environments running thousands or tens of thousands of GPU servers, distributing proprietary software to every node and revalidating it each time the Linux kernel or GPU generation changes is also a significant burden.

Using existing standard technologies is intended to reduce that work.

That does not mean it can be used without any configuration, however.

You will still need to verify supported Linux environments, configure networking, and set up the configuration needed to use GPUDirect Storage. The hardware offered initially is also limited to certified configurations.

The product brief also depicts the storage-side network used by Novus and the back-end network over which GPUs exchange training results as separate networks.

The path that delivers training data from storage to GPUs is distinct from the communication path used to synchronize gradients and other data between GPUs.

Therefore, when evaluating GPU wait times after deploying Novus, it is necessary to separate waits caused by storage I/O from waits caused by GPU-to-GPU communication.

AD

What does "100TB/s" mean?

For the 100TB/s figure NetApp touts, it is necessary to distinguish between what was actually measured and what was projected from it.

NetApp's press release includes an assessment by Tony Palmer of Omdia.

Omdia audited a test in which ONTAP clusters were added to a single namespace, and says it confirmed that performance grew nearly linearly as clusters were added.

A model based on those results predicts that Novus can scale to 100TB/s of sustained read bandwidth.

In other words, the 100TB/s figure is not a benchmark from running actual AI training on a completed system at that scale and measuring its performance.

The same goes for zettabyte-scale support: it indicates the range of scaling the architecture is aiming for, not that a zettabyte-class production system has already been confirmed.

The numbers on GPU counts also need to be read carefully.

In its official blog, NetApp explains that if each GPU requires up to about 2GB/s of storage bandwidth, supplying 50,000 GPUs simultaneously would require a total of 100TB/s.

The product page also lines up the figures "50,000+ GPUs," "about 2GB/s," and "100TB/s."

If total bandwidth were simply divided evenly among the GPUs, it would look like this:

Number of GPUs using bandwidth simultaneously (assumed) Total read bandwidth Bandwidth per GPU
50,000 100TB/s 2GB/s
100,000 100TB/s 1GB/s
1,000,000 100TB/s 0.1GB/s

Here, 1TB is taken as 1,000GB, so 100TB/s is converted to 100,000GB/s. This is a simple calculation that assumes all GPUs use the same bandwidth at the same time, and it does not reproduce an actual deployment environment.

NetApp elsewhere mentions a scale capable of supporting "hundreds of thousands" or even "millions" of GPUs, but this does not mean supplying 2GB/s to every GPU at all times.

The storage bandwidth required varies greatly depending on the model being trained, the data format, how caching is used, the frequency of checkpoints, and other factors.

Storage performance therefore cannot be evaluated by GPU count alone.

Furthermore, sustained read performance of 100TB/s alone is not enough to fully assess the design characteristics of Novus.

What matters is how much latency increases for each workload when opening large numbers of small files and writing large checkpoints are executed at the same time.

If it can be confirmed how far interference between metadata processing and data transfer is held down under such mixed loads, it will be easier to judge how much Novus's disaggregated architecture can actually shorten AI training time.

AFX and Novus separate different things

NetApp already has a disaggregated AI storage product called "AFX."

However, the "disaggregation" described in the official AFX architecture documentation targets something different from the separation in Novus.

Design difference AFX Novus initial configuration
What is mainly separated Storage controller nodes and the storage shelves that house SSDs Metadata processing and the storage that handles the actual data
Data access Controllers access a shared storage pool Clients receive placement information from Data Director and access AFF A90 directly
Approach to scaling Add controllers if you need processing power, shelves if you need capacity Add metadata processing capacity and data-side capacity and bandwidth independently

This table summarizes differences in design philosophy and does not compare the performance of AFX and Novus.

With AFX, the controllers that do the processing are separated from SSD capacity, and each can be added as needed.

Novus works at a higher layer, separating from data transfer the metadata processing that brings multiple ONTAP storage systems together into a single namespace.

Although both use the term "disaggregated," the bottlenecks they try to solve are different.

Novus is said to be available to order as of the September 29 announcement.

In the initial configuration, Novus Data Director runs on certified Supermicro servers, and the data layer uses the NetApp AFF A90.

Meanwhile, the product brief says NetApp is also considering a future software-defined configuration in which both the metadata service and the ONTAP-based data service run on certified third-party hardware.

Configurations available for purchase today need to be considered separately from those planned for the future.

What AI factory operators will want to check before deploying is not just maximum bandwidth but performance under mixed loads using their own actual AI workloads.

That means generating file-opening operations and checkpoint writes at the same time and measuring how much the time GPUs spend waiting on storage changes.

They will also need to check performance when multiple users and jobs use the storage simultaneously, as well as recovery after failures.

If Novus can reduce wait times that simply increasing maximum storage bandwidth could not eliminate, its disaggregated architecture could be a way to keep existing GPUs computing longer before adding more of them.