Samsung Electronics has begun mass-producing its 10th-generation V-NAND, "V10," and has started supplying it to NVIDIA, the Seoul Economic Daily reported on July 20. At the center of this news is Context Memory Storage (CMX), which NVIDIA is preparing for long-context AI inference. CMX offloads KV cache that doesn't fit in the GPU's HBM to network-attached flash, then brings it back and reuses it when needed. What's now being questioned for AI NAND isn't whether SSD capacity can be increased, but how far high-density products can continue to be supplied as part of inference infrastructure.

NVIDIA positions CMX as a shared context layer for long-context and agentic AI inference. If Samsung's V10 supply proceeds as reported, the company's NAND will expand its use case from storage for PCs and general-purpose servers to a tier that affects GPU utilization and inference costs.

AD

CMX: Placing KV Cache Outside the GPU

CMX, as officially described by NVIDIA, is a pod-level shared context layer for long-context, multi-turn, and agentic AI inference. The KV cache it handles is intermediate state that language models retain so they don't have to recompute tokens they've already read. The longer conversation histories grow, and the more agents reference the same context, the larger this volume becomes.

Traditionally, KV cache placed in GPU HBM or server DRAM has to be either discarded once capacity runs out or recomputed every time it's reused. According to NVIDIA's technical blog, CMX inserts an Ethernet-connected flash tier in between, pre-staging KV blocks to GPU or host memory before decoding. The inference frameworks Dynamo and DOCA Memos are designed to manage which node's cache to use.

BlueField-4 handles NVMe SSD management, KV cache integrity protection, and encryption offload. NVIDIA's target is up to 5x improvement in sustained token processing performance and power efficiency compared to conventional storage, for long-context and agentic workloads. This isn't a figure that applies uniformly to all inference, but the intent to reduce KV cache recomputation time and GPU wait time is clear.

Where V10's Density Matters

The Seoul Economic Daily reported that V10 is a roughly 400-layer product, with over 50% higher storage density than the 286-layer V9. An increase in layer count works to increase the amount of KV cache that can be retained within the same SSD form factor, rack footprint, and power and network budgets. For inference handling long contexts simultaneously, this increment translates into more candidates that can be returned to the GPU.

Samsung officially announced the start of mass production of 1Tb TLC 9th-generation V-NAND in April 2024. This latest report indicates that after V9 became the mainstay product, the company has progressed to mass-producing the higher-density V10 and supplying it to NVIDIA. Neither NVIDIA nor Samsung has announced a supply contract for V10 or adoption volumes for CMX, so this remains within the scope of the report.

However, layer count and capacity alone don't determine CMX performance. NVIDIA's CMX operates SSDs together with BlueField-4, Spectrum-X Ethernet, and KV cache placement software as an integrated whole. High-density NAND is the component that increases retention capacity, but whether it can be read with low latency and returned to the GPU depends on a design that includes the network and control systems.

AD

Launching V10 Within a V9-Centered Production System

According to the newspaper, Samsung's V-NAND production capacity exceeds 100,000 wafers per month, with about 60% allocated to V9. V10's share is under 5% of total V-NAND production, while V9's yield has reportedly risen above 80%. Even if supply to NVIDIA expands, the current setup doesn't yet look capable of meeting demand with early-stage V10 production alone.

While V9's yield has risen above 80%, V10's share remains under 5% of total V-NAND output. This is a stage where the mainstay V9 is being stably produced while V10 is being launched, and whether high-density product supply can catch up with demand depends on how far V10's production share and yield can rise. V10 expands the room to configure the same capacity with fewer dies, but actual SSD capacity, controllers, and product configuration also depend on customer-side design.

NVIDIA has not officially disclosed which NAND generation, or how much of it, will be used for CMX. Therefore, the changes confirmable at this stage are twofold: the report that Samsung has begun supplying V10 to NVIDIA, and the fact that NVIDIA has added a dedicated flash tier for KV cache to its product lineup. Rather than jumping ahead to conclusions about supply volume, we need to watch how far V10's production share grows in coming quarters.

Demand Moving From 35 Million TB Toward Over 100 Million TB

The Seoul Economic Daily reported an industry outlook stating that NAND demand required for CMX will grow from 35 million TB in 2026 to over 100 million TB in 2027. These figures aren't Samsung's shipment plans, but a demand outlook assuming CMX adoption spreads. Even so, it's clear that the unit of flash demanded by AI servers is shifting from single-SSD capacity competition toward pod-wide context retention capacity.

In this use case, flash with increased capacity doesn't automatically translate into value. CMX isn't a device for long-term storage of finished inference data; it's a tier for quickly returning KV cache to be reused in the next request. NVIDIA's stated performance claim of up to 5x also presupposes a configuration where cache reuse, pre-staging, and network connectivity are all aligned.

How much of its V10 output Samsung will allocate to CMX has not yet been officially indicated. Whether NVIDIA discloses actual CMX deployment sites and SSD configurations, and whether Samsung can raise V10's production share, will determine whether AI-oriented NAND demand becomes a new procurement standard for inference infrastructure, rather than a temporary supply shortage.