On October 8, 2026, Kioxia announced a technical demonstration of an upgraded version of its GPU-oriented SSD, the KIOXIA GP Series. The demo is planned to show performance above 20 million IOPS in 512-byte random reads. The upgraded drive uses third-generation XL-FLASH, a high-speed flash memory, and more than doubles the read performance of the GP1, which was announced in August with up to 10 million IOPS.
The aim is not faster transfers of large files, but a greater ability to read large numbers of small pieces of data that a GPU needs. In the concept of supplementing HBM (high-bandwidth memory), the high-speed memory used in AI systems, with flash storage, what matters is not only the performance of the memory chips themselves but also how the GPU reads data. This announcement shows how far the technical development toward that goal has progressed.
A technical demo of 512-byte reads exceeding 20 million IOPS
Kioxia plans to show the upgraded GP Series at the OCP Global Summit, to be held October 12–15 in San Jose, California.
According to the company's announcement, the upgraded drive uses third-generation XL-FLASH and will demonstrate performance exceeding 20 million IOPS in 512-byte random reads.
IOPS (Input/Output Operations Per Second) is a metric for the number of input/output operations a storage device can process per second. It is an especially important measure of storage performance in workloads that frequently read and write small pieces of data.
The point of comparison is the GP1, whose development was announced on August 4.
The GP1 combines second-generation XL-FLASH with PCIe 6.0 and was said to achieve up to 10 million IOPS in 512-byte random reads.
With the upgraded version, read performance at the same 512-byte size is more than double.
| Comparison | GP1 (announced in August) | Upgraded GP Series (announced in October) |
|---|---|---|
| Flash memory | Second-generation XL-FLASH | Third-generation XL-FLASH |
| 512-byte random reads | Up to 10 million IOPS | Technical demo of over 20 million IOPS announced |
| Disclosed development/availability status | Evaluation samples to limited customers planned by the end of 2026 | Technical demo planned at OCP Global Summit |
However, these are not independent benchmark results measured on the same equipment or under the same conditions.
The announcement is also a preview of a technical demo; it does not announce mass production or an official commercial launch of the upgraded SSD.
So what does 20 million IOPS mean in terms of the amount of data actually read?
If each read transfers 512 bytes and 20 million reads are processed per second, the amount of data read is as follows:
Using the decimal convention of 1 GB = 1 billion bytes, this is equivalent to 10.24 GB/s.
Because the announced performance exceeds 20 million IOPS, the same conversion would put the data volume above 10.24 GB/s.
However, this figure is calculated from 512-byte read requests; it is not a measured sequential read speed of the SSD. It also does not include incidental data for communication and control.
What the figure shows is a high capacity for reading large numbers of small pieces of data.
A high IOPS count also does not necessarily mean that each individual read completes quickly.
Because an SSD can process multiple read requests in parallel, the response time for each request cannot be derived from the number of operations per second alone.
To evaluate its practicality as GPU storage, read latency needs to be checked in addition to IOPS.
Third-generation XL-FLASH: 3x read performance for the memory, over 2x for the SSD
In Kioxia's comparison material on third-generation XL-FLASH, read performance is said to improve by 200% and write performance by 150% compared with the second generation.
Taking the second generation as the baseline, this means reads are 3 times and writes are 2.5 times as fast.
However, these figures and the 20-million-plus IOPS shown by the upgraded GP Series measure different things.
The former indicates the performance gain of XL-FLASH itself, while the latter indicates the random read performance of the SSD as a whole.
The comparison material also does not give detailed measurement conditions or evaluation methods. It is therefore impossible to judge from these figures alone how much the third-generation gains translate into improvement for the SSD as a whole.
XL-FLASH, to begin with, is designed to prioritize low-latency data access, unlike general NAND flash, which prioritizes high capacity.
According to Kioxia's product description, it achieves fast reads by shortening the word lines and bit lines connected to the memory cells and by optimizing peripheral circuits.
Another characteristic is the configuration of "planes," which divide processing inside the memory.
Existing XL-FLASH products use 16 physical planes. Planes divide the memory into regions that can operate independently, and they are distinct from the number of NAND flash layers.
However, the number of planes and the detailed internal structure of the newly announced third-generation XL-FLASH have not been disclosed.
The structure of existing products is a useful reference for understanding how XL-FLASH operates quickly, but the third generation does not necessarily use the same internal configuration.
Caution is also needed regarding 512-byte reads.
The 512-byte figure supported by the GP Series represents the size of data requested from the SSD by the host.
Meanwhile, Kioxia's product description for existing XL-FLASH states that the physical page size inside the NAND is 4 KB.
In other words, even if the host can request data in 512-byte units, the NAND does not necessarily read only 512 bytes internally.
The specifications alone do not allow us to conclude that all processing involving the internal reading of extra data has been eliminated.
On power efficiency, the third-generation XL-FLASH comparison material says it improves by 40% for reads and 80% for writes compared with the second generation.
However, improved power efficiency is not the same as reduced power consumption.
For example, if more work can be done with the same power, efficiency improves even when power consumption is unchanged.
Therefore, the performance gains and power-efficiency improvements in this material do not reveal how many watts the SSD as a whole operates at, or how much energy each read consumes.
For data center adoption, measured data on the SSD's total power consumption and heat output would be needed.
Direct GPU data reads: Storage-Next and SCADA
The GP Series aims at more than simply building a fast SSD.
It aims to build a storage environment that can supply the data a GPU needs more efficiently.
In current AI systems, the HBM mounted on GPUs plays a key role.
HBM stores a variety of data, including not only AI model weights but also the KV cache that holds intermediate results during inference.
When running large models, processing long texts, or increasing the number of simultaneous users, how to use the limited HBM capacity becomes an issue.
In Kioxia's technical explanation, the concept of supplementing HBM with high-speed flash storage is introduced as the background to developing the GP Series.
If frequently accessed data is kept in HBM and other data can be read from fast SSDs as needed, the amount of data available to the GPU could potentially be increased.
However, simply increasing storage capacity does not improve GPU processing performance.
Actual performance varies greatly depending on which data is placed in HBM and which is read from the SSD.
If the wait for reads from the SSD is long, the GPU's compute capability may not be fully used.
What becomes important, then, is a mechanism for GPUs to access storage efficiently.
NVIDIA's "Storage-Next" is an industry-wide effort to realize such new storage access methods.
According to NVIDIA's announcement about FMS in August 2026, more than 40 companies, including Kioxia and Micron, are participating.
One of its central technologies is "SCADA."
SCADA is a software foundation that uses the GPU's parallel processing capability to efficiently bring data needed by applications from storage into GPU memory.
An important point here is that "transferring data directly to GPU memory" and "the GPU itself issuing read requests to storage" are different things.
With conventional GPUDirect Storage, an extra copy through CPU memory can be avoided when transferring data from storage to GPU memory.
However, the driver responsible for tasks such as setting up the transfers continues to run on the CPU.
In contrast, Storage-Next aims for a mechanism in which the GPU can issue large numbers of small read requests itself.
Kioxia's GP Series pursues high random read performance to support this kind of fine-grained, GPU-driven storage access.
In NVIDIA technical material published on September 30, the "SCADA Server SDK" was introduced.
This SDK is for developing storage servers that respond to requests sent from SCADA clients on the GPU side.
IBM has also demonstrated a prototype system combining SCADA with IBM Storage Scale.
In other words, while Kioxia raises SSD random read performance, NVIDIA and storage-related companies are building the software and communication infrastructure to use that performance from the GPU.
However, even with a mechanism that lets GPUs access storage directly, the roles of the CPU and OS do not disappear entirely.
SCADA separates the part that handles high-speed data processing from the part that manages access permissions to storage.
Even if a GPU can issue large numbers of read requests, a mechanism to prevent access to unauthorized data is still needed.
Kioxia's announcement does not say whether SCADA will be used in the technical demo of over 20 million IOPS.
Which GPU will be used and what software configuration will be used to measure performance have also not been disclosed, so the broader technical development of Storage-Next and the evaluation environment for this SSD should be considered separately.
Targeting 100 million IOPS in 2028, and the challenges to productization
Kioxia also plans further performance improvements for the GP Series.
In the company's roadmap material, a concept is shown that combines third-generation XL-FLASH with PCIe 7.0 and aims for 100 million IOPS in 512-byte random reads in 2028.
This is roughly five times the 20-million-plus IOPS announced this time.
However, the material describes this concept as "under consideration," and product specifications and shipping timing have not been formally decided.
Also, even if the upgraded version and the next-generation SSD targeted for 2028 both use third-generation XL-FLASH, they are at different stages of development.
The generation number of the flash memory also does not necessarily correspond to the generation of the SSD product.
Kioxia has not announced the upgraded version as the "GP2," nor has it disclosed the connection interface.
Caution is also needed so that the previously announced figure of 100 million IOPS is not confused with performance actually achieved by an SSD.
At FMS 2025, Kioxia introduced a method of advancing software development and verification using an emulator that simulates a future high-IOPS environment.
The 100-million-IOPS-class processing performance shown at that time was the result of the emulator and does not indicate that an actual SSD achieved the same performance.
The 2025 emulator-based verification, the 20-million-plus IOPS technical demo announced now, and the concept of 100 million IOPS targeted for 2028 should each be understood as results at different stages.
Kioxia also plans to exhibit its high-capacity LD4 SSD at the OCP Global Summit.
According to the LD4 announcement, it uses 8th-generation BiCS FLASH QLC NAND and is being developed as a high-density data center SSD intended for read-centric workloads.
Samples will be offered in 15.36 TB and 30.72 TB capacities, and the validated architecture can scale capacity up to 122.88 TB.
However, 122.88 TB is the architecture's scalable capacity and does not mean product samples of that capacity are being provided now.
The GP Series and LD4 also emphasize different kinds of performance.
The GP Series pursues high IOPS for reading large numbers of small pieces of data, while the LD4 aims to be high-density storage that can hold large amounts of data.
Both are SSDs for AI infrastructure, but because their uses differ, their performance cannot be simply ranked against each other.
To judge whether the upgraded GP Series can be deployed in real AI systems, more detailed performance data will be needed.
Particularly important are the number of concurrent read requests at which over 20 million IOPS was achieved and the response time for each request.
Whether response delays can be kept low even when large numbers of requests concentrate on the storage will affect GPU utilization.
Performance under mixed read and write loads, as well as SSD capacity and power consumption, also need to be checked.
Once power consumption and heat output are known, it will also be easier to estimate the number of SSDs a data center needs and the burden on cooling equipment.
Furthermore, to take advantage of this hardware performance, combination with compatible software such as SCADA is also important.
The announcement of over 20 million IOPS shows that GPU storage is getting faster, but how much that performance will improve real AI processing is still unclear.
If concrete measurement conditions and compatible software are published, and it is confirmed that GPU wait times can be shortened in real AI workloads, the concept of using high-speed flash memory as a supplementary tier to HBM will move further toward practical use.
Only at that stage will it become possible to concretely evaluate its advantages over conventional memory configurations in terms of both performance and cost.
