IBM's fifth-generation FlashCore Module (FCM5) also uses Everspin's magnetic memory, MRAM. At the SNIA Developer Conference on September 28, Steffen Hellmold, a fellow at Everspin, presented IBM's adoption as a case study. A board photo in IBM's product guide for the FCM5 also shows Everspin chips.
The FCM5 is an enterprise SSD announced in February, with capacities of up to 105.6TB. Built into it, alongside the high-capacity NAND, is a different type of memory that retains its contents when power is cut. IBM has used MRAM since 2018, so the notable point is not that it is a first-time adoption, but why IBM keeps using MRAM even as capacities have grown.
Why NAND being non-volatile still doesn't solve power-loss protection
NAND flash retains written data even when power is off. But not all write data that reaches an SSD is immediately saved to NAND.
It is necessary to protect not only data sitting in intermediate buffers, but also the management information that records which data was written where. Even if the data itself remains in NAND, if the update status and storage locations cannot be tracked correctly, the write state cannot be properly restored after power returns.
In designs that keep this information in volatile DRAM, the controller and memory must keep running from the moment a power loss is detected until the data has been flushed to NAND. Energy storage elements such as supercapacitors temporarily supply the power for this.
The longer the flush takes, the more stored energy is needed, and the more board space it occupies. In enterprise SSDs, this power-loss protection mechanism itself becomes a design constraint.
IBM explained this problem concretely in August 2018, when it moved FlashCore to the standard 2.5-inch U.2 form factor. With the earlier proprietary-form-factor modules, power for flushing data came from the system's batteries.
Once the design moved to U.2, power-loss handling had to be completed within the module itself. Moreover, the FPGA used as the controller in FlashCore at the time consumed a lot of power, making it difficult to flush all the DRAM contents to NAND in the short time the on-board energy storage could keep it running.
IBM's answer was Everspin's STT-MRAM.
Because MRAM retains its contents even when power is lost, it reduces the need to write all the critical information in DRAM out to NAND at every power loss. IBM explained that, in that design, the FPGA could be stopped within milliseconds after a power loss, making it manageable with the on-board energy storage.
In other words, the design changed so that it only needed to secure power for a brief shutdown procedure, rather than power to support a long data flush.
MRAM stores information using the magnetization direction of magnetic material. In Everspin's STT-MRAM, current is used to change the orientation of a magnetic layer in a magnetic tunnel junction. When it is aligned with the reference layer, resistance is low; when opposite, it is high. This resistance difference is used to read 0s and 1s.
It retains information after power loss while also supporting fast reads and writes, making it well suited to frequently updated management information.
In IBM's 2018 explanation, MRAM held journal checkpoints recording update status, as well as information for managing write data according to usage. At power loss, however, tasks such as sending a final command and closing pages in the middle of processing were still necessary.
Adopting MRAM does not make all power-loss protection circuitry and shutdown procedures unnecessary.
What do MRAM, SLC and QLC each do?
On April 30, 2024, Everspin announced that its 1-gigabit STT-MRAM, the PERSYST EMD4E001G, had been adopted for the FCM4. According to the company, the part has a DDR4 interface and 2.7GB/s of bandwidth for both reads and writes.
However, that is the specification of the MRAM part used in the FCM4. It is not a figure for the performance of the FCM5 as a whole, nor does it indicate the MRAM capacity installed in the FCM5. The FCM5 board photo cannot establish that the same part number is used or what the total MRAM capacity is.
Meanwhile, what handles the bulk data storage in the FCM5 is QLC NAND, which stores 4 bits per cell. QLC lends itself to high density, but fast writes and frequent updates require additional techniques.
In IBM's FlashCore, part of the NAND is used as SLC, serving as an area for newly arrived write data and frequently accessed data.
Organizing the designs IBM has published, MRAM, SLC and QLC each play a different role.
| Memory / area | How it stores data | Use described in IBM materials | Source material |
|---|---|---|---|
| MRAM | Retains information via magnetization state | Journal checkpoints; information about write streams | U.2 FlashCore design published in 2018 |
| NAND SLC area | Uses each cell as 1 bit | Temporary storage of new writes; fast access to frequently used data | 2026 FCM product guide |
| NAND QLC area | Stores 4 bits per cell | Dense storage of infrequently accessed data | 2026 FCM product guide |
This table organizes the roles based on IBM's design explanation published in 2018 and the "SLC and Smart Data Placement" section (pages 14–15) of the 2026 product guide.
Not everything about how MRAM is used inside the FCM5 has been made public. For that reason, MRAM uses disclosed in the past should be distinguished from the NAND data placement that is currently documented.
The SLC area is also part of the NAND, so data stored there is retained after power is cut. The difference from MRAM is that responsiveness for reads and writes is improved by how data is placed within the NAND.
So-called "pseudo-SLC" refers to a method of using part of NAND that can natively store multiple bits per cell as 1 bit per cell, like SLC. The storage element itself is different from MRAM.
IBM explains that frequently accessed data is kept in the SLC area, while less frequently used data is moved to the QLC area. However, not every write follows the same path.
For high-bandwidth sequential writes, temporary storage in SLC may be partly skipped to reduce the extra writes generated internally.
For that reason, it is more accurate to see MRAM, SLC and QLC as mechanisms that each handle different problems, rather than as a three-stage buffer through which data flows in order.
Scaling to 105.6TB, and how to read the capacity figures
With the FCM5, the form factor changed from the previous U.2 to EDSFF E3.L, a standard for SSDs.
IBM's product guide cites not only the adoption of high-density QLC NAND but also improvements in component layout, airflow and power delivery. Capacities include 6.6TB, 13.2TB and 26.4TB, as well as 52.8TB and 105.6TB.
In other words, higher capacity is supported not only by NAND storage density. Designs that cool many components within a limited enclosure and supply stable power are also important.
The fact that MRAM reduces the need to flush large amounts of data at power loss works in the direction of easing such implementation constraints.
However, no figure is available showing how much of the FCM5's capacity increase is attributable to adopting MRAM. The larger capacity was achieved through multiple improvements, including high-density QLC NAND and the form factor change, and cannot be explained as an effect of MRAM alone.
Also, because the form factor changed, FCM5 and the earlier U.2 generation cannot be mixed in the same drive enclosure or distributed array.
To understand FCM5 capacity, three kinds of figures need to be distinguished.
The first is the physical capacity that can actually be stored in NAND. The second is the compression ratio, which shows how much data can be shrunk. The third is the upper limit of logical address capacity that can be handled on the system.
If compression and deduplication work well, more original data than the physical capacity can be stored. But that effect varies greatly depending on the content of the data.
The product guide says FCM hardware compression can achieve up to about 3:1 depending on the workload. The upper limit of FCM5's logical address capacity was expanded to about 6:1 relative to physical capacity, compared with about 3:1 for the previous-generation FCM4.
However, the "6:1" logical limit does not mean all data can be compressed to one-sixth. For data on which compression and deduplication have little effect, physical capacity close to the size of the original data is needed.
Treating a large effective-capacity figure as the amount of NAND actually installed could lead to misjudging the storage capacity required.
In the FCM5, the drive itself also takes on some of the processing that improves capacity efficiency.
For deduplication, each FCM provides small hash values for identifying data and information about how compressible it is, and IBM Storage Virtualize looks for data that may be duplicated. The FCM then verifies whether the candidates really have identical content and maps the duplicate data to a single storage location.
This is a distributed design that avoids concentrating huge management tables and all the processing in a central controller, even as storage capacity grows.
However, this feature also has software conditions.
The product guide distinguishes that at the Storage Virtualize 9.1.3 stage only background scanning operates, and full deduplication processing is enabled in subsequent 9.1.3.x updates.
This is a per-version explanation in the product guide, and when deploying, you need to check which features are available in the software version you are running.
The ransomware detection built into the FCM5 is also a separate feature from MRAM's power-loss protection.
In IBM's explanation, the FCM computes I/O statistics, and the storage array controller evaluates them using a model. Having MRAM does not mean drive failures, ransomware attacks, or even data that has not yet reached the SSD from the host are all protected together.
When evaluating an enterprise SSD, you need to check not only how well compression and deduplication work on your actual data, but also up to which stage of write processing it is protected from power loss.
Using large capacity efficiently and safely preserving the state of in-progress writes are separate problems. Only by checking both do the conditions for confidently using high-capacity QLC NAND in business workloads become clear.
