Every time a new GPU is announced, the headline number is TFLOPS (trillions of floating-point operations per second) climbing another notch. But when an AI model actually runs, how much of that compute can be put to use depends on how much data the memory stacked around the GPU can supply. On August 23, 2026, at Hot Chips 2026, which opened at Stanford University, a Micron Fellow presented data that shakes that very premise. The failure of memory to keep pace with compute is, ironically, etched into the record figures Micron announced in its quarterly earnings two months earlier.

AD

The "3x vs. less than 2x" figure presented at Stanford

Hot Chips 2026 was held at Stanford's Memorial Auditorium from August 23 to 25. Speaking in the Day 0 tutorial session was Raghu Sreeramaneni, a Fellow at Micron responsible for HBM design architecture. He said that "AI accelerator TFLOPS performance grows roughly 3x every two years, while the bandwidth of 2.5D-connected memory such as HBM (a packaging approach in which stacked memory sits next to the GPU on a silicon interposer) grows by less than 2x over the same two years."

ServeTheHome, which covered the event on site, reported the remark. The gap between the growth rates of compute and memory bandwidth means the divide widens with each accelerator generation. However fast a single GPU can compute, if HBM cannot deliver the data that computation requires, the compute units end up idle.

Hot Chips is known as a venue where both academia and the semiconductor industry present new processors and memory architectures, and major chip vendors such as NVIDIA, AMD, and Intel speak there every year. That Micron, one of the leading memory vendors across DRAM, NAND, and HBM, put a specific multiple on the gap between compute and memory-bandwidth growth is unusual: a supplier voicing its own alarm in its own numbers.

Why HBM can't keep up with compute

Sreeramaneni also explained the physical reasons for the constraint. He said that memory silicon, including HBM, accounts for roughly 90% of the total silicon area in a package, and that a 12-high stack can reach eight times the area of the GPU die itself. Another example cited was that, in Meta's Llama 3 operations, 17% of unplanned interruptions were attributed to HBM-related factors. These figures come from an article reporting a single presentation and have not been directly corroborated by Micron in public materials, so they should be taken as "as explained."

The presentation also quantified why HBM is indispensable to AI accelerators. Total system bandwidth with HBM3 reaches 5.3TB/s, while DDR5 memory used in typical servers manages only about 300GB/s. With a gap of more than 10x, GPU makers deliberately adopt the costly HBM.

Lining up the numbers independently confirms that a gap exists between compute and memory bandwidth, but the public record doesn't show which GPU and HBM generations Micron compared to arrive at "3x versus less than 2x." Total HBM bandwidth per NVIDIA GPU is projected to grow from 3.35TB/s on the H100 (2022) to 8TB/s on the Blackwell-generation B200 (2024) and 22TB/s on Vera Rubin (due in late 2026), which works out to roughly 2.4–2.8x every two years. Meanwhile, NVIDIA's FP4 compute, on a sparsity-enabled basis, grows about 2.8x from the B200's 18 PFLOPS to Vera Rubin's 50 PFLOPS (50 quadrillion operations per second). For reference, the H100's FP16 performance is given as 1,979 TFLOPS with structured sparsity and 989 TFLOPS for dense matrices.

That per-GPU bandwidth appears to be growing at nearly the same pace as compute is because the number of HBM stacks itself is being increased to make up the difference. Per stack, bandwidth rises only 2.3x in one generation, from HBM3E's 1.2TB/s-plus to HBM4's (Micron's mass-production part) 2.8TB/s-plus, which is why GPU makers have had to adopt a strategy of adding stacks. But adding stacks expands the silicon area HBM occupies in proportion, which ties into Sreeramaneni's explanation that memory silicon accounts for 90% of the package. With limits on area, power, and heat dissipation, there is a ceiling to compensating through stack count, and that is the practical substance of what Micron calls the "memory wall."

To keep this physical wall from widening, Micron explained, HBM has been extended from 4-high to 16-high stacks, with a path to 20-high. HBM4 also doubles interface width from HBM3E's 1024 bits to 2048 bits, and Micron's mass-produced HBM4 exceeds 11Gbps pin speed and 2.8TB/s of bandwidth per stack (36GB, 12-high). The company says this is a 2.3x bandwidth increase over HBM3E, and a 16-high, 48GB product is at the sampling stage.

Note that 16Gbps-class pin speeds are expected for the next-generation HBM4E and belong to a different generation from current HBM4 production parts. The more layers are stacked, the tighter the constraints on wiring and heat, so the cost of increasing bandwidth tends to outweigh the cost of increasing compute. That is the substance of the "memory wall," the divergence between CPU and DRAM speeds that Wulf & McKee raised in their 1995 paper "Hitting the Memory Wall: Implications of the Obvious," now resurfacing in a new form in the relationship between AI accelerators and HBM.

The CPU–DRAM speed divergence was a problem that first surfaced in the 1990s. It was partially eased by deeper cache hierarchies, faster DDR standards, and the arrival of wide-bandwidth memory technologies such as Wide I/O, a forerunner of HBM, but it was never fully resolved. The memory wall of the AI era can be seen as the same physical constraints of wiring density, power, and heat returning in a new form with the GPU–HBM pairing.

AD

Record 3.5x revenue growth underscores the warning

In its fiscal third-quarter 2026 results (quarter ended May 28), Micron reported revenue of $41.46 billion (about ¥6.5863 trillion at $1 = ¥158.86). That is up 345.8% year over year and 74% from the previous quarter. DRAM revenue was $31.3 billion (about ¥4.9723 trillion), 76% of the total, and Micron said it was a record level for a quarter. For the fourth quarter (ending in August), it gave guidance of revenue of around $50 billion (±$1 billion, about ¥7.943 trillion), non-GAAP EPS of around $31.00, and a gross margin of about 86%.

Sumit Sadana, Micron's EVP and Chief Business Officer, said at COMPUTEX 2026 on June 1 (a separate event from Hot Chips) that "AI context lengths are growing at a pace of 30x per year, while memory content per server has only doubled over the past three years." At the KeyBanc Capital Markets Technology Leadership Forum on August 10, he also said that the company can supply only about half of data center demand and that "2027 will be even tighter than 2026." When demand continues to exceed supply capacity, the supply side holds pricing power.

Record revenue growth reflects market expansion, but it is also a sign that Micron is in a position to sell scarce goods at high prices. Gross margin sits at the guided level of about 86%, a figure that shows the supply-demand squeeze in the profit margin as well. The earnings announcement and the Hot Chips warning may look contradictory, but they are two sides of the same phenomenon of tight supply and demand.

Still, Micron's announcements touch on only part of this picture. As far as has been reported, it is not shown how much of the gap between compute and memory-bandwidth growth is being passed through into prices, nor which GPU and HBM generations were compared to derive "3x versus less than 2x." When an announcement stressing tight supply coincides with record revenue and profit growth, outsiders have few means of verifying how severe the squeeze actually is. Without knowing whether revenue growth stems from higher shipment volumes or higher prices, buyers of AI infrastructure cannot accurately gauge how much their cost structure is deteriorating.

HBM4 and HBM4E as prescriptions, and the different paths each company has chosen

As ways for the industry to address this bandwidth shortfall, Micron's presentation cited several packaging technologies. The leading examples are large substrates such as CoWoS-L and CoWoS-R, extended versions of CoWoS (Chip on Wafer on Substrate, an advanced packaging technology developed by TSMC), which allow denser wiring between GPU dies and HBM stacks. Glass substrates, liquid cooling, and hybrid bonding that joins dies directly were also mentioned. All raise wiring density between chip and memory or carry heat away, forming the foundation for converting compute growth directly into memory-bandwidth growth. Simply adding HBM layers without packaging improvements would quickly run into wiring and thermal limits, so these technologies are positioned as prerequisites for increasing bandwidth.

In development looking beyond HBM4, three companies are taking different paths. SK hynix is reported to be considering TSMC's 3nm process for the logic base die for NVIDIA, while Samsung is said to plan to use its own 4nm process. Micron has already settled on outsourcing logic base die manufacturing for both standard and custom products to TSMC, and expects to begin volume production in 2027. This indicates that "custom HBM," in which the logic base die's circuitry is tailored to customer requirements, is becoming the main battlefield.

In May 2026, Samsung announced it had become the first in the industry to begin shipping HBM4E samples, giving nominal specifications of a 14Gbps pin speed in normal operation, extendable to a maximum of 16Gbps, and bandwidth of up to 3.6TB/s. HBM market share figures differ widely by source, so exact ratios can't be stated definitively, but multiple reports agree on the general direction that SK hynix remains the largest player. The strongest driver of HBM4 demand is NVIDIA. Vera Rubin, due in late 2026, is billed as carrying 288GB of HBM4 and reaching total per-GPU bandwidth of 22TB/s (2.8x the previous Blackwell generation and 6.6x the H100). Whether SK hynix, Samsung, and Micron can reliably supply that capacity and bandwidth will be the dividing line in the next round of order competition.

AD

Who gains and who loses from the bandwidth shortage

When supply fails so completely to keep up with demand, the ones holding pricing power are the three major HBM vendors: SK hynix, Samsung, and Micron. The longer the squeeze lasts, the stronger the bargaining power of those who hold the memory. Kioxia, which makes NAND memory, is also thought to be benefiting from a similar tailwind in the same phase of AI demand expansion.

The benefits of short supply extend to the equipment makers supporting HBM manufacturing. Tokyo Electron is a major supplier of temporary bonders and debonders used in forming HBM's TSVs (through-silicon vias, wiring technology that passes vertically through a chip). TSVs are the vertical wiring that electrically connects stacked dies, and the more layers HBM has, the more important this process becomes. DISCO is said to hold a near-monopoly position in HBM manufacturing steps such as dicing and grinding, which thins wafers precisely before stacking. The more AI chips are produced, the more orders for these Japanese companies rise in tandem.

On the other side, those who lose out are the buyers of HBM. Cloud providers and NVIDIA itself will see margins squeezed by higher procurement costs if memory accounts for a growing share of the cost of an AI accelerator. AI inference businesses that leave expensive GPU compute resources underused because of the bandwidth bottleneck are on the same side. Even after installing the latest GPUs with high TFLOPS, if the memory feeding them with data can't keep up, those GPUs keep running without delivering their full performance.

How deep will the memory wall go over the next two years?

HBM4 is set to become the mainstay of volume production during 2026, and Samsung's HBM4E samples have already reached the nominal maximum specification of 3.6TB/s. Micron's volume production of HBM4E logic dies is scheduled to begin in 2027. As long as HBM bandwidth fails to keep up with growth in compute, demand will continue to exceed supply, and pricing power will remain with the supply side.

If Sadana's outlook that "2027 will be even tighter than 2026" is correct, the next focus is how far packaging improvements can lift bandwidth growth once NVIDIA's Vera Rubin generation begins adopting HBM4 in volume. Whether measures such as glass substrates and hybrid bonding can actually push up the bandwidth of HBM4E and next-generation standards should become clear as the shift to volume production proceeds during 2026. Micron's next earnings report, and the timing at which TSMC-made HBM4E logic dies actually move into volume production in 2027, will be the first indicators of how long pricing power lasts.

The key to narrowing the gap between compute and memory bandwidth lies less in GPU makers' designs than in the yields and packaging technology of those who manufacture HBM. How companies tackle the next wall of 20-high stacks will determine whether Sadana's warning that 2027 will be "even tighter" comes true or whether conditions begin to ease.