On September 1, 2026, Nikhil Cherian of Google Cloud spoke at a memory-focused forum at SEMICON Taiwan. He explained that Google is addressing AI's "memory wall" from both the hardware and software sides. The session overview names HBM and advanced DRAM as major cost drivers for AI servers.

Without enough capacity, a system cannot hold models or conversation history. Even with enough capacity, slow data transfer leaves expensive compute units waiting. Google combines TPU designs tailored to specific workloads with a compression technique that shrinks the KV cache to one-sixth. How much cost this actually saves depends on what is stored in memory and where it has to be moved.

AD

A parts bill that is over 75% memory, and a growing KV cache

Shibao Zixun (時報資訊) reported that Cherian said high-performance memory accounts for more than 75% of the bill of materials (BOM) for AI servers. However, the report does not specify the server configuration, the procurement conditions, or the split between HBM and DRAM. The figure cannot be treated as Google's overall capital spending or as an average cost for AI servers in general.

In generative AI inference, the keys and values computed at each layer are kept as a KV cache so that past tokens do not have to be recalculated from scratch every time. The longer a response runs, the more history must be stored. As Hugging Face's explainer shows, the approximate capacity depends on the model's layer count and KV head configuration. That is multiplied by the number of tokens stored, the number of sequences processed at once, and the number of bytes used to store each value. In services that read long documents or handle many conversations at once, this history swells memory use on top of the model weights themselves.

Behavior changes with shared caches, implementations that discard old history, and models with specialized attention mechanisms. Even so, the relationship holds: longer conversations and more concurrent connections mean more capacity, which is the starting point for service design. Quantization reduces the capacity per element, so it acts most directly on the portion occupied by the KV cache.

Weights, the learned parameters, are accounted for separately. Sparse MoE (mixture-of-experts) models, which switch among multiple expert networks, compute only some experts per token to reduce computation. In a typical serving setup that loads all experts into memory, the weights of unselected experts must be stored as well. As an MoE explainer points out, weights that are not used in computation do not disappear from memory. Compressing the KV cache does not automatically shrink weights, training state, or data held on SSDs.

The 75% is a share of parts cost, and the one-sixth is a ratio of KV data capacity. Without knowing how much of the high-performance memory the KV cache uses, or how much the purchase price falls when capacity is reduced, a server cost reduction rate cannot be derived from these two figures.

TPU 8t and 8i use memory differently

The eighth-generation TPU, for which Google published a technical explainer on April 22, takes different approaches to memory in the 8t and 8i. The TPU 8t/8i technical deep dive positions the 8t mainly for large-scale pretraining and the 8i for low-latency inference and post-training. HBM is high-capacity, high-bandwidth memory placed near the compute chip, while SRAM is a small, fast memory area inside the chip. The capacity, bandwidth, and compute figures in the table are published per-chip specifications for the same generation.

Item TPU 8t TPU 8i
Main use Large-scale pretraining Low-latency inference and post-training
HBM capacity 216GB 288GB
HBM bandwidth 6528GB/s 8601GB/s
On-chip SRAM 128MB 384MB
Peak FP4 performance 12.6PFLOPS 10.1PFLOPS
Interconnect and main aim Large-scale connection via 3D torus Shorter communication and aggregation via Boardfly and CAE

The 8t connects 9,600 chips in a 3D torus. Multiplying 216GB per chip by 9,600 gives 2,073,600GB, or 2.0736PB in decimal units, so the total HBM across the connected system is about 2PB by calculation. However, because memory that is split across chips is reached through the interconnect, this does not mean 2PB can be used like the local memory of a single chip. Communication and distributed-placement constraints do not disappear.

The 8i aims to keep more frequently used data in its 384MB of SRAM, and to shorten inter-chip communication and aggregation with the Boardfly interconnect and CAE. CAE is a dedicated engine for aggregation processing. This does not mean, however, that all model weights or every conversation's KV cache fit in SRAM. The specs show that the 8i falls below the 8t in peak FP4 performance but offers more HBM capacity and bandwidth and more SRAM. The emphasis is on delivering data to the compute units so responses are not kept waiting. Peak compute performance alone does not determine real-world inference speed.

TPUDirect Storage, introduced for the training-oriented 8t, is a mechanism that sends data from storage to the TPU while avoiding congestion in the data path through the host CPU. It is neither a feature that increases HBM capacity nor one that compresses the KV cache.

The 8t and 8i were announced in April, and Google Research introduced TurboQuant on March 24. At the September talk, these were presented together as measures against memory constraints. On Google Cloud's product page, checked on September 4, both the 8t and 8i are listed as coming soon.

AD

Reading TurboQuant's 3 bits and 8x separately

What TurboQuant compresses is the KV cache that accumulates during inference. Google Research's introduction explains that it quantized the KV cache to 3 bits while maintaining model quality in testing, and compressed KV memory by at least 6x on a long-context information retrieval task. It can also be applied without retraining or additional fine-tuning of the model.

The compression is not lossless. In a preprint published in April 2025 by Amir Zandieh and colleagues, the authors state explicitly that quantization is inherently a lossy process. TurboQuant smooths out the skew in per-coordinate values with a random rotation, making them easier to reduce to low bit widths. In the version that uses inner products, a 1-bit QJL corrects the residual to reduce bias in the inner-product estimate. Unlike lossless compression that fully restores the original values, it is a technique for reducing storage while preserving the needed computational accuracy.

However, the 3-bit quantization and the up-to-8x speedup were measured under different conditions. The 8x maximum Google reported compares attention-score computation using 4-bit keys on an H100 against 32-bit unquantized keys. Reducing KV capacity and speeding up attention scores differ in both bit width and what is measured. It does not mean overall inference becomes 8x faster, nor was it measured on a TPU.

Table 1 of the preprint also includes measured values showing the relationship between quality and bit width. In an example evaluating Llama-3.1-8B-Instruct on LongBench-V1, the unquantized 16-bit cache and the 3.5-bit setting both scored 50.06 on average, and the 2.5-bit setting scored 49.44. This is a single experiment from the preprint and is not the same test as the 3-bit evaluation in the blog, which includes Gemma and Mistral. With a different model, context length, or type of question, the usable compression ratio changes as well.

How much total memory does 6x compression save?

The reduction in total memory use is determined by the share the KV cache occupied before compression. Let total memory used be M and the KV cache share be f. Assume the KV cache is compressed to one-sixth, all other capacity stays constant, and any extra space needed for compression is ignored. Memory use after compression is then M×(1−f)+M×f/6, so the reduction rate is f×5/6.

KV share before compression (assumed) Reduction in total memory use
20% 16.7%
50% 41.7%
80% 66.7%

If the KV cache is 50% of the total, compressing it to one-sixth cuts total memory use by 41.7%.

The same compression technique has a smaller effect on short requests where weights make up a large share. Conversely, when holding long documents and handling many conversations in parallel, the KV share rises, and the same compression ratio saves more capacity. Each row of the table is a hypothetical comparison with the model, context length, and number of concurrently processed sequences held fixed before and after compression. They are neither Google's measured values nor figures that directly indicate the quantity of HBM purchased, processing speed, or number of concurrent connections.

Translating this into cost requires further conditions. If more conversations are added to the freed HBM, the total installed capacity across a facility may not change even though each request needs less. Compressing the KV cache also reduces the amount of that data transferred, but if reading weights or inter-chip communication account for much of the waiting time, overall responses will not speed up by the same proportion. Software that saves capacity and hardware improvements that deliver data address different constraints.

AD

When compression and rising demand advance together

On August 25, TrendForce forecast that DRAM and NAND will account for 47% of major cloud providers' total capital expenditure in 2026 and 68% in 2027. It also projected that server DRAM contract prices will rise about 270% in 2026 and enterprise SSD prices about 235%.

However, TrendForce's 47% and 68% are forecasts with cloud providers' capital spending as the denominator, while the 75% reported by Shibao Zixun is a share of AI server parts cost. Because the capex forecast also counts NAND, the two cover different scopes, and the 75% has not been corroborated by another survey.

If compression lowers the memory burden per request and compute and bandwidth have headroom, providers can handle longer contexts and more concurrent connections on the same equipment. If that headroom is filled by new uses, total demand is maintained and may even grow. According to Nextapple (壹蘋新聞網), Cherian said that even with the KV cache cut to one-sixth, there still would not be enough memory for the scale Google is aiming for. Being able to run on less memory is a separate matter from a company buying less memory overall.

To judge the effect of deploying the 8t and 8i, one needs to compare total memory use, including non-KV components, and cost per unit of processing in real operation, with the same model and response quality. If the freed capacity can be used without increasing response latency and the number of requests handled rises, the gains from compression can be turned into the ability to provide more AI services on the same equipment.