On August 7, NVIDIA responded to observations surrounding the memory specifications of its next-generation CPU, "Vera." In a statement to Wccftech, the company explained that it continuously optimizes compute, networking, and memory, and reiterated the official specification that Vera's modular SOCAMM configuration supports up to 1.5TB. This is not an outright denial of reports that shipping units would have reduced capacity. While retaining the maximum capacity figure, NVIDIA leaves room for offering configurations with varying capacities depending on supply availability and customer use cases.
The point of contention isn't whether "1.5TB has disappeared," but rather which capacity will become standard in mass production. On June 10, TrendForce reported that the SOCAMM capacity in the Vera Rubin Superchip had been halved. While NVIDIA's response and official pages indicate an upper limit, they haven't clarified the standard configuration or whether 96GB modules will be adopted. Taken together, these sources suggest we need to distinguish between the maximum value listed in spec sheets and the configuration that will actually be widely available in the market.
"Up to 1.5TB" Has Not Been Retracted
According to NVIDIA's official specifications, the Vera Rubin Superchip consists of one Vera CPU and two Rubin GPUs. The Vera side features up to 1.5TB of LPDDR5X, while the Rubin side offers up to 288GB of HBM4 per GPU, totaling 576GB across two GPUs. HBM4 bandwidth reaches up to 22TB/s per GPU, and the NVLink-C2C connecting Vera and Rubin provides 1.8TB/s. However, NVIDIA explicitly states that all these figures are preliminary, represent "maximum" values, and are subject to change.
While LPDDR5X and HBM4 are both categorized as "memory," they serve different roles. HBM4 sits close to the GPU, rapidly feeding data for computation, while Vera's LPDDR5X serves as large-capacity system memory handled by the CPU. NVIDIA designed these two to connect coherently, allowing applications to treat them as a single address space. This enables offloading KV cache to the CPU side during inference or reducing data movement when running multiple models.
Vera's 1.5TB represents a significant leap compared to the previous generation, Grace. NVIDIA's technical documentation lists Grace at a maximum of 480GB and 512GB/s, versus Vera's maximum of 1.5TB and 1.2TB/s—triple the capacity and 2.4 times the bandwidth. However, this comparison is between maximum configurations on both sides; the bandwidth for the 768GB configuration mentioned in reports has not been disclosed.
192GB Mass Production and a 60% Supply Gap
On the supply side, large-capacity SOCAMM2 products are moving toward commercialization. On April 20, SK hynix announced the start of mass production of 192GB SOCAMM2 modules for Vera Rubin. Using 1c-nm LPDDR5X, equivalent to the sixth generation 10nm-class process, the company states this delivers over double the bandwidth and more than 75% better power efficiency compared to conventional RDIMMs. The 192GB capacity indicates that components capable of implementing Vera's 1.5TB configuration have entered mass production.
Micron also began sampling 256GB SOCAMM2 modules to customers on March 3. Using 32Gb monolithic LPDDR5X dies, an 8-channel CPU configuration can achieve 2TB. Micron states this reduces power consumption and footprint to one-third compared to standard RDIMMs, and its testing showed a 2.3x speedup in time-to-first-token generation for long-context LLMs. However, sample shipments don't guarantee that NVIDIA will secure the volumes it requires.
Even with mass-producible components available, allocated volumes may not meet demand. According to TrendForce, the LPDRAM allocation NVIDIA will receive from Samsung, SK hynix, and Micron in early 2027 production plans amounts to only about 60% of the company's estimated demand. TrendForce analyzed that NVIDIA is likely reducing the per-unit memory capacity while increasing Vera CPU shipment volumes. This isn't a decrease in memory demand, but rather a decision to distribute a limited bit supply across more systems.
Why Capacity Reduction and SKU Diversity Can Coexist
Wccftech reported observations suggesting a shift from a 1.5TB configuration using 192GB modules to a 768GB configuration using 96GB modules. However, what NVIDIA confirmed to the outlet was the upper limit of "up to 1.5TB"—not the adoption of a 96GB configuration or plans to make 768GB the standard SKU. Therefore, 768GB remains a candidate figure from reporting and cannot be treated as a confirmed specification.
TrendForce's analysis suggests NVIDIA's goal is to reduce per-unit capacity while increasing Vera CPU production volumes. Halving the per-unit capacity allows the same total amount of LPDDR5X to be distributed across more systems. If capacity-tiered SKUs are indeed offered, customers could choose configurations matching their required CPU memory capacity. However, pricing, power consumption, and total cost of ownership across these tiers remain uncomparable at this point.
According to official documentation, Vera's LPDDR5X and Rubin's HBM4 form a single address space, which can also be used for KV cache eviction. Consequently, how much CPU memory is actually utilized will vary by workload. NVIDIA has not confirmed a 768GB configuration, nor has it disclosed bandwidth figures or benchmarks by capacity tier. The performance impact cannot be quantified until per-SKU specifications and measurement conditions are released.
Numbers to Verify With Rubin Ultra
NVIDIA's response this time addressed the SOCAMM ceiling for Vera specifically, and did not directly confirm the HBM configuration for Rubin Ultra. Reducing the LPDDR5X connected to the CPU is a separate matter from reducing the HBM4E on the GPU package. It would be a mistake to conclude, based on NVIDIA's response regarding Vera, that observations about Rubin Ultra's capacity have also been refuted.
Published Rubin specifications indicate up to 288GB of HBM4 per GPU with 22TB/s bandwidth. The first thing to verify is how closely these maximum values align with the primary SKU once Rubin enters mass production. Rubin Ultra's turn will come next. NVIDIA will need to formally disclose how it determines HBM4E capacity and stack count per GPU, along with the bandwidth figures and mass production timeline.
NVIDIA treats all Vera Rubin figures as preliminary. Given that mass production of components supporting 1.5TB configurations is underway alongside supply projections meeting only about 60% of demand, the maximum values listed on product pages alone cannot predict what configurations will actually reach the market. Once major server manufacturers disclose capacity-tiered SKUs and pricing, we'll be able to determine whether this flexibility expands customer choice or represents a compromise absorbing supply shortages.
