At the AI Infra Summit on September 15, 2026, NVIDIA backed its claim of boosting compute without increasing power with actual measured data. Cloud provider Lambda ran a 5-rack, 19-node cluster while keeping it within the power envelope of just 16 nodes, achieving a 24% increase in token throughput. The same day, NVIDIA presented another figure: a demonstration in which its own AI factory automatically reduced power consumption from 4 megawatts to 3 megawatts in response to a utility signal. Yet the number most likely to make headlines is neither of these—it's a separate figure of "up to 40%." So which layer of the power grid is actually being moved by these three numbers, and how confident can we be in each?
Why 19 Nodes Ran Within the Power Budget of 16
What Lambda tested was a cluster of 5 racks and 19 nodes built on NVIDIA HGX B200 hardware. All 19 nodes ran simultaneously while staying within the same power envelope required to run 16 nodes at full power. According to the announcement from NVIDIA and Lambda, overall cluster token throughput increased by 24%, and performance-per-watt improved by 23%. This is described as the first case study validating DSX MaxLPS on HGX B200 servers.
What's driving this is power reallocation. DSX MaxLPS monitors power consumption at both the GPU and rack level, redistributing surplus capacity across nodes depending on workload type. NVIDIA explains that because training and inference draw power differently, facilities running both workloads create room to tighten allocation. Under a static allocation scheme, that slack simply sits unused within the budget.
The two published ratios can be cross-checked against each other. Since performance-per-watt equals throughput divided by power consumption, the change in power consumption can be derived by dividing the throughput ratio by the performance-per-watt ratio. Backing out the change in total power consumption from NVIDIA's two published figures gives 1.24 ÷ 1.23, or approximately 1.008—meaning power consumption rose by only about 0.8%. The claim of "increasing compute while keeping the power budget the same" is internally consistent with the published figures.
However, this only confirms that the two published numbers don't contradict each other—it is not an independent measurement of power consumption. Both the 24% and 23% figures are self-reported by NVIDIA and Lambda and have not undergone third-party peer review such as MLPerf. The scope is also limited to a single test on one cluster of 5 racks and 19 nodes.
Dave Ward, President of Cloud Services at Lambda, said of the validation that his team believes "we've broken through the constraint of a fixed power budget," explaining that DSX MaxLPS opens a path to converting stranded, unused capacity into actual usable capacity. What's being recovered is power that a facility has secured but isn't fully utilizing.
From Boosting Compute to Cutting Power

On a humid evening in August 2026, during a time window when cooling demand spikes, Silicon Valley Power—the municipal utility of Santa Clara, California—sent a demand response signal to an AI factory. The signal was received by software called Conductor from Emerald AI, which, following a pre-established workload priority order, had the lowest-priority jobs cede their power allocation. High-priority inference workloads kept running uninterrupted, while the facility's total power consumption dropped from 4 megawatts to 3 megawatts. No human intervened.
Silicon Valley Power has since continued sending demand signals to the same facility, and all of the more than 200 signals reportedly succeeded. NVIDIA states that the time between receiving the signal and Conductor's response was under one minute. The company positions this Flexible Load Interconnect Program as the first commercial utility program designed to treat AI factories as a grid-dispatchable resource capable of adjusting output on command.
The testbed for this demonstration was not a customer data center—it was NVIDIA's own Eos AI factory. The results were obtained in an environment where NVIDIA's own operations team manages its own hardware, so whether the same responsiveness would hold at another company's facility remains to be verified separately.
Regarding the benefit to the utility, Dion Harris, Senior Director of HPC and AI Hyperscale Infrastructure Solutions at NVIDIA, offered an explanation: if a grid operator or transmission line owner knows that output can be curtailed within a given time window, they no longer need to be as conservative in capacity allocation, since they don't need to stack as much margin. In other words, what this mechanism is going after isn't power stranded inside the equipment—it's the margin the utility had set aside just in case.
This trend didn't start here. In August 2025, Google signed demand response agreements with Indiana Michigan Power and the Tennessee Valley Authority, agreeing to defer machine learning workloads when the grid is under strain. The Santa Clara pilot was announced on April 21, 2026; NVIDIA unveiled the DSX platform at GTC Taipei on May 31; and Emerald AI announced its deployment claim on June 1. The movement to treat AI data centers as a grid balancing resource began with Google's demand response agreement in August 2025, passed through NVIDIA's DSX announcement in May 2026, and took about one year and one month to reach the point where it was disclosed in September 2026 as a measured result of 1 megawatt at a single facility.
Silicon Valley Power is a municipal utility owned by the City of Santa Clara, and this demonstration exists between one utility and one customer. Bringing the same model to another region—including Japan—would require that region's utility to establish a framework for issuing curtailment commands and offering something in return. What NVIDIA has shown is that its hardware and software can respond to such commands—not that the institutional groundwork is in place.
What the "Up to 40%" Figure Actually Refers To
The number most likely to make headlines isn't the 24% or the 1 megawatt—it's "up to 40%." For its next-generation Vera Rubin NVL72-generation AI factories, NVIDIA projects up to 40% more GPU capacity and up to 35% higher token throughput within the same megawatt-scale power envelope. This figure comes with the qualifiers "up to" and "in appropriate deployment environments," and it refers to a different hardware generation than the HGX B200 generation Lambda actually measured.
| Figure | Category | Generation/Scope | Qualifiers |
|---|---|---|---|
| 24% increase in token throughput | Measured | HGX B200 / Lambda's 5-rack, 19-node cluster (1 cluster) | None (self-reported) |
| 23% improvement in performance-per-watt | Measured | Same as above | None (self-reported) |
| Power consumption dropped from 4MW to 3MW | Measured | NVIDIA Eos AI factory (1 facility) | Upon receiving demand signal |
| Up to 40% increase in GPU capacity | Projected | Vera Rubin NVL72 generation | "Up to," "in appropriate deployment environments" |
| Up to 35% increase in token throughput | Projected | Same as above | "Up to" |
| Dedicated DSX Flex deployment at 96MW facility | Planned | Vera Rubin AI factory, Manassas, Virginia | Operation not yet confirmed |
| ~20MW unused out of a 100MW facility | Sourceless generalization | Description attributed to a The Register reporter | The author himself notes the ratio varies by facility |
Of the power-related figures NVIDIA presented, only two are actual measurements: the 24% increase on a 5-rack, 19-node cluster, and the 1-megawatt reduction (from 4MW to 3MW) at a single facility. The headline-grabbing "up to 40%" and "up to 35%" figures are conditional projections for the next-generation Vera Rubin NVL72 platform. If this distinction is lost, the effect verified on a single cluster ends up looking just as certain as projections for a generation of hardware that hasn't even been installed yet.
The bottom row of the table is of a different nature from the sources above it. The claim that a 100-megawatt data center uses at most about 80% of its capacity for critical compute loads, leaving roughly 20 megawatts unused, is a general assumption written by The Register as a broad statement, with no cited source. The reporter himself cautions that the actual ratio varies by facility.
There are also figures that remain at the planning stage. The first commercial deployment of dedicated DSX Flex is planned for a 96-megawatt Vera Rubin AI factory at the NVIDIA AI Factory Research Center in Manassas, Virginia, which NVIDIA says builds on five prior demonstrations spanning two continents. Its operation has not yet been confirmed.
Whose DSX Flex Was the Santa Clara Demonstration?
Regarding the same Santa Clara demonstration, NVIDIA itself writes, "This is not a DSX Flex deployment," while Emerald AI announced it as "the first commercial, multi-megawatt-scale deployment of DSX Flex," and The Register wrote that "NVIDIA has demonstrated that the DSX Flex platform can be used." The three parties are describing different things.
According to NVIDIA's official blog explanation, the Santa Clara case is not a deployment of DSX Flex itself, but rather proof that the underlying concept works at commercial scale; Emerald AI Conductor will be integrated into DSX Flex as the platform matures. In other words, the product name is catching up after the fact, rather than the demonstration following the product. Emerald AI's announcement came on June 1, 2026, while NVIDIA's statement that this was not a deployment came on September 15—the later explanation narrows the scope.
This discrepancy directly affects how many DSX Flex deployments get counted. NVIDIA's description of the Manassas project as the first commercial deployment of dedicated DSX Flex aligns with its position of excluding Santa Clara from the count. By Emerald AI's counting, the first deployment has already happened. There's no publicly available information to determine which framing is correct.
The Register's phrasing may simply have been compressed in the process of summarizing, so it can't be dismissed as an outright error. What is certain is that the same single event has been given three different names, while the substance of that event—that power consumption dropped by 1 megawatt in response to a utility signal—is something all three parties agree on. Tracking a track record by product name simply doesn't work yet in this space.
Two Layers That Moved, and the Grid Itself That Didn't

What DSX is moving isn't grid capacity itself—it's the two layers in front of it. One is the slack left unused within a data center's site due to static power allocation. The other is the margin utilities set aside just in case. What can actually be confirmed through measurement remains bounded: 24% for the former, at the scale of a single cluster, and 1 megawatt for the latter, at the scale of a single facility.
A sense of the scale involved becomes clearer when placed alongside grid-side estimates. A February 2025 report from Duke University's Nicholas Institute estimated that among the top 22 balancing authorities accounting for 95% of U.S. load, if flexible loads could curtail between 0.25% and 1% of peak operating hours, existing grids could absorb between 76 and 126 gigawatts of new demand. The authors themselves characterize this as a first-order approximation, noting that transmission constraints were not modeled and that whether data centers can actually deliver the assumed flexibility remains unknown. That estimate is stated in gigawatts, whereas what actually moved in Santa Clara was 1 megawatt.
There are also boundaries to how far this effect extends. Things get more complicated when third-party accelerators from companies like d-Matrix or SambaNova enter the mix. Regarding d-Matrix, which sits within the NVLink Fusion ecosystem, Harris sees a path toward expanding support, but this requires software integration, and not every chip opens up the same fine-grained control externally as NVIDIA's own hardware does. The Register writes that operators using AMD's Instinct GPUs would need to look for a separate data center management system to get equivalent functionality to DSX.
Separately, startups like GridCARE are pursuing efforts to analytically uncover idle capacity dormant within the grid itself. What NVIDIA chose, before reaching toward that layer, was to recover the surplus left within the walls of its own equipment and within the contract between those facilities and their utility. Since sales of its own hardware are ultimately capped by grid supply, improving customers' performance-per-watt is, in effect, also a way of protecting NVIDIA's own revenue.
The next set of things to watch can be narrowed down: whether Lambda's 24% figure gets reproduced by independent measurement; whether the same demand response holds up at facilities run by operating teams other than NVIDIA's own; and how many other utilities emerge with frameworks similar to Silicon Valley Power's. Once Vera Rubin NVL72 is actually installed, the projected "up to 40%" figure can finally be checked against real measurements. Once those pieces are in place, operators currently stuck waiting in line for grid interconnection will be able to estimate, using numbers from their own facilities, how much compute they can add without waiting on transmission lines.
