On September 16, 2026, NVIDIA, Google and Emerald AI launched the AI Energy Management Alliance (AEMA), an industry group aimed at adjusting the power AI data centers draw according to grid conditions. Large AI loads are waiting to connect, while grids must build equipment to cover a limited number of peak hours each year. The idea is to temporarily curb some computing and hand spare capacity back to the grid when it is needed. But there is still a wide gap between showing a "up to 40% reduction" in a pilot and guaranteeing a grid operator a fixed number of megawatts. AEMA is trying to close that gap.

AD

AEMA standardizes response performance, not a percentage

AEMA's founding members are Emerald AI, Google and NVIDIA. At launch, 18 companies and organizations joined as partners, including Anthropic, Analog Devices, AES, Constellation, National Grid, NRG Energy and RWE. The lineup spans not only semiconductors and cloud but also generation, transmission and utilities. Building control capabilities on the data center side alone would not be enough to reach commercial operation unless grid interconnection and tariff rules change too.

AEMA's stated approach is performance-based: rather than mandating particular equipment, it evaluates measurable outcomes. What matters is how many seconds it takes to shed load after a command, how many hours the reduction can be sustained, and how reliably the promised amount can be delivered. The alliance also proposes settling, before interconnection, the conditions for riding through brief grid disturbances, emergency response, and how operational data is shared.

This design turns data center flexibility from a voluntary act, "we'll help if we have room," into a verifiable service. Utilities can offer large loads with confirmed flexibility a faster interconnection path that prices in the risk. Data center operators, for their part, are expected to reflect the benefits of avoided grid upgrades and ramp control in their connection costs.

Still, AEMA is a newly formed industry and policy coalition. It has not published binding standards, a certification scheme, rate schedules or penalty terms for shortfalls. What has been presented so far is a set of issues for measuring performance and a policy direction, not a finished commercial contract.

Translating grid signals into GPU workload priorities

There are three broad ways a data center can reduce its grid draw: move computing with slack in its deadline to another time or region, discharge batteries, or switch to on-site generation. Each can cut the instantaneous power taken from the grid, but none necessarily reduces total computation or daily energy consumption.

On the compute side, a layer is needed to translate commands from grid operators into cluster jobs. NVIDIA's DSX Flex and Emerald AI's Emerald Conductor reorder workload priorities, adjust GPU power caps, and pause and resume jobs when necessary. The grid asks for responses expressed in megawatts and seconds, whereas computing infrastructure deals in job deadlines, checkpoints and customer service levels. Putting both into the same control loop is the core of the technology.

How movable a load is depends not on the facility but on what is running at that moment. Long AI training runs are easy to resume if intermediate state can be saved, and batch inference and video processing can be shifted in time. Interactive inference, search, maps and healthcare-related services, by contrast, have strict requirements for latency and availability. Cloud usage in which customers fix the location and time of processing cannot easily be moved to another region either.

Google says that in 2025 it included machine learning workloads for the first time in demand response agreements with Indiana Michigan Power and the Tennessee Valley Authority. The year before, in the Omaha Public Power District service area, it reduced machine learning power use in response to three grid events. The company itself acknowledges, however, that the regions where this is possible are limited and that workloads requiring high availability, such as search, maps and healthcare, face constraints.

AD

Cutting 25% for three hours, 40% in under a minute

Emerald AI and NVIDIA showed in two demonstrations that flexibility is more than a control-software concept. In Phoenix, a cluster with 256 GPUs cut its power draw by 25% for three hours during an actual grid stress event. According to a related preprint that has not been peer reviewed, the workloads involved stayed within their service levels.

At Nebius's facility in the UK, 96 NVIDIA Blackwell Ultra GPUs responded to more than 200 simulated grid events. NVIDIA's case study says the system reached up to a 40% load reduction in under a minute, about 30% within 30 seconds, and tracked reduction requests lasting up to 10 hours. What matters is that responsiveness to short commands and endurance over long periods were handled in the same set of tests.

The two figures cannot, however, be treated directly as a capability sheet for a commercial facility. Phoenix was a 256-GPU test responding to an actual grid stress event, while the UK test was a simulated event using 96 GPUs running a representative mix of AI workloads. The workload mixes and measurement scopes differ. Moreover, tracking a target in a limited test does not prove that the same reduction can be kept on call at all times in long-term operation, as seasons and customer demand change.

The demonstrations answered the question "can it be controlled?" A contract asks "can it deliver the promised amount when needed?"

Flexibility can't be contracted as a single percentage

A preprint tracking 155,410 GPUs over 185 days found that, against average facility demand of 55.8 MW, the power classified as immediately curtailable averaged only 3.55 MW, or 6.35% of facility demand. The reduction that could be sustained for one hour with 95% probability was 2.51 MW, falling to 1.95 MW for 24 hours.

These figures were reconstructed from operating records of 37,707 Alibaba servers across 17 clusters, in a preprint that has not been peer reviewed published on September 4, 2026. It is not a study that measured facility consumption directly with power meters. GPU utilization and published measurements were fed into a power model, and the 3.55 MW is not a reduction achieved in response to actual commands but an average of power classified as eligible for immediate curtailment based on workload type and priority.

The comparison conditions also need to be read narrowly. The 2.51 MW and 1.95 MW values assume that the power classified as curtailable can be realized 100%, evaluated over the same 185 days, the same power model and 95% availability. If the share that can actually be controlled falls below 100%, the contractable megawatts shrink accordingly. What the study shows is not a particular operator's guaranteed capacity but a structure in which, even for the same facility, the amount that can be reliably delivered decreases as the contract duration lengthens.

Applying the 6.35% average uniformly across all hours would, within the study's model, overestimate the reduction by 17% for one hour, 25% for four hours and 47% for 24 hours. An average does not mean capacity that is available at any time. If job arrival times and priorities are correlated across clusters, aggregating multiple sites does not fully cancel out the variability. What grid operators can buy is not a maximum reduction rate but megawatts specified by duration and confidence level.

AD

Deferred computing returns as rebound demand

When a load-reduction command ends, deferred training and batch jobs resume. Power draw can then exceed normal levels, producing what is known as rebound demand. Even if the three peak hours are weathered, the grid faces a new problem if there is another constraint in the hours when recovery begins.

The 155,410-GPU study calculated that, under illustrative conditions of a four-hour reduction, 100% rebound and a one-hour checkpoint interval, net energy savings including the recovery period could turn negative. This is not an upper bound measured at a commercial facility. It is a warning that, unless contracts specify how quickly deferred work is brought back and how much computation is lost when jobs are stopped, benefits will be overstated by counting only the hours when load was cut.

Peak shaving, deferral of capital investment, energy savings and decarbonization therefore have to be evaluated separately. Temporarily lowering power draw helps with peaks. But if the same computation is processed later, total energy does not fall, and depending on the generation mix at the destination or time of day, emissions could even rise. AEMA's performance criteria need room to record, in addition to megawatts during curtailment, the ramp rate on recovery, the rest period required before the next command, and changes in total energy.

Where FERC reform and AEMA meet: interconnection contracts

About three months before AEMA's launch, the Federal Energy Regulatory Commission (FERC) had already elevated the treatment of large loads to a national regulatory issue. On June 18, FERC ordered the six regional grid operators, PJM, MISO, SPP, CAISO, ISO-NE and NYISO, to justify their existing tariffs or submit revisions within 60 days. Topics under review include transmission service for flexible large loads, mechanisms to prevent costs from being shifted to other customers, co-located generation, and interconnection review.

Under a scheme that speeds interconnection in exchange for flexibility, a data center swaps part of its firm connection capacity for a lower-priority connection that can be curtailed when the grid needs it. If load can be reliably shed during hours when transmission capacity is short, it may be possible to begin operations without waiting for grid upgrades. Conversely, failure to curtail would undermine reliability. Contracts would need a way to determine baseline normal power, metering boundaries, third-party verification and payments for shortfalls.

Priority against customer service levels is also unavoidable. If cloud customers' workloads are delayed to honor a commitment to the utility, another breach of contract could result. Data center operators must separate the workloads they can adjust themselves from those whose timing customers specify, and decide which layer bears the risk of a shortfall. AEMA's standardization needs to reach beyond power-control APIs to operating rules that keep the two contracts from colliding.

What it takes to turn 100 GW into real interconnection capacity

AEMA argues that using flexible large loads could allow US grids to accept around 100 GW of additional load. The figure draws on a Duke University analysis that estimated system-level headroom assuming short-duration load reductions. It is not capacity secured by AEMA members, nor a number that guarantees interconnection in any individual region. Constraints on transmission lines, substations and generation margins remain location by location.

Turning 100 GW into capacity usable in interconnection applications requires at least four kinds of track record. First, separate the movable power by workload type. Second, measure megawatts specified by response time and duration across seasons. Third, have a third party verify delivery reliability including rebound demand, and build responsibility for shortfalls into rates. Finally, the performance has to be matched to specific transmission and distribution constraints.

Whether AEMA becomes an established common specification will not be decided by the number of participating companies or the maximum reduction rate. It depends on whether large commercial facilities can return the promised megawatts at the needed place and time while keeping their promises to customers. When that record is built up in an auditable form, AI data center flexibility will become a power product that speeds up interconnection.