Prices for ChatGPT and Claude subscriptions have kept climbing, yet some API prices are now actually falling. Over the same period, Amazon, Alphabet, Meta and Microsoft announced a combined 2026 capital expenditure of $725 billion, up 77% from $410 billion the previous year. Investment keeps swelling even as some prices are cut. The mismatch became even starker in August 2026, when DeepSeek, long the spearhead of the price war, reversed course and raised prices, reportedly amid tight supply. So who can actually see a race to the bottom through to the end?

AD

Why token prices keep vanishing behind $725 billion in capex

A balance scale with a CPU and a falling coin

Combined 2026 capital expenditure at Amazon, Alphabet, Meta and Microsoft is expected to reach $725 billion. That is a 77% increase from $410 billion the year before, and it follows several of the companies raising their investment plans in their first-quarter 2026 earnings. Meanwhile, OpenAI's API pricing has fallen about 83%, from $30 per million input tokens when GPT-4 launched in March 2023 to $5 for GPT-5.5.

On August 16, 2026, DeepSeek raised the output price of its reasoning model V4-Pro by 355%, from $0.87 to as much as $3.96 per million tokens, but only during peak-demand hours. DeepSeek itself explained that it set off-peak rates 50% below peak rates so that "users can flexibly choose when to run their workloads." Several media outlets, however, attributed the move to tight GPU supply and surging inference demand, a constraint the main players in the price war had not previously discussed. The combination described at the outset, swelling investment alongside falling prices, can be read as the result of a tug-of-war over power and chips.

Well-capitalized camps such as Anthropic, OpenAI and Google had not stopped cutting prices as of August 2026, while DeepSeek turned to a hike in the same month. Only companies that can buy GPUs in bulk and secure electricity can keep fighting the price war; those without that stamina are shifting to price increases. This sorting is happening at the same time as, and behind, the seemingly contradictory moves of expanding investment and cutting prices.

The week NVIDIA lost $589 billion in a single day

On January 20, 2025, DeepSeek released R1, a reasoning-focused model. Its API was priced at $0.55 per million input tokens ($0.14 on a cache hit) and $2.19 per million output tokens, roughly one-27th of the output price of OpenAI's o1 at the time ($15 input, $60 output). Performance was judged comparable to OpenAI's top models, yet the price was orders of magnitude lower, and that combination instantly drew the industry's attention.

It took less than a week for that cheapness to move Wall Street. On January 27, 2025, NVIDIA shares fell 17%, wiping out $589 billion in market capitalization in a single day, the largest one-day loss of market value in U.S. corporate history. The same day, Microsoft CEO Satya Nadella posted on LinkedIn that "Jevons paradox strikes again," suggesting that greater efficiency would expand demand.

Jevons paradox refers to an observation by the 19th-century economist William Stanley Jevons about coal use: the more efficiently a resource is used, the more total consumption can rise, because lower prices expand usage. Applied to AI inference, the lower the price per token, the more use cases open up, and total power consumption and GPU demand actually grow. A survey compiled by the private research firm Menlo Ventures in December 2025 found that generative AI spending among about 500 U.S. companies grew roughly 3.2-fold, from $11.5 billion in 2024 to $37 billion in 2025, while per-token prices fell by about a factor of 10 over the same period. Price cuts did not extinguish demand; they poured fuel on it.

OpenAI's Sam Altman reportedly called DeepSeek's model "very good" and suggested OpenAI would now maintain a smaller lead than before. Altman also disclosed that OpenAI cut the o3 API price by 80% in June 2025, so the DeepSeek shock directly affected OpenAI's own pricing strategy.

The figure often cited as the origin of this cheapness is the "$5.576 million training cost." But that is the estimate in the DeepSeek-V3 technical report, released the previous month in December 2024, and covers only rental costs (at $2 per GPU-hour) for pre-training on 14.8 trillion tokens using 2,048 H800 GPUs, plus context extension and post-training. It does not represent R1's own training cost. It also excludes R&D and investment in existing infrastructure (DeepSeek is reported to have spent more than $500 million cumulatively on GPUs). The surprise that a strong model can be built at low cost is real, but reading a single training-cost figure in isolation calls for caution.

AD

What the $0.78 GPU-hour breakdown reveals about price cuts

gpu-cost-breakdown-illustration.webp

According to Stanford HAI's AI Index Report 2025, the inference cost of a model with GPT-3.5-level performance fell from $20 per million tokens in November 2022 to $0.07 in October 2024, a drop of more than 280-fold in 23 months. Epoch AI's tally shows that for GPT-4-class performance, defined as an MMLU (Massive Multitask Language Understanding benchmark) score of 86 or higher, the price fell 208-fold in two years, from $37.50 in March 2023 to $0.18 in February 2025. What both have in common is that prices keep falling by orders of magnitude while performance is held constant.

An MIT paper (arXiv:2511.23455) published in March 2026 breaks this price decline into three factors. It puts algorithmic efficiency gains at about 3x per year and hardware performance improvement at about 30% per year (roughly 1.43x per year in cost-efficiency terms). It gives no explicit multiplier for the contribution of inter-company price competition itself, but says a residual of roughly 1.5 to 2x per year can be read from the observed decline. Of the three, only algorithmic efficiency stems purely from technological progress; the other two presuppose either massive capital investment or the stamina to sacrifice profit.

A paper from October 2025 (arXiv:2510.26136) calculated the cost of running an A800 80GB GPU for one hour at $0.78, broken down into depreciation of $0.64, electricity of $0.08 and maintenance of $0.06. Electricity accounts for only about 10% of the total, and the main cost driver is depreciation of the GPU itself. Looking at this breakdown alone, the simplification that "electricity determines the cost of AI" seems not to hold.

But depreciation itself depends on the capital spending needed to keep buying new GPUs. According to NVIDIA (February 2026), moving to the Blackwell generation, combined with the low-precision format NVFP4 and software optimization, let Sully.ai cut costs by 90%, Latitude to one-quarter (a 75% reduction) and Decagon to one-sixth (roughly 83% lower). These are results of combined optimization rather than the GPU generation change alone, but the foundation is investment in the latest GPUs, and only companies able to replace their fleets with the next generation in large volumes can benefit. The figure of electricity being 10% of cost refers to the unit cost of running one GPU for one hour; at the scale of data centers running hundreds of thousands of units, the absolute amount of power and the very quantity that can be procured determine whether the investment succeeds.

Apply the depreciation share shown in the A800 cost breakdown (82%) to the $315 billion capex increase (the difference between $725 billion and $410 billion), and the scale equivalent to GPU purchases and replacement comes to roughly $258 billion. This is only a rough estimate, since data center construction also involves land, buildings and cooling equipment that are not captured in the depreciation share, but the picture of most investment going toward GPU replacement does not change. The up-to-90% unit-cost improvement NVIDIA cites is a direct fruit of replacement combined with low-precision computing and software optimization, and both the roughly 1.43x annual hardware progress in the MIT paper and the 1.5 to 2x annual price-competition residual are supported by the investment capacity to keep replacing hardware. But the roughly 10% electricity share is only a breakdown of the operating cost of running one GPU for one hour, and its meaning changes once you build and operate data centers with hundreds of thousands of units. If you cannot procure the electricity to keep the replaced GPUs running, the cost improvement from depreciation ends up a pie in the sky.

According to The Information (February 2026), OpenAI's gross margin fell from 40% in 2024 to 33% in 2025, while Anthropic improved from negative 94% to about 40%, though both fall short of internal targets. Keeping the size of losses within what investors will tolerate while continuing to cut prices makes managing gross margin a tightrope walk. Price cuts made to project cheapness are quietly eroding companies' financial strength.

From Opus to DeepSeek: the map of the price war

ai-price-reversal-arrow.webp

Anthropic's latest model Opus 5/4.8 costs $5 per million input tokens and $25 for output, while its mainstay Sonnet 5 costs $2 for input and $10 for output. OpenAI's GPT-5 is $1.25 for input and $10 for output, and its cheapest tier, GPT-5.6-luna, has been cut to $0.20 for input and $1.20 for output. Google's Gemini 3.7 Flash is $0.75 for input and $3.75 for output. Prices differ by tens of times between the top-end and cheapest models, and the differences in stamina for the price war within the same industry are already etched into the price lists.

With Opus 4.5, released on November 24, 2025, Anthropic offered $5 for input and $25 for output, a 67% cut from the old pricing ($15 input, $75 output), and has maintained that level through 4.8. Sonnet 5 has also carried $2 for input and $10 for output since its release on June 30, 2026, and on August 10, 2026 Anthropic announced it was scrapping its plan to raise prices to $3 for input and $15 for output in September, making current pricing permanent. Less than six weeks passed from release to the withdrawal announcement, and the decision to pull back a price hike so soon after announcing it underscores how strong competitors' price pressure is.

On July 30, 2026, OpenAI cut the price of Luna, the cheapest model in the GPT-5.6 family, by 80% (from $1 input and $6 output to $0.20 input and $1.20 output). The direct reasons the company cited were delivery efficiency and software and GPU optimization, but competition with Anthropic and Chinese open-weight models is also reported to be in the background. In OpenRouter's aggregated data, the token-consumption share of Chinese models jumped from under 2% at the end of 2024 to mid-2026, reported at anywhere from 46% to 61% depending on the method of tallying. The greater the presence of inexpensive Chinese models, the more incumbents are forced to move on price.

xAI's Grok 4.6 costs $2 for input and $6 for output below 200,000 tokens, and Alibaba's Qwen3.8-Max international edition is at nearly the same level, $2 for input and $6 for output. Prices offered in mainland China are set 60 to 70% below the international edition, and a trend of splitting price tiers by region is spreading. Whether American or Chinese, most major companies are moving in step toward lower prices.

In the same month of August, DeepSeek moved in exactly the opposite direction. On August 16, 2026, it introduced separate peak-hour pricing (weekdays 01:00 to 04:00 and 06:00 to 10:00 UTC) for V4-Pro and V4-Flash. For the higher-end V4-Pro, the output price rose from $0.87 to as much as $3.96, and the input cache-miss price from $0.435 to as much as $1.32. Multiple outlets, including InfoWorld, reported that some prices "jumped more than tenfold." The company that had been the standard-bearer of price cuts turned to a hike itself, citing supply constraints.

Anthropic, OpenAI and Google can keep investing tens of billions of dollars of capital through large fundraises, and the decline in gross margin from price cuts is tolerated as a strategic cost they can explain to shareholders and investors. A company like DeepSeek, whose GPU procurement itself is constrained, loses the resources to sustain price cuts the moment demand surges. Although all are described with the same phrase "price war," the depth of the ground they fight on is entirely different. What divides them is, in the end, the difference in ability to raise capital.

AD

MoE and quantization cut costs, and the truth about the TPU myth

MoE (mixture of experts) architecture is a design that activates only a portion of all parameters for each input token. According to Epoch AI's analysis, a sparse configuration with eight experts runs at a compute cost close to that of a dense model with half the total parameters, with the savings varying by processing conditions. Mixtral 8x7B has about 47 billion parameters, but uses only two experts per token, so its inference cost is close to that of a 14-billion-parameter model. Such model-side technical progress is one of the sources that create room for price cuts.

Quantization is a technique that reduces computation and memory by lowering the numerical precision used to represent a model's weights. When precision is lowered to 8-bit integers (INT8), large-scale evaluations have reported accuracy loss generally within 1 to 3%, with memory usage cut by about 50%. Going down to 4-bit integers (INT4), weight-only quantization methods can in some cases preserve accuracy comparable to INT8, though the extent of accuracy loss varies widely with the quantization method and the target task, and memory is reduced by about 75%. In one real-world example, INT4 quantization of a 70-billion-parameter model reportedly cut GPU costs by 83%, from $24,000 to $4,000. Speculative decoding also typically improves throughput by 2 to 3x, and NVIDIA has demonstrated a 3.6x speedup on H200 GPUs.

According to Google Cloud (August 2025), a median Gemini prompt consumes 0.24 Wh (watt-hours) of electricity, and over the previous 12 months energy consumption fell 33-fold and CO2 emissions 44-fold. Efficiency techniques such as MoE and quantization reduce not only costs but also actual measured power consumption. The decline in per-token prices is backed by this improvement in power efficiency.

Beyond such techniques, a common belief is that Google's in-house TPU (Tensor Processing Unit), a chip designed for inference and training, offers lower inference costs than NVIDIA GPUs. An anecdote of one service's monthly cost dropping from $2.1 million to under $700,000 after migrating is often repeated, and the figure that TPUs are 4x cheaper than GPUs has taken on a life of its own.

But measurements by the independent benchmarking organization Artificial Analysis show the opposite. For Llama 3.3 70B inference at a fixed interactive speed of 30 tokens per second, the cost per million tokens was $1.06 on NVIDIA H100, $2.24 on AMD MI300X and $5.13 on Google TPU v6e, giving NVIDIA about a 5x advantage over TPU v6e in cost per token. A blog (junyi.dev) that examined the "TPUs are cheaper" articles pointed out that MLPerf, the cited basis, has no category reporting cost per inference performance at all, and that the articles disagree with each other on figures such as "4x" and "4.7x." Cost differences between hardware swing widely depending on which company measured under what conditions, and in independent measurements NVIDIA actually comes out ahead.

Still, investment in custom inference chips continues. Groq delivers 394 tokens per second on a custom chip at under $1 in cost, and Cerebras achieves 2,100 tokens per second with its WSE-3 chip. Microsoft-backed d-Matrix is reported to claim a 90% efficiency improvement over existing approaches. Which chip wins on absolute cost can flip depending on measurement conditions, but what is certain is the investment direction of moving away from GPU-only reliance, and that too is an option available only to companies that can pour in huge development funds.

Nuclear PPAs and gas turbines: selecting the survivors

ai-power-bottleneck-datacenter.webp

NVIDIA CEO Jensen Huang said at GTC 2026 in March 2026 that "power is the bottleneck," presenting the formula that "revenue is determined by tokens per watt times available gigawatts." Microsoft CEO Satya Nadella reportedly said at Davos in January 2026 that energy costs will determine which countries win the AI race. Two leaders at the top of an industry that has treated GPU performance as the main battlefield were speaking in unison about electricity as a constraint at the same time.

Microsoft and Constellation Energy have signed a 20-year power purchase agreement (PPA) to restart Unit 1 of the shuttered Three Mile Island nuclear plant (835 megawatts), and on June 1, 2026 they obtained approval from the U.S. Federal Energy Regulatory Commission (FERC) for the transmission exemption related to that contract. Nuclear-related power contracts signed by Microsoft, Amazon, Meta and Google total about 7.5 gigawatts (Meta is the largest, including investments in TerraPower and Oklo, and some reports put the upper limit at 6.6 gigawatts). Across the industry, data center contracts using nuclear power reach 9.8 gigawatts in 13 deals by seven companies, equivalent to the electricity of about 7 million households. Because new nuclear plants take a decade or more to build, restarting existing reactors and locking up power through long-term contracts are, as Huang and Nadella suggest, the realistic solutions for now.

Nuclear alone is not enough, and money is also flowing into on-site natural gas generation. Behind-the-meter (directly connected, bypassing the grid) gas generation capacity planned by data center operators in the United States came to about 101 gigawatts as of May 2026, of which more than 57 gigawatts had reportedly already been ordered. The plans themselves keep piling up month after month.

GE Vernova, the largest gas turbine maker, has an order backlog filled through 2029, and in October 2025 the power infrastructure company VoltaGrid and Oracle signed a 2.3-gigawatt on-site gas generation contract for an AI data center in Texas. In some regions, the wait for a new grid connection is said to exceed five years, and operators who cannot wait for grid upgrades are increasingly dependent on self-generation. The move to build their own power plants rather than wait for utilities to reinforce the grid makes up part of the content of the hyperscalers' $725 billion in capex.

Both nuclear PPAs and gas-turbine self-generation require years of upfront investment from groundbreaking to operation and credit lines worth tens of billions of dollars. The four companies that can provide this can secure power while continuing to cut prices, while those that cannot are forced to choose between raising prices and remaining dependent on other companies' infrastructure. This structure is not fixed, however. Camps with less capital, such as DeepSeek, could change strategy again if the environment for procuring power and chips changes. The war of attrition over cheapness is in the process of screening participants by the yardstick of power and chip procurement ability, and the outcome is not yet settled.

Japan, with its ¥22 electricity, and where it fits in

Japan's industrial electricity price is about ¥22 per kilowatt-hour, said to be higher than in the United States, China and South Korea. China's industrial electricity price is reportedly only about one-third of Japan's. Data centers in countries with cheap electricity have a base from which to offer AI services of the same performance at lower cost, so the outcome of this war of attrition directly rebounds on Japan.

According to an estimate presented in January 2025 by the Organization for Cross-regional Coordination of Transmission Operators, Japan (OCCTO), electricity demand for data centers is expected to reach 44 terawatt-hours in fiscal 2034, or 51 terawatt-hours including semiconductors, accounting for about 14% of industrial and other demand. AI's power demand is beginning to carry a weight that cannot be ignored even within Japan's industrial structure.

NTT's proprietary lightweight language model "tsuzumi," announced in November 2023, is said to hold inference costs to as little as 1/70 of GPT-3. For Japan, which faces a disadvantage on electricity prices, model-side efficiency is one of the few countermeasures.

On April 27, 2026, NTT announced "AIOWN," an AI-native infrastructure initiative, laying out a plan to expand the IT power capacity of domestic data centers more than threefold, from about 300 megawatts at present to about 1 gigawatt by fiscal 2033. Expanding IT power capacity means building up, over years, negotiations with utilities over grid reinforcement and new generation contracts, and AIOWN's figures represent the progress of those negotiations themselves. Whereas U.S. PPAs made use of existing assets through nuclear restarts, Japan's expansion of IT power capacity requires reinforcing the grid itself, so its timeline is longer.

Meanwhile, SoftBank Group announced on May 30, 2026 that it would invest up to €75 billion (about ¥14 trillion) in France to build one of Europe's largest AI data centers. The plan calls for 3.1 gigawatts by 2031 and eventual expansion to 5 gigawatts. Japanese companies themselves are choosing to bypass domestic power constraints and direct investment overseas where power conditions are better.

The two figures, $725 billion in capex and the 83% decline in OpenAI's input price, can be read as two sides of the same sorting mechanism. Camps that can procure power and chips in large volumes keep cutting prices, and those that cannot turn to price hikes, as DeepSeek did. For Japan, the next fork in the road depends on how far NTT can bring forward its expected expansion to 1 gigawatt of IT power capacity by fiscal 2033, and whether a path can be drawn for domestic electricity prices to move closer to those of major countries. Whether Japan can keep enjoying cheap AI services will be decided by the outcome of this tug-of-war over electricity, which lies outside the performance race.