In July 2026 data flowing through Vercel's AI Gateway, Anthropic processed 29.8% of all tokens while accounting for 65.1% of spend when converted to public list prices. The average price per token was 4.4 times the average of all other labs combined. Usage volume rose 59% month over month, and list-price-based spend grew 37%, yet the average unit price across the entire Gateway fell 13.6%. Within the same month's aggregate figures, we can see both a shift of workloads toward cheaper models and continued spending on expensive models for other tasks.
Vercel routes tens of trillions of tokens each month between production applications and AI labs. The "AI Gateway Production Index August 2026," published on August 11, is based on anonymized and aggregated data through the end of July. This is not a measure of share across the entire AI API market. Even so, it does reveal how development teams that actually switch between multiple models are responding to price differences.
The question is why Anthropic's per-token price premium widened from 3.4x in June to 4.4x in July, even as the overall unit price was falling. The answer isn't a single quality assessment. After examining the measurement methodology and the supply of low-cost tiers, we need to separately consider coding-agent use cases and the operational pattern of customers who frequently reconfigure their model mix.
What Does the 4.4x Figure Actually Compare?
The 4.4x figure is not the billed amount when the same prompt is sent to each company's models. Vercel calculates "Spend" as the total of each request valued at the lab's public list price. "Price per token" is that total spend divided by total token volume—a weighted average that mixes differently priced input and output tokens, along with caching and long-context usage. The AI Gateway states that it uses providers' list prices as its baseline and does not mark up token prices.
The relationship between the numbers is straightforward. Looking at Anthropic alone, dividing its 65.1% spend share by its 29.8% token share yields a relative index of 2.1846. Other labs processed 70.2% of tokens for 34.9% of spend, giving a relative index of 0.4972. Dividing the former by the latter yields approximately 4.39, matching the 4.4x figure Vercel reported.
However, this ratio does not mean "Claude is always 4.4 times more expensive for the same work." The mix of models used differs across labs, as does the ratio of input to output tokens and the presence or absence of long context. Because this is a list-price metric, it does not reflect the prices actually applied through direct contracts or cloud channels, including enterprise volume discounts or subscriptions. Nor does it represent what customers actually paid or each company's actual revenue.
A high spend share alone also cannot prove a causal link to quality. Customer profiles and use cases differ, and factors like model availability and the reintroduction of Fable 5 are mixed into the same aggregate. The 4.4x figure should be read as an indicator of the combination of workloads chosen on the Gateway and public pricing—nothing more.
Cheap Models Capture Volume, Claude Captures Dollars
In July, DeepSeek accounted for roughly 25% of all token volume, overtaking Google's 10.7% to claim second place. DeepSeek V4 Flash alone processed more tokens than all of Google's models combined, handling about one-fifth of the entire Gateway's volume. Across all open-weight models, the token share was 36% against a spend share of only about 9%. This trend of shifting bulk processing toward the low-price tier is directly reflected in the falling unit price.
The shift at OpenAI is even clearer. According to Vercel, 85% of the token volume added to OpenAI in July flowed to GPT-5 Nano, pulling OpenAI's average per-token price down to 58% of its June level. GPT-5 Nano is priced at $0.05 per million input tokens, $0.005 for cached input, and $0.40 for output. Vercel calculates that even Anthropic's cheapest model, Haiku 4.5, costs about two-thirds of the Gateway-wide average, while GPT-5 Nano runs about one-sixth, and DeepSeek V4 Flash about one-sixteenth.
Anthropic, meanwhile, has no product in the market's rock-bottom price tier. In Vercel's data, Anthropic has maintained a spend share above 60% in every month observed, reaching 65.1% in July. A product lineup that drops prices to capture large volumes of lightweight tasks, and a lineup that commands higher per-token prices for other work, sit side by side on the same usage base.
That's why the 13.6% drop in the overall unit price doesn't necessarily reflect a simultaneous cut in list prices across all models. Vercel explains that if June's model mix were held constant, the unit price would have been roughly flat. The main driver of the decline is that users shifted their routing toward cheaper models. Even as volume in cheap processing rises, spend on the premium tier can persist—these two trends are not mutually exclusive.
The Persistent Choice of Premium Pricing in Coding Agents
The largest use case for tokens on the Vercel Gateway is coding agents. In July, DeepSeek handled roughly one-third of token volume in this use case. Anthropic, by contrast, accounted for over 80% of list-price-equivalent spend in the same category. In other words, while cheap models carried much of the token volume, list-price-equivalent spend concentrated on Claude.
The reintroduction of Fable 5 cannot be overlooked here. Fable 5 was halted on June 12 due to export controls and relaunched globally on July 1 after the restrictions were lifted. By July, it had grown to account for 13.2% of total Gateway spend, ranking second behind Opus 4.8. Reportedly, 90% of teams using Fable 5 had not used the model before. While we don't know their usage history outside the Gateway or the true net increase in customer count, the pool of teams using it in July was substantially different from the pre-suspension user base.
Fable 5's API list price is $10 per million input tokens and $50 for output. Opus 4.8's standard price is $5 for input and $25 for output. Anthropic markets Fable 5 around its capacity for long-horizon autonomous work and token efficiency, and has cited customer case studies. However, Vercel's data alone cannot confirm whether that marketing claim has been proven in terms of quality difference or return on investment.
Development teams weigh per-token price alongside the cost of failure, processing speed, and the amount of rework needed to reach completion. Vercel explains that customers assess task difficulty and the cost of failure, choosing models according to speed and price. This data doesn't reveal how much premium-tier models actually reduce manual rework or failures. Still, the fact that Anthropic's list-price-equivalent spend share exceeds 80% in coding shows that allocation isn't determined by price alone.
Model Mixes Aren't Fixed—They Shift Month to Month
Of the tokens that passed through the Gateway in July, 81% were processed by models that didn't exist there six months earlier. Among teams that used more than 10 million tokens in both June and July, 75% changed their model mix by more than 10%, and 60% changed it by more than 25%. The median team saw its unit price fall by 2.9%, but a quarter of teams cut their price by more than 30%, while another quarter raised it by more than 20%.
Averages alone can't explain individual teams' decisions. Some teams shifted much of their workload to cheaper models, while others moved processing toward models with higher per-token prices. Vercel's claim that models can be "swapped with a single line" refers to the switching cost at the API integration level. In actual operations, this requires adjusting prompts, evaluating outputs, and absorbing differences through auditing—and there's no confirmation that this burden disappears.
Ramp's June 2026 data on U.S. corporate spending is consistent with this broadening pattern of concurrent usage. 42.4% of companies used Anthropic, and 39.5% used OpenAI. Companies using model-provisioning platforms made up only 5.8% of AI-spending companies, but among those, 93.2% also used Anthropic and 85.8% also used OpenAI. Because Vercel and Ramp differ in population and unit of measurement, their figures cannot be combined directly. Still, this is consistent with an explanation in which usage increasingly combines multiple price tiers by use case, rather than standardizing entirely on a single company's models.
Among companies using model-provisioning platforms in Ramp's data, median AI spend was $248 per employee—about 23 times the $10.59 median across all AI-spending companies. Ramp suggests this may indicate that cheap models are increasing overall AI processing volume rather than reducing total spend. This is Ramp's own interpretation and doesn't directly explain Vercel's falling unit price. But the fact that adopting low-price models doesn't automatically mean shrinking AI budgets is a condition worth confirming in practice.
To gauge whether Anthropic can sustain a 65.1% list-price-equivalent spend share, we'll need to track not only how far cheap models expand within coding use cases, but also Anthropic's own product lineup, pricing, and token composition by use case. Which workloads do teams that switched models eventually shift back to the premium tier? Only once companies publish not just per-token prices but the actual cost—and amount of rework—required to complete real tasks will it be possible to compare the economics behind these choices.
