On OpenRouter, AI agent traffic now consumes about five times as many tokens as human-driven traffic. Data presented by US venture capital firm Andreessen Horowitz (a16z) on August 21 shows that agentic usage has grown roughly 14-fold since February, with more than 85% of it made up of cached input.

Daniel Newman, CEO of Futurum Group, also suggested in an X post on September 30 that this ratio could widen to 10x or more.

But five times the tokens does not mean five times the inference cost or five times the bill. Cached input, which reuses the same context, is priced differently from new input and output. And as more agents repeatedly handle large amounts of context, the industry needs not only compute but also mechanisms to hold and reuse that information.

To understand how demand from AI agents is growing, it helps to look separately at the volume of tokens processed, the actual billed amount, and the memory and storage that support reused context.

AD

What OpenRouter's "about 5x" actually means

A16z's chart classifies token usage on OpenRouter into "agentic," "human," and "mixed," and shows it as a 7-day moving average.

As of August, agentic usage was 7.3 trillion tokens and human usage was about 1.4 trillion, a gap of roughly five times.

The key point is that this figure does not compare one AI agent with one user doing the same job.

It compares the tokens processed by API usage classified as agentic with those from usage classified as human, across all of OpenRouter. It does not directly show the number of users, the number of agents, or how efficiently a single task is completed.

Also, the 7.3 trillion tokens is not a total for the past seven days. It is a usage level shown as a 7-day moving average.

According to Peter Walker of OpenRouter, who explained the data, agentic token volume overtook human volume on February 6, 2026. In the roughly six months since, agentic usage has grown 14-fold. Human usage also grew 2.8-fold over the same period.

In other words, human use of AI has not declined. What the data shows is that a style of use, in which a human sets a goal and software then keeps calling models repeatedly to carry out the work, has expanded even faster.

The classification method also deserves attention.

According to Walker's explanation and OpenRouter's official write-up, seven indicators, including tool-call frequency, number of conversation turns, and response intervals, are combined to classify each API key as agentic, mixed, or human.

Individual requests are not judged one by one as "human" or "agent." If one API key handles several purposes, it is classified according to the dominant behavior pattern of that key as a whole.

So "about 5x" does not tell us that five agents are running per user, or that a single job requires five times the compute. What it shows is that the composition of tokens flowing through OpenRouter is rapidly tilting toward agentic use.

"85%+" and "about 70%" use different aggregation conditions

A16z says more than 85% of agentic token usage is cached prompts. Walker, meanwhile, says that for the average agentic request, about 70% is cached.

From public information alone, the two cannot be fully mapped onto each other.

Public explanation Cached share What the figure refers to
a16z article, August 21 85%+ Agentic token usage observed on OpenRouter
Peter Walker's explanation About 70% Total tokens of the average agentic request

Sources: a16z article; Walker's post

A share calculated over all tokens and a per-request average can weight large requests with long contexts differently. However, it has not been confirmed that this explains the gap between 85% and 70%.

Rather than treating one as wrong, it is safer to read them as figures with different aggregation conditions.

What the two share is that agents refer back to past context at a very high rate.

For example, a coding agent that fixes code keeps the initial task instructions, repository information, tool definitions, and the changes made so far, while reading new test results and deciding what to do next.

As turns accumulate, earlier context, not just newly added information, takes up a large part of the input.

However, "cache" here does not mean a mechanism that returns a previous answer as is.

OpenRouter also has a separate "Response Caching" that returns an earlier response for an identical request, but the statistics here center on Prompt Caching.

With Prompt Caching, input that is the same as before, such as the system prompt, tool definitions, and past conversation, is reused, and the model generates its next output including the newly added information.

So a high number of cached tokens alone does not show that an agent is pointlessly repeating the same failures. The chart does not include task success rates or the quality of the output.

AD

14x the tokens does not mean 14x the bill

According to OpenRouter's Prompt Caching pricing explainer, published July 21, the unit price of input tokens read from cache is 0.1 to 0.5 times that of regular input, depending on the provider.

If the long system prompts, tool definitions, JSON schemas, and work history that an agent sends every time can be reused, it is cheaper than processing the same amount of input from scratch each time.

For the same model and pricing conditions, the higher the share of cached tokens, the more easily total token volume and the bill diverge in growth.

But the savings apply only to the portion read from cache.

Newly added input and the output generated by the model are charged at regular rates. Some providers also charge more than regular input for the "write" that first creates the cache.

For example, in OpenRouter's explanation, for Anthropic a cache write held for 5 minutes costs 1.25 times regular input, and one held for 1 hour costs 2 times. If a cache is created once and never reused, it can end up costing more than regular input.

So a16z's "more than 85% cached" cannot be read directly as "costs can be cut by 85%."

There are also conditions for the cache to work.

It cannot be reused if the retention period expires or if the beginning of the cacheable prompt changes. And on a service like OpenRouter that routes requests across multiple inference providers, a request sent to a different provider than last time may find no usable cache there.

To reduce this, OpenRouter uses "Sticky Routing," which preferentially sends requests after Prompt Caching has been used back to a provider holding the same cache.

Developers can also specify a session_id to make it more likely that a session is sent to the same provider from the start.

However, if that provider becomes unavailable, requests fall back to another one, so the cache is not guaranteed to be usable for every request.

To manage an AI agent budget, it is necessary to look not only at total tokens but at these categories separately:

  • New input
  • Input read from cache
  • Cache writes
  • Output

Doing so makes it easier to separate why token volume rose from why actual spending rose.

Prompt Caching and KV cache are not the same number

A large volume of cached tokens is also related to memory demand in inference infrastructure. But there is an important distinction here.

The "cached tokens" shown in OpenRouter's statistics count input tokens that were processed using the cache, from an API and pricing perspective.

When an LLM actually runs, by contrast, a KV cache is used, which holds the Keys and Values obtained in attention computation to avoid recomputation.

With Prompt Caching, a provider may reduce processing by reusing such computed state. The specific storage method, however, differs by model and provider implementation.

So the fact that 85% of tokens on OpenRouter were cached does not mean the information corresponding to that 85% is stored as is in GPU HBM.

Nor can the required HBM capacity be calculated directly from the 7.3 trillion tokens, which is a cumulative-style processing volume.

Referring to the same context many times increases the number of tokens processed as cached. But if the same information is being reused, the data actually stored does not necessarily keep growing at the same rate.

The amount of memory needed depends on factors such as:

  • The number of agents and sessions running at the same time
  • The context length of each session
  • The model's number of layers and attention structure
  • The data type of the KV cache
  • Mechanisms for sharing and compressing the cache
  • How long the cache is retained

AD

Long contexts that HBM alone cannot support

When agentic processing runs for long periods and holds large amounts of context simultaneously, it becomes difficult to keep everything in the GPU's high-speed memory alone.

In a technical explainer published March 16, NVIDIA set out an approach of placing context such as KV cache across multiple tiers according to how often it is used and how fast it is needed.

Storage tier Main role described by NVIDIA
GPU HBM (G1) The fastest KV cache, used directly during token generation
System DRAM (G2) Staging and buffering for KV cache evicted from HBM
Local SSD (G3) "Warm" KV cache reused after a short interval
Shared storage (G4) Persistent data such as history and artifacts, away from immediate inference

The closer to the GPU, the faster the access but the more limited the capacity. Moving to more distant DRAM, SSD, or shared storage increases capacity, but the transfer time back to the GPU becomes an issue.

NVIDIA's CMX adds a network-attached flash tier it calls "G3.5" between G3 and G4.

The aim is to hold large amounts of KV cache while pre-transferring information that will be needed again to the GPU or host memory, reducing the time generation stalls waiting for data.

The important point is that simply adding capacity is not enough.

If the context an agent needs cannot be delivered to the GPU by the time it is next used, inference has to wait even if there is plenty of capacity. As agents increase, the ability to store context and move it where it is needed becomes a performance factor in infrastructure, alongside compute.

However, this is a design NVIDIA presents for its own AI infrastructure, and it does not mean OpenRouter stores its cache in the same configuration.

OpenRouter's usage statistics and NVIDIA's infrastructure design cannot be directly linked to estimate the amount of HBM or SSD required.

Still, it shows why the growth of agentic AI cannot be thought of as demand for GPU compute alone. Long-running processes need both the ability to compute and mechanisms to reuse past context efficiently.

Going from 5x to 10x does not mean 10x the value

Daniel Newman suggests that the gap between agentic and human usage, now about five times, could widen to 10 times and beyond.

This differs from the ratio OpenRouter currently observes; it is a forecast about the future. Since no timeline or calculation model was given, it cannot be read as meaning that 10x is certain.

What OpenRouter's data shows is that it is becoming harder to measure enterprise AI demand simply by "how many people use AI."

If a human gives a model one instruction, and behind it an agent searches, writes code, runs tools, reads results, fixes problems, and calls the model again, a large volume of inference occurs with almost no increase in the number of human interactions.

Even with the same 100 people using AI, the compute required differs greatly between those who only use chat and those who run multiple agents for long periods.

On the other hand, the sheer number of tokens consumed cannot tell us whether the investment created value.

An agent that completed a job with 100,000 tokens and an agent that used 1 million tokens and failed midway cannot be judged as the latter having greater demand or value.

The cost to measure is per completed task

As the use of agentic AI grows, the metrics companies should watch change too.

Beyond simple per-token prices or monthly cost per user, they need to look at how much it ultimately cost to complete a single task successfully.

For example, measure together:

  • Total tokens until the task is completed
  • Actual inference cost after caching
  • Time to completion
  • Success rate
  • Time needed for human review and correction
  • Cost spent on failures and retries

Even if caching lets large amounts of context be reused cheaply, it is not efficient if the agent fails repeatedly and humans have to fix the results. Conversely, even with a large token volume, if it reliably finishes work that took a human hours in a short time, it may be worth the cost.

The fact that agentic token volume on OpenRouter has reached about five times human volume shows that the center of AI demand can no longer be captured by simple "human-AI conversations" alone.

The longer agents run autonomously, the more costs are shaped not only by model compute but also by cache pricing design, the memory that holds context, and the mechanisms for storing and moving it.

If that scale expands further, what will matter is not only how many tokens were used, but how many tasks those tokens actually completed.