While an AI agent executes code, searches the web, or pulls information from a database, the GPU that handles large language model (LLM) inference can sit idle.
In today's AI systems, much of this tool processing either runs on the CPU or is controlled by it. The usual explanation is that "GPUs are bad at conditional branching," but that is not the only reason.
According to NVIDIA's CUDA Programming Guide, GPUs can execute conditional branches. However, when threads in the same warp take different paths, the hardware temporarily disables the threads that are not on the current path and processes each path in turn, which can reduce parallel efficiency.
The tools AI agents use include code execution via Bash or Python, web access, external APIs, file operations, and large search indexes. Reasons for placing this work on the CPU also include integration with the OS and existing software, I/O, and memory capacity.
A paper by researchers at the Georgia Institute of Technology and Intel, "Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective," first posted in November 2025 and updated in April 2026, uses five agentic AI workloads to measure how much of a bottleneck CPU-side tool processing actually is.
The results show that in some cases tool processing dominates overall latency, while in others LLM inference is still the main cost.
In other words, the single sentence "the CPU becomes the bottleneck in AI agents" is not enough. You need to know which tools were used, which CPU and GPU were paired, and what the figures actually measured.
What actually runs when an AI agent calls a tool?
An AI agent does not necessarily finish after the LLM reasons once and returns an answer.
Partway through, it decides "I'll use this tool," receives the result, and then the LLM decides the next action. A coding agent, for example, repeats a cycle: open a file, rewrite code, run Python or tests, read the results, and decide on the next fix.
The paper divides these execution styles into two broad types.
In one, the LLM itself decides which tool to use next. In the other, host-side code, such as Python, decides which tool or LLM to call next.
The research team selected and measured five workloads that use different kinds of tools:
- RAG with Haystack
- Toolformer
- LangChain with web search
- SWE-Agent
- ChemCrow
SWE-Agent, for example, involves file operations through Bash and execution of Python code, which resembles the processing in current coding agents.
The paper says that while AI model inference itself runs mainly on the GPU, much of the tool processing, including Bash and Python execution, web search, LexRank summarization, and searches of large databases, runs on the CPU or is controlled from it.
That does not mean tools must run on the CPU simply because they are tools.
For RAG retrieval, for example, the team used a 115GB corpus of C4 documents. Both GPUs used in the measurements had 96GB of memory, so the search index could not fit in GPU memory, and the team ran exact nearest neighbor search (ENNS) with CPU-based FAISS.
In this case, the choice of CPU was driven directly by the size of the data, not by the GPU's branching performance.
The paper also contains no experiment that implemented the same tool on both CPU and GPU and compared them.
These results alone therefore do not support the conclusion that moving tool processing to the GPU would not make it faster.
GPUs aren't unable to branch

A common explanation of GPUs is that they excel at running the same computation across huge amounts of data but are weak at conditional branching.
That is an easy way to grasp the general direction, but more precisely, GPUs can execute conditional branches. Efficiency tends to drop when threads that run simultaneously split onto different paths.
The SMs (Streaming Multiprocessors) of NVIDIA GPUs execute threads in groups of 32 called warps.
Because a warp executes a common instruction at a time, it is most efficient when all 32 threads follow the same path. This execution model is called SIMT (Single Instruction, Multiple Threads).
To simplify, it is like 32 workers following the same instruction and working in unison.
If everyone does the same task, this is efficient. But if the instruction splits midway, with those who meet a condition doing A and those who don't doing B, the warp can no longer have the same instruction active for all 32 threads.
The CUDA Programming Guide explains that when threads in a warp take different paths because of a data-dependent conditional branch, the threads not on the path being processed are disabled.
This "warp divergence" leaves some of the GPU's arithmetic units unused for a time and lowers execution efficiency.
But this applies within a single warp.
Separate warps are scheduled independently, so divergence in one warp does not stall every thread on the GPU in the same way.
Furthermore, starting with compute capability 7.0, corresponding to the Volta generation, "Independent Thread Scheduling" was introduced.
The GPU keeps execution state such as the program counter and call stack for each thread, and can regroup threads at a finer granularity even within the same warp.
So the picture from older GPUs, where an entire warp always advances in perfect lockstep, doesn't map accurately onto current GPUs either.
Still, the basic problem remains: if many threads in the same warp do different work, SIMT parallelism cannot be fully exploited.
The CUDA Programming Guide also explains that, unlike a CPU core, an SM issues instructions in order and does not perform the branch prediction or speculative execution common on CPUs.
However, these are explanations of the GPU execution model. They do not state that AI agent tools should be run on the CPU.
Across five workloads, the CPU dominated in some cases and not in others
The research team ran experiments on two hardware configurations.
Sys 1 pairs a 64-core Intel Xeon Granite Rapids with an NVIDIA RTX Pro 6000 Blackwell.
Sys 2 is a GH200 Grace Hopper system with a 72-core NVIDIA Grace CPU and an H200 GPU.
Both GPUs have 96GB of memory. The software was PyTorch 2.8.0 and vLLM 0.14.0, and each workload was run five times.
The measurements do not produce the simple result that the CPU always dominates in AI agents.
| Workload | Main tool | What accounted for the largest share of total latency |
|---|---|---|
| Haystack RAG | Exact nearest neighbor search (ENNS) | Retrieval: 81–83% on Sys 1, up to 89% on Sys 2 |
| Toolformer | WolframAlpha API | LLM inference: about 88% on Sys 1, 77% on Sys 2 |
| Web-augmented LangChain | Web search, LexRank summarization | LexRank summarization: 48–55% on Sys 1, 40–45% on Sys 2 |
| SWE-Agent | Bash, Python execution | Tool execution: 25–38% on Sys 1, up to 65% on Sys 2 |
| ChemCrow | 3D conformer generation with RDKit | For heavy molecules, 85% (Sys 1) and 88% (Sys 2) |
For Haystack RAG retrieval, and for ChemCrow when handling computationally heavy molecules, tool processing alone accounted for more than 80% of total latency.
In Toolformer, which uses the WolframAlpha API, LLM inference accounted for the larger share.
Even among "AI agents that use tools," the real bottleneck varies greatly by tool.
LangChain with web search has another factor at play.
Fetching information from URLs involves network communication, so latency varied widely.
A long wait does not necessarily come from CPU compute performance alone.
Speeding up the GPU can make the CPU's share larger
Comparing the two systems also showed cases where faster GPU-side inference made CPU-side processing relatively more prominent.
In Toolformer, the share of total latency taken by LLM inference fell from about 88% on Sys 1 to 77% on Sys 2, which uses the H200.
In SWE-Agent, conversely, the share taken by Bash and Python tool execution rose from 25–38% on Sys 1 to as much as 65% on Sys 2.
The research team explains that when a high-performance GPU shortens LLM inference, CPU-side tool processing that was previously unremarkable can emerge as a new bottleneck.
However, Sys 1 and Sys 2 differ in CPU as well as GPU.
It is therefore not possible to attribute the entire difference between the two systems to switching to the H200. The paper itself evaluates each as a whole system combining a different CPU and GPU.
Doubling the batch from 64 to 128 sharply increased CPU-side waiting

The researchers also measured performance as they increased the number of requests processed in parallel.
For LLM inference on a GPU, raising the batch size up to a point makes use of its parallel capacity and increases throughput.
As batches grow, however, the KV cache consumes more GPU memory, and gains gradually slow because of capacity and memory bandwidth limits.
CPU-side tool processing showed a different problem.
When the batch size in SWE-Agent was doubled from 64 to 128, average latency increased as follows.
| Process | Increase in average latency |
|---|---|
| LLM inference on H200 | 1.06x |
| LLM inference on RTX Pro 6000 | 1.18x |
| Multi-process Bash execution on Intel Granite Rapids | 1.53x |
| Multi-process Bash execution on NVIDIA Grace | 1.94x |
The same tendency appeared in LangChain's LexRank summarization.
When the batch size was raised from 64 to 128, average latency of the CPU-run summarization rose 2.0x on Sys 1 and 1.9x on Sys 2. Meanwhile, GPU LLM inference time barely changed.
The research team says that in the CPU-heavy LangChain, SWE-Agent, and ChemCrow workloads, CPU cores became oversubscribed around a batch size of 128 and throughput saturated.
RAG retrieval hit its limit even earlier: beyond a batch size of 16 or 32, gains slowed because of load on the last-level cache (LLC) and disk I/O contention.
In Figure 4a of the paper, the periods of high CPU and GPU load are also noticeably offset from each other.
While a CPU-heavy tool is running, the GPU waits, and while the GPU is running LLM inference, CPU load is limited to orchestration and runtime data management.
In other words, speeding up only the CPU or only the GPU leaves time during which the other is waiting.
COMB: overlapping CPU and GPU work
As a countermeasure, the research team proposes CPU-Aware Overlapped Micro-Batching (COMB).
Instead of handing a large number of requests to the CPU at once, it splits them into small micro-batches.
The paper shows a method of capping micro-batch size at roughly 1 to 2 times the number of CPUs, based on CPU-side parallel efficiency.
The key is not simply making batches smaller.
Once the CPU processing of one micro-batch finishes, the GPU processes its result, and while it does, the CPU begins processing the next micro-batch.
By overlapping CPU and GPU time in this way, the approach reduces the time either side spends waiting.
In v3, the paper reports that in open-loop load tests, COMB cut service latency by up to 2.9x at P50 and up to 3.9x at P90, and total latency by up to 1.8x.
However, these results come from the specific workloads and hardware the research team built, and the same improvement will not necessarily be obtained for every AI agent.
This study alone doesn't show that CPUs are worse at parallel processing than GPUs
The paper states that, under these experimental conditions, parallelizing LLM inference on the GPU was more efficient than multi-process execution on the CPU.
The differences of 1.06x, 1.18x, 1.53x, and 1.94x when raising SWE-Agent's batch size from 64 to 128 are one example.
However, this was not an experiment that gave the CPU and GPU exactly the same work to compare the hardware's inherent parallel performance.
The GPU side ran LLM inference with vLLM, while the CPU side ran multiple Bash processes. Both the workloads and the software stacks differ.
The authors also include Intel researchers, and as of October 2, 2026 the paper is still a pre-peer-review arXiv preprint.
The appropriate reading is therefore that, for the agentic workloads measured, CPU-side tool processing reached its parallelization limit sooner than GPU inference did.
It is not a result proving a performance gap between CPUs and GPUs as processors in general.
NVIDIA puts more than 22,500 CPU environments in one rack
Products have appeared with this growing CPU demand in mind.
On March 16, 2026, NVIDIA announced the Vera CPU, which it positions for agentic AI and reinforcement learning.
The liquid-cooled Vera CPU Rack announced at the same time holds up to 256 Vera CPUs per rack and, NVIDIA says, can run more than 22,500 CPU environments concurrently.
According to NVIDIA, each environment runs independently and can execute compilers, scripts, runtimes, and the tools that agents call.
Cursor, which develops a coding agent, is also named as a company adopting Vera.
At the announcement, NVIDIA CEO Jensen Huang said that the CPU is no longer merely an aid to the model, and that its role in running the AI system itself is growing.
That said, "more than 22,500 CPU environments" is NVIDIA's own product specification and performance claim, not the result of independent third-party verification.
Forecasts also point to changing CPU-to-GPU ratios
The growth in CPU demand also appears in market forecasts.
On September 29, 2026, TrendForce explained that today's AI data centers commonly use a configuration of roughly 4 to 8 GPUs per CPU.
Citing Arm's estimates, the firm said that while a conventional AI data center needed about 30 million CPU cores per gigawatt, agentic AI could quadruple that to about 120 million cores.
It also said the CPU-to-GPU ratio could move toward roughly 1:1 to 1:2 in the future.
This does not mean that current AI data centers have already shifted to 1:1.
The 120 million core figure is an Arm estimate, and the 1:1 to 1:2 ratio is also a future projection. Actual configurations will vary with the work a data center handles, such as training, inference, and agent execution.
The arrival of NVIDIA's Vera CPU Rack and these market forecasts indicate that attention is returning to CPU-side processing power, not just GPUs, in AI infrastructure.
Which version does the "up to 90.6%" figure come from?
One thing to watch for when reading about this field is the claim that "tool processing accounts for up to 90.6% of total latency."
That figure appears in v1 of the paper, published on November 1, 2025.
The abstract and body of v1 say that tool processing running on the CPU accounted for up to 90.6% of total latency.
In the current v3, published on April 16, 2026, the introduction instead says "up to 88% in workloads where tool processing dominates."
Looking at v3's individual results, RAG retrieval accounts for 81–83% on Sys 1 and up to 89% on Sys 2.
In other words, even under the same arXiv number, the title, authors, proposed optimization methods, and descriptions of performance figures changed between versions.
v1 called the optimization techniques "CGAM" and "MAWS," while v3 renames them "COMB" and "MAS."
The "up to 90.6%" figure that TrendForce cited in its September 29, 2026 article comes from v1. Reading the latest arXiv version, v3, you instead encounter figures such as 88% and 89%.
v3 does not explain why the figure changed from v1.
It is therefore not possible to infer that "performance improved from 90.6% to 88%" or that "remeasurement showed 90.6% was wrong."
The numbers may not have been simply corrected under identical measurement conditions.
Rather than reading only a headline saying "CPU processing accounts for up to 90.6% in AI agents," you need to check:
- which workload the figure was measured on
- which CPU and GPU were used
- which version of the paper is being cited
The same applies to finding out whether the CPU is a bottleneck in your own AI agent.
First, measure separately the time spent on LLM inference and the time spent on tool processing such as search, code execution, and file operations.
What the paper shows is that in workloads where tool-processing time is not negligible, speeding up only GPU inference can make CPU-side processing a relatively larger bottleneck.
It isn't that "CPUs clog up because it's an AI agent." You need to measure where the time goes, then decide whether to improve the CPU, GPU, I/O, or network.
