
AI Agents Put CPUs Back in the Spotlight: What a New Study Shows About Tool Calls and GPU Idle Time
Why do so many AI agent tool calls run on the CPU? Drawing on NVIDIA's CUDA guide and a research paper: GPUs can execute branches, and divergent branches run one after another. In SWE-Agent, doubling the batch from 64 to 128 raised latency 1.06–1.18x for GPU inference and 1.53–1.94x for CPU Bash execution.








