A research team from NVIDIA, Nanyang Technological University, MIT, and others reported "SoL-Pi," a system for making coding agent execution more efficient, in a preprint dated September 17. The team has also published a Pi extension on GitHub.

Rather than retraining the model itself, this research focuses on streamlining the back-and-forth exchanges through which an AI edits code, checks test results, and decides on its next action. In long-running tasks, repeated round trips between the model and tools, along with repeated resending of lengthy past output logs, accumulate over time. The research team compared SoL-Pi against a standard version of Pi using the same underlying model, reporting a reduction in total token volume of up to 49.0%.

However, cutting token usage by nearly half did not come without any cost to task performance. It's important to separate what baseline the "49% reduction" figure was calculated against, and how much of that result the publicly available extension can actually reproduce.

AD

Tokens Dropped 49%, But So Did the Average Score

The paper's central comparison used 51 public tasks from EdgeBench, which include long-running coding work. The team ran GPT-5.6 Sol at the same high reasoning setting (xhigh), comparing standard Pi against SoL-Pi with all four mechanisms enabled.

Summing the input tokens, cache reads/writes, and output tokens recorded by the authors, the total dropped from 2.1538B for Pi to 1.0990B for SoL-Pi (B = billion). Model usage cost, calculated based on API pricing, also fell from $1,339 to $894.

EdgeBench 51 public tasks, GPT-5.6 Sol Pi SoL-Pi (all 4 features enabled)
Recorded total token volume (in billions) 2.1538 1.0990
Model cost based on API pricing $1,339 $894
Average score 44.833 42.003

With all four features enabled on GPT-5.6 Sol, SoL-Pi cut total token volume by 49.0% compared to Pi, while the average score dropped by 6.3%. The cost reduction was 33.2%.

The 49.0% figure here is calculated as the difference between the two systems relative to Pi's total token volume. The cost figures are estimates based on API pricing as of August 17, 2026, not a comparison of actual billed amounts for real users.

While the average score represents 93.7% of Pi's score, this does not mean equivalent results were achieved across every task.

The research team separated the tasks used to search for improvement candidates from the tasks used for EdgeBench evaluation, deciding on the mechanisms to adopt before conducting the final evaluation. However, not all 51 public tasks remained completely untouched until the final evaluation. Eleven tasks were used once, after the mechanisms had been decided, to confirm whether they should be adopted, leaving the remaining 40 tasks for the final evaluation.

The team also ran a test applying the mechanism—built based on GPT-5.6 Sol's execution history—directly to Opus 5, a model not used during the development phase. In this case as well, total token volume dropped by 44.7% compared to Pi, with the average score reaching 94.3% of Pi's.

While a similar trend was confirmed with another model, the paper only verified results on two models: GPT-5.6 Sol and Opus 5.

What the Four Mechanisms Actually Cut

The research team devised 152 improvement candidates and tested them across 535 executable environments. This breaks down into 495 tasks built from GitHub issues and their corresponding fix PRs, plus 40 synthetic tasks that can be run directly to judge success or failure.

After more than 3,000 trial runs, four mechanisms remained. Note that these figures—152 candidates and 535 environments—are not the basis for the 49% token reduction calculation. The team's approach was to adopt only those candidates that improved efficiency without significantly degrading a pre-set performance threshold.

The first mechanism, "Action Fusion," combines file edits or writes with the test or build step that immediately follows into a single tool call.

When what to do after an edit is already predetermined, there's no need to loop the model back in between. This is not used, however, for tasks where the next command depends on checking the edit's results first.

What gets eliminated isn't the testing or verification itself, but the extra round trip to the model that occurs between operations whose results don't need to be checked.

"ObservationPack" is a mechanism that saves tool output exceeding 10KiB locally. The first two times, the full text is sent to the model, but after that, only a short excerpt and a reference identifier are passed along. The original content can be reloaded page by page whenever needed.

This curbs token consumption that arises from long logs or file contents being repeatedly included in context across every subsequent turn.

The problem of already-completed work lingering in the conversation history is addressed by "Online Context Compact."

At each milestone in the work plan, the system compares the cost required to compress the context against the benefit of shortening subsequent inputs, and uses Pi's compression feature when the benefit outweighs the cost. The cost of rebuilding the cache is also factored into this decision.

In the GPT-5.6 Sol comparison, cache read tokens for SoL-Pi with all four features enabled fell to roughly half of Pi's. Cache write tokens, on the other hand, increased. This means looking at just one part of the cache figures isn't enough to judge overall cost.

The final mechanism, "Evidence-Preserving Reducer," has a low-cost auxiliary model read diagnostic logs of 4KiB or more, output from designated builds or tests, first.

The short record produced by the auxiliary model is then mechanically checked against the original log—verifying the hash, exit status, and whether quoted sections match the source text exactly. If verification fails, the original log is passed to the model as-is instead of the summary.

Rather than simply truncating long logs, this mechanism's distinguishing feature is that it verifies summary content and always allows a return to the original text when needed.

AD

The Public Version Is a Pi Extension, With All Four Features Disabled by Default

The SoL-Pi published on GitHub is an independent extension that operates using Pi's extension API. It can be added without modifying Pi itself. The paper verified results using Pi version 0.85.1, which requires Node.js 22.19 or later.

One important distinction to note here: the "automated research loop," in which the AI itself searched for improvement candidates in the paper, is separate from the SoL-Pi extension that general users can install from GitHub.

Installing SoL-Pi does not automatically switch the Codex or Claude Code that users work with over to the same mechanisms.

Also, all four features are disabled by default.

The official README presents a relatively cautious configuration example, enabling only Action Fusion and ObservationPack—which don't require additional model calls and are less likely to interfere with regular work.

The maximum 49.0% reduction reported in the paper, however, is the result when all four features are enabled. This means enabling only Action Fusion and ObservationPack is not guaranteed to achieve the same reduction rate.

Indeed, in a separate configuration in the paper, adding only ObservationPack to GPT-5.6 Sol raised the average score from 44.833 to 47.208, but the reduction in total token volume was limited to just 6.1%.

In other words, the ideal configuration depends on whether the priority is maximizing token reduction or minimizing the impact on performance.

For enterprise use, how logs are handled also matters.

According to the official security documentation, SoL-Pi operates with the same file, command, and network permissions as the Pi process itself, and does not provide its own sandbox or isolated environment.

ObservationPack and similar features store the original output within the session. And if the Reducer is enabled, the target logs may be sent to a configured auxiliary model.

A mechanism exists to detect strings that appear to be secrets and avoid sending them, but this is not a complete safeguard against information leaks. For environments handling diagnostic logs that cannot be sent to external services, the official documentation itself recommends not enabling the Reducer.

Adoption Decisions Should Weigh Both "Tasks Solved" and "Total Cost"

In evaluations outside of EdgeBench, the trade-off between efficiency and task completion rate becomes even clearer.

In a test using 63 CPU-only tasks from Terminal-Bench 4, Pi and Codex each solved 18 tasks, while SoL-Pi solved 15.

Meanwhile, model usage costs for Pi and SoL-Pi were $286.45 and $211.12 respectively, with SoL-Pi being cheaper. While the cost per successful task also decreased, the actual number of tasks solved declined. This test did not include tasks requiring a GPU.

Therefore, the paper's finding that "token usage could be reduced in long-running coding tasks" cannot be extended to mean that costs can be cut "for any kind of development work while maintaining quality."

Organizations considering adoption should test using the tasks they actually handle, and first establish quality benchmarks—such as the number of tasks solved or the success rate on regression tests. Only after that should they compare total cost or cost per successful task.

In addition, operational rules need to be established regarding where logs can be stored and which models or external services they may be sent to.

Only once these conditions are met—reducing unnecessary round trips between the model and tools, and cutting down on repeated resending of the same logs—will the efficiency gains from SoL-Pi translate into actual cost savings.