On September 29, OpenAI announced a new AI model, GPT-6.1 Sol. It keeps the standard input and output API prices of the previous GPT-6 Sol while improving performance in areas such as code fixes and business automation, and it halves the price of cached input.

OpenAI positions the model as offering performance close to the higher-tier GPT-6 Astra at one-fifth of Astra's input and output prices. However, looking at the published results by reasoning setting, some tasks score higher at lower settings than at maximum compute. In some areas, the gap to Astra and Claude Opus 5.5 also remains large, so this is not simply a "cheaper Astra."

AD

In code fixes, "high" scores better than "max"

On DeepSWE 1.1, a benchmark of software development ability, GPT-6.1 Sol scored 75.2% with its reasoning setting at high. In the graph published by OpenAI, the cost per task was $0.65. That is 6.4 points above the previous Sol's best score of 68.8%, which cost $2.74 per task, far more.

DeepSWE uses real codebases and asks models to carry out software changes that require multiple steps. Datacurve, which runs it, says it prepared 113 tasks, created without directly reusing existing fix histories, spanning 91 repositories and five programming languages. Rather than filling in short snippets of code, it is close to real development work: reading existing code, modifying it, and verifying the result.

Reasoning setting GPT-6 Sol score Previous Sol cost GPT-6.1 Sol score New Sol cost
low 37.2% $0.16 64.4% $0.17
medium 56.6% $0.38 73.0% $0.42
high 65.3% $0.64 75.2% $0.65
xhigh 66.6% $1.00 71.9% $0.79
max 68.8% $2.74 71.9% $1.57

Costs are per task on DeepSWE 1.1. Figures are based on graphs in OpenAI's announcement and come from evaluations using the company's research environment or the API.

Even at medium, the new Sol beats the previous Sol's max. Work that used to require a high reasoning setting may now be handled at a lower one.

On the other hand, comparing the same settings, costs from low through high are nearly unchanged from the previous Sol. The model update did not make every use case uniformly cheaper.

DeepSWE 1.1: Raising the reasoning setting does not always raise the score横棒グラフ。カテゴリ 5 件、系列: Evaluation score(単位: %)GPT-6 Sol (max)GPT-6 Sol (max)GPT-6 Sol (max) — Evaluation score: 68.8%68.8GPT-6.1 Sol (medium)GPT-6.1 Sol (medi…GPT-6.1 Sol (medium) — Evaluation score: 73%73GPT-6.1 Sol (high)GPT-6.1 Sol (high…GPT-6.1 Sol (high) — Evaluation score: 75.2%75.2GPT-6.1 Sol (max)GPT-6.1 Sol (max)GPT-6.1 Sol (max) — Evaluation score: 71.9%71.9GPT-6 Astra (xhigh)GPT-6 Astra (xhig…GPT-6 Astra (xhigh) — Evaluation score: 74.1%74.1単位: %
データを表で見る
Evaluation score (%)
GPT-6 Sol (max)68.8
GPT-6.1 Sol (medium)73
GPT-6.1 Sol (high)75.2
GPT-6.1 Sol (max)71.9
GPT-6 Astra (xhigh)74.1
DeepSWE 1.1: Raising the reasoning setting does not always raise the scoreFigures published by OpenAI. Reasoning settings differ by model. Statistical significance of the score differences cannot be judged from this graph alone.出典: OpenAI, "Introducing GPT-6.1 Sol", DeepSWE 1.1

For the new Sol, high scored 75.2% while max scored 71.9%. That is 3.3 points lower, yet the cost rises about 2.4 times, from $0.65 to $1.57.

In other words, more reasoning does not always produce better results. Rather than fixing max for every use, it makes more sense to find the setting that meets the quality your actual work requires.

That said, these results alone do not support concluding that making the model think longer causes performance to drop.

Astra's xhigh also scored 74.1% at $4.43, which on paper is below the new Sol's high. But this small gap alone is not enough to say the new Sol is better overall.

These results are better read not as a ranking of the models' abilities but as material for asking which setting achieves the required quality at the lowest cost.

The gap to Astra varies across PDFs, business automation, and screen operation

On GDP.pdf, which asks models to read complex PDFs and answer questions, the new Sol at high scored 32.0% at $0.35. Astra at the same setting scored 31.0% at $1.79, and Opus 5.5 scored 28.8% at $0.83. In OpenAI's comparison, the new Sol comes close to the higher-tier model at a lower cost.

Surge AI's GDP.pdf uses PDFs from real work in 10 fields, including finance, healthcare, and law. Models must not only read text but combine tables, charts, diagrams, and information in small print to answer.

It is a different evaluation from "GDPval-AA," which compares actual deliverables such as documents and spreadsheets.

Evaluation / setting GPT-6 Sol GPT-6.1 Sol GPT-6 Astra Claude Opus 5.5
GDP.pdf / high 28.0% / $0.35 32.0% / $0.35 31.0% / $1.79 28.8% / $0.83
AutomationBench 1.0.6 / medium 26.9% / $0.21 31.7% / $0.19 34.1% / $1.27 29.5% / $0.65
AutomationBench 1.0.6 / max 32.0% / $0.34 36.1% / $0.30 41.4% / $1.73 42.5% / $1.44
OSWorld 2.0 offline / max 64.4% / $3.37 71.4% / $1.27 73.5% / $9.44 Not shown in this chart

Each cell shows score / cost per task, based on the comparison graphs in OpenAI's announcement. Opus 5.5 figures include conditions with fallback to another model, and competitor values are cited from public reports. These are not independent tests measuring all models simultaneously in the same environment, and identical setting names do not necessarily mean equal amounts of compute.

In business automation, the relative position against competing models also changes with the reasoning setting.

Zapier's AutomationBench evaluates whether a model can carry out tasks in sales, marketing, support, finance, and more through to completion using 47 kinds of tools.

At medium, the new Sol beats Opus 5.5 by 2.2 points at about one-third the cost per task. At max, however, Opus 5.5 scores 42.5% versus the new Sol's 36.1%, a 6.4-point deficit.

In other words, strength when used at relatively low cost is not the same as strength when spending more to chase the highest score.

On OSWorld, which involves operating on-screen applications, the new Sol beat the previous Sol by 7.0 points and narrowed the gap to Astra to 2.1 points.

The tasks used were the offline tasks in the August 8, 2026 version of OSWorld 2.0, and partial credit is given for what was achieved along the way, not only for complete success.

Therefore, the 71.4% figure cannot be interpreted as "71.4% of the work was fully completed."

Big gains on science tasks, but a gap to top models remains

On the science benchmark Terminal-Bench Science 0.1, the new Sol at max scored 57.0% versus 27.6% for the previous Sol. The score more than doubled, and average cost fell from $12.18 to $5.47.

Astra, meanwhile, scored 68.1% and Opus 5.5 scored 63.3%.

Terminal-Bench Science 0.1: Big improvement for the new Sol, but a gap to top models remains横棒グラフ。カテゴリ 4 件、系列: Score at max setting(単位: %)GPT-6 SolGPT-6 SolGPT-6 Sol — Score at max setting: 27.6%27.6GPT-6.1 SolGPT-6.1 SolGPT-6.1 Sol — Score at max setting: 57%57Claude Opus 5.5Claude Opus 5.5Claude Opus 5.5 — Score at max setting: 63.3%63.3GPT-6 AstraGPT-6 AstraGPT-6 Astra — Score at max setting: 68.1%68.1単位: %
データを表で見る
Score at max setting (%)
GPT-6 Sol27.6
GPT-6.1 Sol57
Claude Opus 5.563.3
GPT-6 Astra68.1
Terminal-Bench Science 0.1: Big improvement for the new Sol, but a gap to top models remainsComparison announced by OpenAI. Opus 5.5 includes fallback to another model. The agent environment used also differs by model.出典: OpenAI, "Introducing GPT-6.1 Sol", Terminal-Bench Science 0.1

The gap to Astra is 11.1 points, not as close as in the screen-operation benchmark OSWorld. For those who need top-level results on difficult science tasks, there are still reasons to use higher-tier models such as Astra.

Model (all at max) Average cost per task
GPT-6 Sol $12.18
GPT-6.1 Sol $5.47
Claude Opus 5.5 $23.21
GPT-6 Astra $23.80

The new Sol's cost is about 77.0% lower than Astra's and about 76.4% lower than Opus 5.5's.

Considering this price gap together with the score gap, the suitable model differs between repeatedly processing large numbers of tasks and concentrating compute on a small number of difficult ones.

The evaluation's official site lists 70 tasks, including data analysis and simulation. In the results posted there for Astra and Opus 5.5, Codex and Claude Code were used respectively, so this is not a test isolating the capability difference between models alone.

It is also a different evaluation from Terminal-Bench 4.0, which was used in announcements such as Sonnet 5.5's.

Accuracy and safety: not every metric improved

The share of answers containing factual errors improved substantially at low reasoning settings.

In OpenAI's evaluation, the previous Sol at low was at 11.4%, versus 7.7% for the new Sol. At max, both models come in at 4.6%.

Reasoning setting GPT-6 Sol error rate GPT-6.1 Sol error rate GPT-6 Astra error rate
low 11.4% 7.7% 6.3%
xhigh 4.5% 4.1% 4.0%
max 4.6% 4.6% 3.9%

Values from graphs in OpenAI's announcement. This is the share of answers containing at least one factual error; lower is better.

The evaluation used ChatGPT conversations in which earlier models gave wrong answers and users pointed out the errors, with personally identifiable information removed.

Because the questions were selected as ones prone to errors, these figures cannot be taken as the error rate in everyday ChatGPT use.

The results show improvement at low reasoning settings, but do not mean errors fall by the same proportion in every use.

On safety evaluations as well, some metrics improved and others worsened slightly.

In section 7.4.2 of the System Card, the rate at which the model failed to tell the user in its final answer that it had been unable to use a search tool fell from 4.92% for the previous Sol to 2.08% for the new Sol.

On the other hand, in section 7.4.1's coding work, the rate of "misrepresentation," where the model describes something different from what it actually did, was 1.50% for the new Sol, 1.30% for the previous Sol, and 0.51% for Astra.

Both are difficult evaluations designed to deliberately elicit problem behavior, and they do not show that the same rates of problems would occur in real usage environments.

Still, when handing large numbers of tasks to a low-cost model, you need to evaluate separately not only whether the final answer is correct but also whether the model accurately reports when it fails along the way.

AD

Standard pricing matches Sonnet 5.5; cached input is half price

GPT-6.1 Sol and Claude Sonnet 5.5 have the same standard API pricing: $2 per million input tokens and $10 per million output tokens.

GPT-6.1 Sol also lowers cached input to $0.10, half the $0.20 charged for the previous GPT-6 Sol and Sonnet 5.5.

Model Standard input Output Cache read
GPT-6 Sol $2 $10 $0.20
GPT-6.1 Sol $2 $10 $0.10
GPT-6 Astra $10 $50 $1
GPT-6 Luna $0.10 $0.50 $0.01
Claude Sonnet 5.5 $2 $10 $0.20
Claude Opus 5.5 $4 $20 $0.20

Units are US dollars per million tokens. Standard prices based on OpenAI's new Sol announcement, the previous Sol's model documentation, and Anthropic's published pricing. Surcharges for long inputs, Fast mode, cache write fees, and the like are not included.

The "one-fifth" of Astra that OpenAI cites refers to standard input and output prices. For cached input, the new Sol's $0.10 versus Astra's $1 is one-tenth.

The more a workload repeatedly reuses the same long instructions or materials, the greater the benefit of cheaper cached input.

However, this price table alone cannot tell you which is better in performance, the new Sol or Sonnet 5.5.

The evaluations Anthropic has published for Sonnet 5.5 include OSWorld 2.1 and GDPval-AA, which differ from the OSWorld 2.0 and GDP.pdf that OpenAI presented this time. Placing benchmark figures that measure different things side by side does not make a direct performance comparison.

In our article on Sonnet 5.5 as well, what mattered was not only the API unit price but how many tokens and how much money are actually spent to finish the same job.

For this Sol too, actual cost depends not only on the price per million tokens but on how little processing it takes to reach the required quality.

Half-price cache reads do not halve total cost

Prompt caching is a mechanism that reuses the common part of input already processed in subsequent requests.

According to OpenAI's model documentation, the new Sol's cache write price is $2.50 per million tokens, 1.25 times the standard input price of $2. This $2.50 is not a surcharge on top of the standard input price; it is the unit price applied to the input portion being saved to the cache.

Assuming a fixed input of 100,000 tokens saved to the cache first, used 10 times in total, with 1,000 tokens of output each time, the total cost reduction from the previous Sol is about 17% even though the cached input price is halved.

This is a simple estimate using the published prices of the new and previous Sol as of September 30, not a measurement from running the actual models.

It assumes the entire input is saved to the cache once on the first request and fully reused from the second request on, with no cache expiry or changes to the input, and without adding output text to the next input.

Times the same fixed input is used GPT-6 Sol GPT-6.1 Sol No cache (both models)
1 $0.26 $0.26 $0.21
2 $0.29 $0.28 $0.42
10 $0.53 $0.44 $2.10
100 $3.23 $2.24 $21.00

The table shows cumulative totals up to each number of uses, including the initial cache write and output fees for all uses. Tool usage fees, Fast mode, taxes, and discounts are not included.

The formula is:

0.1 × cache write price + 0.1 × (number of uses − 1) × cache read price + 0.001 × number of uses × output price

For 10 uses, the new Sol comes to 0.25 + 0.09 + 0.10 = $0.44, and the previous Sol to 0.25 + 0.18 + 0.10 = $0.53. The reduction is about 17%.

Looking only at the portion read from the cache, the price is halved, but the initial write fee and the output fees are unchanged, so the total does not fall by half.

For input used only once, saving it to the cache actually costs more. For uses that repeatedly draw on the same long materials or system prompt, the effect is larger.

According to the prompt caching developer documentation, reuse requires that the beginning of the input match. Changes to the model, settings, or prompt structure may also cause cache misses.

Simply continuing the same conversation does not guarantee that all input is billed at cache prices.

AD

Reasoning settings and long-input pricing to check before adopting

GPT-6.1 Sol is available to Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. It can also be used through the API as gpt-6.1-sol, but it is not yet available in regular Chat.

In other words, having a subscription to an eligible ChatGPT plan does not mean you can select GPT-6.1 Sol from every screen.

For developers, differences in settings also matter when migrating from the previous Sol.

According to the model documentation, the new Sol's reasoning.effort supports low, medium, high, xhigh, and max, with medium as the default.

The none setting available on the previous Sol cannot be used with GPT-6.1 Sol.

Tool calling requires the Responses API; the Chat Completions API does not support tool calling with GPT-6.1 Sol.

The context window is 1.05 million tokens, and the maximum output is 128,000 tokens.

However, for requests with more than 272,000 input tokens, long-input pricing applies to the entire request: input and cache-related charges double and output charges rise 1.5 times.

When feeding in very large materials at once, you cannot estimate costs using only the standard $2 input and $10 output prices.

Fast mode also costs twice the standard price, so additional charges for large contexts or high-speed processing need to be taken into account.

Regarding the faster GPT-6.1 Sol Ultrafast, as of September 30 there is a difference in wording between OpenAI's Japanese and English announcements.

The Japanese version says GPT-6.1 Sol Ultrafast is "now available," while the English version says it will come to Codex "in the next few days."

It is best to check separately that the standard GPT-6.1 Sol is already available and whether Ultrafast can actually be used.

When adopting the model, first decide what your everyday work must satisfy to count as a pass.

For code fixes, that means test results and scope of changes; for PDF reading, accuracy of evidence; for business automation, whether the required processing was completed correctly to the end. Then compare cost and processing time by reasoning setting.

With the new Sol, more jobs may meet the required quality even at low reasoning settings. If you give those jobs to Sol and reserve higher-tier models for the difficult tasks that need the higher performance of Astra or Opus 5.5, the same budget can cover more processing.