On September 28, Anthropic announced a new AI model, Claude Sonnet 5.5. According to the company, it keeps API token prices unchanged from the previous generation, Sonnet 5, while increasing output speed by more than 30% and cutting the cost of completing the same task by up to 30%.

The update is aimed at handling everyday work, such as drafting documents and slides or making well-scoped code fixes, in less time and with fewer tokens. However, the published evaluation results include cases where a setting one step below the maximum reasoning level scored higher. When comparing Sonnet 5.5 with the higher-tier Opus 5.5, or with GPT-6 Sol, which has the same standard API pricing, you need to look beyond benchmark scores to reasoning settings and actual processing costs.

AD

Same prices, faster output and better token efficiency

Sonnet 5.5's standard API pricing is $2 per million input tokens and $10 per million output tokens. The read price for reusing cached input is $0.20, and all of these are the same as Sonnet 5. According to Anthropic's announcement, the cost reduction does not come from a price cut but from reducing the number of tokens needed to complete the same task.

The "more than 30% faster" figure also refers to the speed at which the model generates output. It does not mean that the entire process, including time spent waiting for search or external tool responses, is uniformly shortened by 30%.

The "up to 30% cost reduction" is likewise based on Anthropic's own testing. Actual bills will vary with the volume of input material, the number of reasoning tokens used, the number of tool calls, and other factors.

Anthropic positions Sonnet 5.5 as suited to tasks with relatively clear goals and scope, such as bug fixes and everyday document work. Opus 5.5, by contrast, is intended for complex work that requires sustained, careful judgment over long periods.

Sonnet 5.5 is now available on platforms including AWS, Google Cloud, and Microsoft Azure. Haiku 5.5, a low-cost model for high-volume processing, is due to join within the next few weeks.

Performance compared with the previous Sonnet, Opus, and GPT-6 Sol

On Terminal-Bench 4.0, which tests multi-step work on the command line, Sonnet 5.5 scored 70.6% versus 10.3% for Sonnet 5, a large improvement over the older model.

That said, the fact that Opus 5.5 scored 66.4% in the same table does not mean Sonnet 5.5 is better across all software development.

The table below organizes the evaluations and conditions based on the comparison table Anthropic published. Claude figures are generally from the maximum reasoning setting, max, though there are some exceptions. The table also includes results measured by outside organizations, so it is not a comparison of all models under identical environments and conditions.

Evaluation / what it measures Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
Terminal-Bench 4.0 / multi-step command-line work 70.6% 10.3% 66.4% (xhigh) Not listed
FrontierCode 1.1 Main / code change quality 46.2% (max), 52.1% (xhigh) 42.4% 54.4% 49.3%
CursorBench 4.0 / work close to real development 55.5% 34.1% 57.8% Not listed
GDPval-AA v2.1 / occupation-based deliverables, Elo 1844 1449 1846 1487
AA-Briefcase v1.1 / extended knowledge work, Elo 1811 1359 1822 1483
Humanity's Last Exam / cross-domain reasoning, with tools 64.5% 54.9% 67.7% Not listed
OSWorld 2.1 / screen operation, partial credit 80.1% 57.0% 81.8% Not listed
Chartography / reading specialized charts, no tools 61.6% 15.6% 64.4% 53.6%

Elo is a relative rating calculated by comparing deliverables produced by multiple models; it is not an accuracy rate. "Not listed" means no figure appears in Anthropic's comparison table. It does not mean the model failed the evaluation or does not support it.

The GPT-6 Sol figures are drawn from externally published evaluation results and were not necessarily measured with the same amount of reasoning or execution conditions as Claude.

According to Section 8.5 of the system card, Terminal-Bench 4.0 used 66 tasks, with Sonnet 5.5 and Opus 5.5 each run five times, for 330 runs in total.

Sonnet 5.5 was run at max and Opus 5.5 at the one-step-lower xhigh. Assuming each trial is independent, the standard errors were ±2.5 points and ±2.6 points, respectively. The score gap is 70.6 − 66.4 = 4.2 points, but because the reasoning settings and measurement error differ, this result alone cannot establish which model is better overall.

The evaluation used an environment with no external network access and with the necessary resources prepared in advance. For Sonnet 5.5, safety features triggered on 1.2% of requests, and cases where the system switched to an older model to respond are included in the results. It is also worth noting that this is not a test that faithfully reproduces real development environments.

When comparing with other companies' models, timing of measurement also matters.

On September 25, OpenAI fixed an image-processing bug that had degraded GPT-6 Sol's image understanding. It cannot be confirmed whether the 53.6% Chartography figure in the table was measured before or after the fix, so it is difficult to treat it as a direct comparison of current visual performance.

In addition, no GPT-6 Sol result has been published for Terminal-Bench 4.0. Filling that gap with the older GPT-5.6 Sol's figure would not be appropriate either.

AD

Does maximum reasoning always mean better performance?

On FrontierCode 1.1 Main, Sonnet 5.5 scored 52.1% at xhigh, but 46.2% at max, which uses even more reasoning. That is a gap of 5.9 points.

In other words, making the model think more does not always raise the score.

FrontierCode 1.1 Main:推論設定で変わる得点横棒グラフ。カテゴリ 5 件、系列: 評価スコア(単位: %)Sonnet 5(max)Sonnet 5(max)Sonnet 5(max) — 評価スコア: 42.4%42.4Sonnet 5.5(max)Sonnet 5.5(max)Sonnet 5.5(max) — 評価スコア: 46.2%46.2Sonnet 5.5(xhigh)Sonnet 5.5(xhigh…Sonnet 5.5(xhigh) — 評価スコア: 52.1%52.1Opus 5.5(max)Opus 5.5(max)Opus 5.5(max) — 評価スコア: 54.4%54.4GPT-6 Sol(公表値)GPT-6 Sol(公表値…GPT-6 Sol(公表値) — 評価スコア: 49.3%49.3単位: %
データを表で見る
評価スコア (%)
Sonnet 5(max)42.4
Sonnet 5.5(max)46.2
Sonnet 5.5(xhigh)52.1
Opus 5.5(max)54.4
GPT-6 Sol(公表値)49.3
FrontierCode 1.1 Main:推論設定で変わる得点Cognitionによる評価。ClaudeはClaude Code、GPTはCodex CLIを使用。推論量をそろえた実測比較ではない。出典: Anthropic Sonnet 5.5 System Card 8.1・8.4節

At max, Sonnet 5.5 falls below GPT-6 Sol's published figure, while at xhigh it exceeds it. A simple ranking that lines up only model names hides differences like these caused by reasoning settings.

FrontierCode, developed by Cognition, evaluates 150 software development tasks on not only whether features work correctly but also whether changes stay within the requested scope and whether the code meets a certain quality bar.

Changes beyond the requested scope are penalized even if the changes themselves are useful. This mechanism alone cannot fully explain why Sonnet 5.5 scored lower at max, but the test is not one where spending more time and tokens to make more changes earns a higher rating.

In evaluations such as document creation, on the other hand, more reasoning pays off in a different way.

In GDPval-AA and AA-Briefcase, which Artificial Analysis ran independently and Anthropic included in the system card, Sonnet 5.5 running at max came close to Opus 5.5.

  • GDPval-AA v2.1
  • AA-Briefcase v1.1
知識労働の評価:最大設定でOpus 5.5に接近横棒グラフ。カテゴリ 4 件、系列: GDPval-AA v2.1, AA-Briefcase v1.1(単位: Elo)Sonnet 5Sonnet 5Sonnet 5 — GDPval-AA v2.1: 1,449Elo1,449Sonnet 5 — AA-Briefcase v1.1: 1,359Elo1,359Sonnet 5.5Sonnet 5.5Sonnet 5.5 — GDPval-AA v2.1: 1,844Elo1,844Sonnet 5.5 — AA-Briefcase v1.1: 1,811Elo1,811Opus 5.5Opus 5.5Opus 5.5 — GDPval-AA v2.1: 1,846Elo1,846Opus 5.5 — AA-Briefcase v1.1: 1,822Elo1,822GPT-6 SolGPT-6 SolGPT-6 Sol — GDPval-AA v2.1: 1,487Elo1,487GPT-6 Sol — AA-Briefcase v1.1: 1,483Elo1,483単位: Elo
データを表で見る
GDPval-AA v2.1 (Elo)AA-Briefcase v1.1 (Elo)
Sonnet 51,4491,359
Sonnet 5.51,8441,811
Opus 5.51,8461,822
GPT-6 Sol1,4871,483
知識労働の評価:最大設定でOpus 5.5に接近Artificial Analysisの結果をAnthropicが掲載。Claudeはmax。Eloは各評価内の相対値で、異なる評価間の点差・倍率を比べる指標ではない。出典: Anthropic発表・Sonnet 5.5 System Card 8.14.3〜8.14.4節

On GDPval-AA the gap between Sonnet 5.5 and Opus 5.5 is 2 points, and on AA-Briefcase Sonnet 5.5 scored 1811 against Opus 5.5's 1822. However, an Elo gap is not a difference in accuracy, and it does not mean the two models produce equivalent results on every real-world task.

GDPval-AA compares deliverables such as documents and spreadsheets across 220 tasks spanning 44 occupations and 9 industries.

Lowering Sonnet 5.5 to xhigh gives an Elo of 1725, while output tokens fall by about 67% compared with max. On AA-Briefcase, xhigh scored 1746 against 1811 at max, while cutting output tokens by about 61%.

The right reasoning setting depends on whether you are aiming for the highest possible score or want to handle more work at lower cost while maintaining sufficient quality.

Anthropic also notes that Sonnet 5.5 in both evaluations was measured in a pre-release environment that had a structured-output bug. The bug has since been fixed, but the scores listed here are not values re-measured after the fix.

Half the API price doesn't necessarily mean half the processing cost

GPT-6 Sol, which appeared on September 22, and Sonnet 5.5 have identical standard API prices for input, output, and cache reads.

Opus 5.5, meanwhile, charges twice as much as Sonnet 5.5 for standard input and output, but its cache read price is the same $0.20 per million tokens.

For workloads that repeatedly reuse long documents or conversation history, this difference in pricing structure matters. According to Claude's model overview, both Sonnet 5.5 and Opus 5.5 support a context window of up to 1 million tokens.

Model Input Output Cache read
Claude Sonnet 4.6 $3 $15 $0.30
Claude Sonnet 5 $2 $10 $0.20
Claude Sonnet 5.5 $2 $10 $0.20
Claude Opus 5 $5 $25 $0.50
Claude Opus 5.5 $4 $20 $0.20
GPT-6 Sol $2 $10 $0.20

Units are US dollars per million tokens. These are standard API prices based on Claude's pricing page, the Sonnet 5.5 announcement, and GPT-6 Sol's model documentation, not monthly subscription plan prices. Cache writes, tool use, and regional surcharges or discounts are not included. For GPT-6 Sol, the prices shown apply when input is 272,000 tokens or fewer.

Assuming 1 million input tokens and 10,000 output tokens in total across multiple calls, if 90% of the input is read from cache, Opus 5.5 costs about 1.63 times as much as Sonnet 5.5, not twice as much.

This is an estimate using API prices as of September 28, not a result from having both models do actual work. Input across multiple API calls is fixed at 1 million tokens in total and output at 10,000 tokens, and only the share of input read from cache is varied.

Share of input read from cache Sonnet 5.5 calculated cost Opus 5.5 calculated cost Opus / Sonnet
0% $2.10 $4.20 2x
90% $0.480 $0.780 about 1.63x
99% $0.318 $0.438 about 1.38x

The formula is: input price × uncached share + cache read price × cached share + output price × 0.01.

For the 90% row, Sonnet 5.5 is

2×(1−0.9)+0.20×0.9+10×0.01=$0.48

and Opus 5.5 is

4×(1−0.9)+0.20×0.9+20×0.01=$0.78

making the ratio

0.78÷0.48=1.625x.

However, this estimate does not include the cost of initially creating or refreshing the cache, so it does not represent actual total billing.

Also, even if you give both models the same job, they will not necessarily use the same number of input and output tokens or produce final deliverables of the same quality.

Still, it shows that even though standard input and output prices are twice as high, the real cost of using Opus 5.5 is not always double that of Sonnet 5.5.

Conversely, for workloads that barely use caching and generate large amounts of output, Sonnet 5.5's lower standard prices tend to translate directly into savings.

GPT-6 Sol also has a pricing threshold. When input in a single request exceeds 272,000 tokens, input and cache prices double and the output price rises 1.5 times for the entire request.

For uses that feed in very long documents or large volumes of material at once, you cannot compare costs by looking at standard prices alone.

Caution is also needed when simply comparing prices if migrating from the older Sonnet 4.6.

According to the migration guide, Sonnet 5.5 uses the same tokenization as Sonnet 5, which means the same text yields roughly 30% more tokens than with Sonnet 4.6.

The rate of increase varies with the content of the text, so when comparing actual costs in a migration from an older generation, you need to re-measure token counts using your own data.

AD

Migrating your API: watch reasoning settings and safety features

Sonnet 5.5's default reasoning setting is medium in the Claude app and Claude Code, and high in the Claude Platform API. Even with the same model, the default differs depending on the environment.

The migration guide recommends starting with medium for coding with clear goals and scope and for multi-step tool use, and trying high for harder or longer-running work.

If your existing code disabled reasoning, you cannot migrate just by changing the model name.

Sending

thinking: {"type": "disabled"},

which worked with Sonnet 5, directly to Sonnet 5.5 returns a 400 error.

The new setting for no upfront reasoning is between_tools, which can be combined with the low, medium, and high reasoning settings. It cannot be used with xhigh or max.

Also, in tool-using workflows, thinking blocks may be returned between tool calls. max_tokens covers both the tokens used for reasoning and the final answer, and reasoning tokens are billed as output tokens.

As capabilities have improved, safety features in the cybersecurity area have also been strengthened.

Anthropic assesses that Sonnet 5.5's cyber capabilities have reached a level close to Opus 5, and has introduced a mechanism that switches to Sonnet 5 to respond to some requests judged high-risk.

Anthropic says this does not affect normal software development, but for security-related uses, the specified Sonnet 5.5 will not necessarily respond directly to every request.

If you are considering adoption, it would be sensible to prepare the same materials and acceptance criteria and compare output quality, time to completion, and actual billing for each reasoning setting.

If you find tasks where Sonnet 5.5 at low or medium still maintains the required quality, the time and cost saved there can be redirected to work that requires more careful judgment. The value of Sonnet 5.5 lies less in simply getting closer to Opus-level scores than in being able to finely choose the balance of quality, speed, and cost to suit the job.