On October 7, 2026, Anthropic announced Claude Haiku 5.5, a small AI model. For inputs of 100,000 tokens or fewer, its API input and output prices are one-tenth those of Haiku 4.5, putting it on par with OpenAI's GPT-6 Luna.
In evaluations published by Anthropic, the model also improved substantially over the previous Haiku in computer use and knowledge work. However, once an input exceeds 100,000 tokens, the price rises fivefold, and when migrating from the older model, the same text is counted as a different number of tokens.
The range of work that can be entrusted to a low-priced model has grown. But deciding between Luna and the higher-tier Sonnet requires looking beyond the simple API price to the length of your inputs and the difficulty of the work.
What's the difference between "90% lower unit prices" and "75% lower average costs"?
Haiku 5.5's standard API pricing starts at $0.10 per million input tokens and $0.50 per million output tokens.
A token is the unit an AI uses to process text, and it does not match character or word counts. Anthropic's announcement says it cut prices by 90% from the previous Haiku for inputs of 100,000 tokens or fewer, and by 50% for inputs above 100,000 tokens.
| Model and input condition | Input | Output | Cache read |
|---|---|---|---|
| Haiku 5.5 / up to 100K tokens | $0.10 | $0.50 | $0.01 |
| Haiku 5.5 / over 100K tokens | $0.50 | $2.50 | $0.05 |
| Haiku 4.5 | $1 | $5 | $0.10 |
| GPT-6 Luna / up to 272K tokens | $0.10 | $0.50 | $0.01 |
| Claude Sonnet 5.5 | $2 | $10 | $0.10 |
| Claude Opus 5.5 | $4 | $20 | $0.20 |
| GPT-6.1 Sol / up to 272K tokens | $2 | $10 | $0.10 |
Units are US dollars per million tokens. Prices are standard API rates as of October 8, 2026. Claude figures are based on the announcement and pricing documentation; OpenAI figures are based on the model documentation for Luna and Sol. Cache writes, region selection, search, and other add-on charges are not included. Sonnet's cache read price reflects the cut made on October 7.
For inputs of 100,000 tokens or fewer, Haiku's input and output prices are one-twentieth of Sonnet's. Above 100,000 tokens, Haiku's prices are one-quarter of Sonnet's and one-eighth of Opus's.
To judge how cheap it really is, you first need to check which pricing tier your inputs fall into.
Meanwhile, the "average cost reduction of about 75%" that Anthropic cites means something different from the 90% unit price cut.
According to the company, about 90% of requests to the previous Haiku had inputs of 100,000 tokens or fewer. But the average cost calculation also reflects changes in how many tokens the new model uses to complete the same task.
Even if 90% of requests are short inputs, they do not necessarily account for 90% of total billing. That means the "75% average reduction" cannot be applied directly to your own workload.
According to the model specifications, Haiku 5.5 has a context window of 1 million tokens and a standard maximum output of 128,000 tokens. That is up from 200,000 and 64,000 for the previous Haiku.
It can now handle large volumes of material at once, but that does not mean the lowest price applies all the way up to 1 million tokens.
Same unit price as GPT-6 Luna, but the gap widens with long inputs
For GPT-6 Luna, a surcharge applies once input exceeds 272,000 tokens.
Beyond that threshold, input and cache prices double and the output price rises 1.5 times for the entire request, to $0.20 for input and $0.75 for output.
Haiku's pricing threshold is 100,000 tokens, versus 272,000 for Luna. Even among models in the same low-price tier, this difference has a large effect on the cost of handling long inputs.
Using Anthropic's pricing and OpenAI's pricing terms and varying only input length, we get the following estimates.
Output is fixed at 20,000 billable tokens, including reasoning, for a single standard API call with no caching.
| Input tokens per call | Haiku 5.5 | GPT-6 Luna | Haiku / Luna |
|---|---|---|---|
| 100K | $0.020 | $0.020 | 1x |
| 200K | $0.15 | $0.03 | 5x |
| 300K | $0.20 | $0.075 | about 2.67x |
For a single standard API call using 200,000 input and 20,000 output tokens, Haiku 5.5 costs $0.15 and GPT-6 Luna costs $0.03, so by this estimate Haiku is five times more expensive.
The formula is: input price × input tokens ÷ 1 million + output price × 20,000 ÷ 1 million.
For Haiku with 200,000 input tokens:
0.50 × 0.20 + 2.50 × 0.02 = $0.15
For Luna:
0.10 × 0.20 + 0.50 × 0.02 = $0.03
When input grows to 300,000 tokens, Luna's surcharge also applies, so the gap narrows.
However, this is a price calculation assuming identical token counts, not a measured result from running the same task on both models.
Even the same text is counted as different numbers of tokens by different models, and the tokens used for answers and reasoning also vary. Search fees, cache writes, region selection, and other add-on charges are not included either.
To compare the actual cost per task, you need to check the number of failed attempts and reruns as well.
When migrating from the previous Haiku, the way input tokens are counted also changes.
According to the migration guide, the new tokenizer counts the same input text as about 30% more tokens than the older model.
The actual increase varies by content, but for inputs near a pricing threshold, this difference can change which pricing tier applies.
For example, suppose a text that was 80,000 tokens on the previous Haiku grows 30% to 104,000 tokens under the new tokenizer.
Looking at input cost alone, the old model costs:
1 × 0.08 = $0.08
The new model, because the text exceeds 100,000 tokens, costs:
0.50 × 0.104 = $0.052
The reduction is:
1 − 0.052 ÷ 0.08 = 35%
This is a hypothetical example that excludes output and reasoning costs, but the savings are far smaller than the "90% lower unit price" figure might suggest.
When migrating, don't reuse input lengths measured on the old model; recount them with Haiku 5.5's tokenizer.
For agents that accumulate not just document text but also tool descriptions and conversation history, repeated processing can also push usage past the 100,000-token pricing threshold.
Some published figures beat Luna, but a gap with Sonnet remains in complex development
In the comparison table Anthropic published on October 7, Haiku 5.5 scored higher than Luna on every evaluation the two models share.
In OSWorld 2.1, which evaluates computer use, Haiku 5.5 scored 72.4% on a subset of offline tasks, 23.5 points above Luna's 48.9%.
However, the gap with higher-tier models varies widely by type of work.
| Evaluation and what it measures | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 / relative rating of knowledge work | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 / relative rating of knowledge work | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1 / computer use, subset of offline tasks | 72.4% | 15.7% | 48.9% | 83.9% |
| Terminal-Bench 4.0 / multi-step terminal work | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 Main / code changes | 46.4% | Not listed | 42.4% | 52.1% (xhigh) |
| Chartography / chart and figure reasoning, no tools | 46.4% | 6.4% | 29.1% | 61.6% |
| Humanity's Last Exam / cross-domain reasoning, no tools | 45.9% | 10.2% | Not listed | 56.9% |
| Humanity's Last Exam / with tools | 57.4% | 18.7% | Not listed | 64.5% |
Source: Anthropic's Haiku 5.5 announcement.
The GDPval-AA and AA-Briefcase figures are relative ratings, not accuracy rates or the percentage of tasks a model can handle. "Not listed" means the figure does not appear in Anthropic's announcement table.
Sonnet's FrontierCode result was also measured with the xhigh reasoning setting, so it is not a comparison at the same reasoning effort across all models.
In Terminal-Bench 4.0, the gap between Haiku 5.5 and Sonnet 5.5 widens to 31.4 points.
Haiku 5.5's 39.2% shows a major improvement over the previous Haiku's 0%, but that 0% is a result on this benchmark only and does not mean the older model "couldn't write code."
Anthropic itself recommends Sonnet 5.5 or Opus 5.5 for work such as long-running, complex software development.
So the fact that Haiku 5.5 has published figures above Luna's is a separate matter from whether it can handle every kind of development work in place of Sonnet.
If it meets the quality you need for tasks such as reading charts, classifying information under set conditions, or performing computer operations with limited steps, there is considerable room to take advantage of Haiku's low price.
The timing of the comparison also deserves attention.
On September 25, OpenAI fixed an image-processing bug that had impaired image understanding in GPT-6 Luna and GPT-6 Sol.
It is not stated whether the Luna vision results in Anthropic's table were measured before or after that fix.
We also cannot confirm that conditions such as reasoning settings and measurement error were all aligned. The table therefore cannot be treated as a ranking of the models for every current use case.
The same goes for speed.
Anthropic describes Haiku 5.5 as the "fastest" only in comparison with the standard modes of Claude models, and a footnote says it is slower than Opus's Fast Mode.
How quickly a model generates text and the total processing time for a job, including tool operations and waiting, should be evaluated separately.
How to choose between low-priced Haiku and price-cut Sonnet
Haiku 5.5 is the first model in the Haiku series to let users adjust reasoning strength.
The default is medium, under which the model reasons as needed.
Lowering the reasoning setting reduces how much the model thinks, and for simple tasks it may skip reasoning altogether. At high or below, reasoning can also be disabled.
There is no need to use the same reasoning setting for simple classification or extraction as for summarization that involves judgment.
Anthropic lists summarization, compressing conversation history, and database queries among the intended uses.
It also envisions setups in which a higher-tier model decides the development approach while Haiku is assigned to investigating target files or extracting information into a set format.
To make such a division of labor practical, you need to decide in advance how much to hand off to Haiku and under what conditions to return processing to a higher-tier model.
Tasks that can be completed with short instructions are likely to stay in the low-price tier. But if you pass the full conversation history or all materials to Haiku every time, inputs will easily exceed 100,000 tokens, and the pricing advantage shrinks.
Sonnet also received a price cut on the same day.
The cache read price was halved from $0.20 to $0.10 per million tokens, so Sonnet now matches GPT-6.1 Sol's standard pricing on all three items: input, output, and cache read.
Using the $0.20 Sonnet figure listed in past comparison tables would lead to misreading the current price gap.
However, halving the cache read price does not mean total costs for a job are halved.
Caching is a mechanism that stores input that has already been processed so it can be reused in later requests.
Consider a process spanning multiple calls in total, in which 90% of 1 million input tokens are read from cache and 10,000 output tokens are used. The initial cost of writing to the cache is excluded.
At the old price:
2 × 0.10 + 0.20 × 0.90 + 10 × 0.01 = $0.48
At the new price:
2 × 0.10 + 0.10 × 0.90 + 10 × 0.01 = $0.39
That is a reduction of 18.75%.
This is close to Anthropic's figure of "about 20% savings for many agent tasks," but it is not a reproduction of the company's actual measurements.
If caching is not used at all, the cost for the same input and output is $2.10 both before and after, so the price cut has no effect.
The cheaper a model's per-token price becomes, the larger the share of the bill taken up by separate charges such as search.
Claude's search pricing is $10 per 1,000 searches, and token charges for processing the retrieved content are billed separately.
That comes to $0.01 per search.
The Haiku cost of 100,000 input and 20,000 output tokens calculated above is $0.02, so a single search adds an extra charge equal to half of that.
As token prices fall, the number of searches and the number of times the same material is reloaded have a greater impact on total cost.
Settings to review when migrating the API, and monthly credits for Max and Team
Haiku 5.5 is available on AWS, Google Cloud, and Microsoft Azure in addition to the Claude Platform.
The model name used in the Claude API is claude-haiku-5-5.
However, if you only change the model name in code that called the previous Haiku, it may produce an error depending on your settings.
The migration guide asks users to move from the older approach of specifying budget_tokens in thinking to the new reasoning settings.
For temperature, top_p, and top_k, it recommends omitting them rather than carrying over the old values.
max_tokens, which sets the maximum response size, includes reasoning tokens. If the limit is set too low, reasoning alone can use up the allowance and no final answer text may be returned.
The model may also refuse to respond because of safety measures, so apps need to handle stop_reason: "refusal" appropriately.
Haiku 5.5 has no mechanism for automatically passing refused requests to another model on the server side.
When adopting it, you need to check not only the success rate under normal conditions but also how to handle refusals and insufficient output.
Monthly API credits are also being provided to Max and Team plans.
Max 5x gets $100 and Max 20x gets $200.
For example, with three Standard seats and two Premium seats:
20 × 3 + 100 × 2 = $260
Not every Team contract receives a flat $500.
Credits are linked to and received through an organization in the Claude Console.
They apply to API charges for things like custom-built apps and agents, and cannot be used for interactive Claude Code or for additional usage charges beyond a plan's usage allowance.
The credits are rolled out gradually over several days, and conditions apply for the notice to appear, such as having used an eligible plan for seven days.
They are credits to support development and testing with the API, and do not mean the amount available in regular Claude chat increases as is.
When adopting Haiku 5.5, a straightforward approach is to prepare the same materials and acceptance criteria, first vary the input length, then vary the reasoning strength, and compare quality and actual billed amounts.
Then decide which tasks to entrust to Haiku and under what conditions to return to Sonnet or Opus.
Doing so can turn Haiku 5.5's sharp price cut into more work that can be processed at the required quality, rather than merely a cheaper token price.
