Claude Haiku 5.5: the small model just got serious
Anthropic's new Haiku is up to 90% cheaper than Haiku 4.5, gets a 1M context window and adaptive thinking - and brings a few breaking changes worth knowing before you switch.
Anthropic released Claude Haiku 5.5 on October 7, 2026. Haiku has always been the "fast and cheap" tier - the model you reach for when you need to classify, route or summarise a lot of things without thinking too hard about the bill. This release keeps that job description but moves the ceiling a long way up.
The short version
- Price: $0.10 / $0.50 per million input / output tokens for prompts up to 100K tokens. Above 100K it's $0.50 / $2.50. Haiku 4.5 was $1 / $5.
- Context: 1M tokens (up from 200K), and up to 128K output tokens (up from 64K).
- Thinking: adaptive thinking is on by default, steered with the
effortparameter - a first for a Haiku model. Default effort ismedium. - Model ID:
claude-haiku-5-5- the same ID on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS;anthropic.claude-haiku-5-5on Amazon Bedrock.
Anthropic puts the saving at "around 75% less to run" on average, and 90% lower list prices for requests up to 100K tokens.
How much better is it?
Anthropic's own benchmark table, against Haiku 4.5 and the mid-tier Sonnet 5.5:
| Benchmark | Haiku 5.5 | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| OSWorld 2.1 (computer use) | 72.4% | 15.7% | 83.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 70.6% |
| Humanity's Last Exam (no tools) | 45.9% | 10.2% | 56.9% |
| GDPval-AA v2.1 | 1620 | 735 | 1840 |
The computer-use jump is the one that stands out: from barely usable to within about 11 points of Sonnet 5.5. Terminal-style agentic coding is still clearly Sonnet territory, though.
Where I'd use it
- Subagents. Let Opus 5.5 or Sonnet 5.5 plan and judge, and hand the reading-heavy, repetitive work - searching files, extracting fields, summarising pages - to Haiku. Anthropic positions it exactly for this.
- Classification, routing and extraction at volume, where latency matters more than depth.
- Browser and computer-use tasks that used to need a bigger model.
Do the maths with the new tokenizer
Haiku 5.5 uses the newer tokenizer from Claude 4.7 onwards, so the same text counts as about 30% more tokens than on Haiku 4.5. Thinking tokens are billed as output, too. Take a daily job of 10M input and 2M output tokens (measured on Haiku 4.5), all in prompts under 100K:
- Haiku 4.5: 10 × $1 + 2 × $5 = $20.00
- Haiku 5.5, before thinking: 13 × $0.10 + 2.6 × $0.50 = $2.60
Thinking at higher effort will add to that, which is why Anthropic's average is 75% rather than 90%. Measure on your own traffic before you promise anyone a number.
Breaking changes from Haiku 4.5
If you're switching an existing integration, check these first:
- No
budget_tokens. Manual extended thinking returns an error. Use adaptive thinking and seteffortinstead. - No sampling parameters. A non-default
temperature,top_portop_kreturns a 400. Remove them. - No assistant prefill. End
messageswith a user turn; use structured outputs if you were prefilling to force JSON. - Responses can start with a
thinkingblock. Select content blocks bytype, not by position. - Thinking text is omitted by default. Set
thinking.displayto"summarized"if you want to see it. - Computer use needs
computer_toolset_20260801on the Claude API and Google Cloud. - Keep conversations append-only if you send thinking blocks back - editing earlier turns invalidates them.
- Handle
stop_reason: "refusal". Safety classifiers can now decline a request, and server-side fallback isn't available for Haiku.
A minimal TypeScript call that respects all of the above:
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const res = await client.messages.create({
model: "claude-haiku-5-5",
max_tokens: 4096, // thinking tokens count toward this
output_config: { effort: "low" }, // cheap and fast for routing work
messages: [{ role: "user", content: "Classify this ticket: ..." }],
});
if (res.stop_reason === "refusal") {
// handle the decline
}
const text = res.content.find((b) => b.type === "text");
Scenario: moving a ticket router to Haiku 5.5
Say you run a support inbox where every new ticket is tagged as billing, bug or how-to and routed to a queue. It runs on Haiku 4.5 today.
- Pick a sample. Export 200 recent tickets with the tag a human finally gave them. That's your eval set.
- Swap the model. Change the model to
claude-haiku-5-5, removetemperatureand any prefill, and setoutput_config: { effort: "low" }. Routing doesn't need deep thought. - Fix the parsing. Read the answer with
res.content.find((b) => b.type === "text"), since the first block may now bethinking. - Compare. Run all 200 tickets through both models. Check accuracy, p95 latency and the cost from
res.usage. - Decide. If accuracy holds at
low, ship it. If a category slips, trymediumon just that route before reaching for Sonnet. - Roll out gradually. Send 10% of live traffic to Haiku 5.5 for a day, watch refusals and misroutes, then move the rest.
My take
For a long time the advice was "Haiku for toys, Sonnet for real work." With 1M context, adaptive thinking and computer-use scores that close most of the gap, Haiku 5.5 is the first small model I'd put in front of production traffic by default - and only step up to Sonnet or Opus where the evals say I have to.
Sources: Anthropic announcement · Haiku 5.5 docs · What's new in Haiku 5.5