Skip to content
Amal Hashim
All posts
News

Claude Haiku 5.5: the small model just got serious

Anthropic's new Haiku is up to 90% cheaper than Haiku 4.5, gets a 1M context window and adaptive thinking - and brings a few breaking changes worth knowing before you switch.

5 min read#AI#Anthropic#Claude

Anthropic released Claude Haiku 5.5 on October 7, 2026. Haiku has always been the "fast and cheap" tier - the model you reach for when you need to classify, route or summarise a lot of things without thinking too hard about the bill. This release keeps that job description but moves the ceiling a long way up.

The short version

  • Price: $0.10 / $0.50 per million input / output tokens for prompts up to 100K tokens. Above 100K it's $0.50 / $2.50. Haiku 4.5 was $1 / $5.
  • Context: 1M tokens (up from 200K), and up to 128K output tokens (up from 64K).
  • Thinking: adaptive thinking is on by default, steered with the effort parameter - a first for a Haiku model. Default effort is medium.
  • Model ID: claude-haiku-5-5 - the same ID on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS; anthropic.claude-haiku-5-5 on Amazon Bedrock.

Anthropic puts the saving at "around 75% less to run" on average, and 90% lower list prices for requests up to 100K tokens.

How much better is it?

Anthropic's own benchmark table, against Haiku 4.5 and the mid-tier Sonnet 5.5:

BenchmarkHaiku 5.5Haiku 4.5Sonnet 5.5
OSWorld 2.1 (computer use)72.4%15.7%83.9%
Terminal-Bench 4.039.2%0.0%70.6%
Humanity's Last Exam (no tools)45.9%10.2%56.9%
GDPval-AA v2.116207351840

The computer-use jump is the one that stands out: from barely usable to within about 11 points of Sonnet 5.5. Terminal-style agentic coding is still clearly Sonnet territory, though.

Where I'd use it

  • Subagents. Let Opus 5.5 or Sonnet 5.5 plan and judge, and hand the reading-heavy, repetitive work - searching files, extracting fields, summarising pages - to Haiku. Anthropic positions it exactly for this.
  • Classification, routing and extraction at volume, where latency matters more than depth.
  • Browser and computer-use tasks that used to need a bigger model.

Do the maths with the new tokenizer

Haiku 5.5 uses the newer tokenizer from Claude 4.7 onwards, so the same text counts as about 30% more tokens than on Haiku 4.5. Thinking tokens are billed as output, too. Take a daily job of 10M input and 2M output tokens (measured on Haiku 4.5), all in prompts under 100K:

  • Haiku 4.5: 10 × $1 + 2 × $5 = $20.00
  • Haiku 5.5, before thinking: 13 × $0.10 + 2.6 × $0.50 = $2.60

Thinking at higher effort will add to that, which is why Anthropic's average is 75% rather than 90%. Measure on your own traffic before you promise anyone a number.

Breaking changes from Haiku 4.5

If you're switching an existing integration, check these first:

  1. No budget_tokens. Manual extended thinking returns an error. Use adaptive thinking and set effort instead.
  2. No sampling parameters. A non-default temperature, top_p or top_k returns a 400. Remove them.
  3. No assistant prefill. End messages with a user turn; use structured outputs if you were prefilling to force JSON.
  4. Responses can start with a thinking block. Select content blocks by type, not by position.
  5. Thinking text is omitted by default. Set thinking.display to "summarized" if you want to see it.
  6. Computer use needs computer_toolset_20260801 on the Claude API and Google Cloud.
  7. Keep conversations append-only if you send thinking blocks back - editing earlier turns invalidates them.
  8. Handle stop_reason: "refusal". Safety classifiers can now decline a request, and server-side fallback isn't available for Haiku.

A minimal TypeScript call that respects all of the above:

ts
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();

const res = await client.messages.create({
  model: "claude-haiku-5-5",
  max_tokens: 4096, // thinking tokens count toward this
  output_config: { effort: "low" }, // cheap and fast for routing work
  messages: [{ role: "user", content: "Classify this ticket: ..." }],
});

if (res.stop_reason === "refusal") {
  // handle the decline
}
const text = res.content.find((b) => b.type === "text");

Scenario: moving a ticket router to Haiku 5.5

Say you run a support inbox where every new ticket is tagged as billing, bug or how-to and routed to a queue. It runs on Haiku 4.5 today.

  1. Pick a sample. Export 200 recent tickets with the tag a human finally gave them. That's your eval set.
  2. Swap the model. Change the model to claude-haiku-5-5, remove temperature and any prefill, and set output_config: { effort: "low" }. Routing doesn't need deep thought.
  3. Fix the parsing. Read the answer with res.content.find((b) => b.type === "text"), since the first block may now be thinking.
  4. Compare. Run all 200 tickets through both models. Check accuracy, p95 latency and the cost from res.usage.
  5. Decide. If accuracy holds at low, ship it. If a category slips, try medium on just that route before reaching for Sonnet.
  6. Roll out gradually. Send 10% of live traffic to Haiku 5.5 for a day, watch refusals and misroutes, then move the rest.

My take

For a long time the advice was "Haiku for toys, Sonnet for real work." With 1M context, adaptive thinking and computer-use scores that close most of the gap, Haiku 5.5 is the first small model I'd put in front of production traffic by default - and only step up to Sonnet or Opus where the evals say I have to.

Sources: Anthropic announcement · Haiku 5.5 docs · What's new in Haiku 5.5

Original source · anthropic.com