Anthropic shipped Claude Haiku 5.5 on October 7, model ID claude-haiku-5-5. The small model gets a 1M token context window, 128,000 max output tokens, and the first adjustable effort setting in the Haiku line. Running it costs about 75% less than Haiku 4.5.
Two price tiers instead of one
Haiku 5.5 bills differently depending on whether a prompt sits under or over 100,000 tokens. Below that line, a million input tokens cost $0.10; above it, $0.50. Output runs $0.50 or $2.50. Cache reads are a cent or five cents per million.
Haiku 4.5 charged a flat $1.00 for input and $5.00 for output. So for requests under 100,000 tokens Haiku 5.5 is 90% cheaper, and 50% cheaper above. Anthropic says 90% of all Haiku 4.5 requests fell into the cheaper bracket.
There’s a catch in the footnotes. Haiku 5.5 uses the newer tokenizer from Sonnet 5.5 and Opus 5.5, so the same text counts as slightly more tokens than before. The 75% figure already accounts for that.
The jump over Haiku 4.5 is big
These numbers read like a generational change, not a point release:
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 | 1620 | 735 | 1437 | 1840 |
| OSWorld 2.1 | 72.4% | 15.7% | 48.9% | 83.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| Humanity’s Last Exam | 45.9% | 10.2% | — | 56.9% |
| Chartography | 46.4% | 6.4% | 29.1% | 61.6% |
Terminal-Bench is the outlier: Haiku 4.5 solved nothing there at all. On OSWorld, which measures driving a real computer, Haiku 5.5 lands closer to Sonnet 5.5 than to its own predecessor.
Early customers say much the same. Asana saw latency drop by more than 30% and inference up to 2.5x faster per agent turn. HubSpot hit 92.8% on a CRM test suite, which it calls the best score any model has posted there. Box reports 11 points above Haiku 4.5 at roughly half the latency.
What breaks when you migrate
Code written for Haiku 4.5 won’t necessarily keep working. Manual extended thinking via budget_tokens now returns a 400. Adaptive thinking is on by default, so a response can open with thinking blocks. And then there’s the tokenizer change.
On safeguards, Haiku 5.5 sits between generations: stricter than Haiku 4.5, looser than the recent large models. In cyber it permits a wider range of defensive work than Sonnet 5.5 does, but still blocks penetration testing. The biology restrictions match Sonnet 5, Sonnet 5.5 and Opus 5.
It’s live on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry.
Haiku 5.5 is not an Opus replacement
Anthropic says so itself, and the numbers back it: for complex agentic coding, Sonnet 5.5 and Opus 5.5 remain the better pick. 39.2% on Terminal-Bench against 70.6% is not a substitute.
Where Haiku 5.5 gets interesting is the work nobody was doing, because it cost too much. Compaction, summaries, classification, subagents sent into a document to find one number. At ten cents a million tokens you can run a model where you used to write a regex. That’s the real shift here — not the benchmark jump, but the threshold at which calling a model is worth it at all.