Models & Research

Claude Sonnet 5.5: same price, 30% faster, five breaking changes

3 min read AI-generated

Artificial Analysis clocked roughly 193,000 output tokens per task at max effort. That's the heaviest token use they have ever measured, and about seven times what GPT-6 Astra needs.

Featured image for "Claude Sonnet 5.5: same price, 30% faster, five breaking changes"

Six days after Opus 5.5, Anthropic shipped Claude Sonnet 5.5 on Monday, the second model in the 5.5 family. The model ID is claude-sonnet-5-5, and it runs on the Claude API, Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Haiku 5.5 is due in the coming weeks.

Pricing is untouched: $2 per million input tokens, $10 per million output, 20 cents for cache reads. Exactly Sonnet 5’s numbers. Anthropic still claims up to 30% lower cost per task, on the grounds that the model needs fewer tokens to get there. Output also arrives 30% faster.

What the benchmarks say

On Terminal-Bench 4.0, an agentic command-line evaluation, Sonnet 5.5 jumps from 10.3% to 70.6% and lands above Opus 5.5 at 66.4%. On GDPval-AA, which covers real work across 44 occupations, it sits two Elo points behind Opus 5.5 – 1844 against 1846, up from 1449 for Sonnet 5. It is also the first Sonnet to beat Pokémon Red from screenshots alone.

Epic Games says it handled tens of thousands of lines of gameplay code and held up across multi-hour tasks. Balyasny ran 2,441 finance tasks through it: about 121,000 tokens per answer, where Sonnet 5 burned 497,000. Base44 measured 3.6 iterations per app build across 118 builds, against 7.7 for Opus 5.

Five things that break existing code

Anthropic lists them itself, and they are uncomfortably specific:

  1. thinking: {"type": "disabled"} now returns a 400. To switch off up-front thinking you send between_tools, and that only works at effort high or below.
  2. Forced tool use is gone. tool_choice set to any or tool returns a 400.
  3. Thinking blocks are bound to a model and a conversation. Sonnet 5.5 reads blocks from Sonnet 5, Opus 4.8 and Haiku 4.5, but none from Opus 5, Opus 5.5, Fable or Mythos. And no other model reads Sonnet 5.5’s blocks at all.
  4. On the Claude API and Google Cloud, computer_20251124 is rejected in favour of computer_toolset_20260801. On Bedrock the old tool still works.
  5. Opus 4.8, Opus 4.7 and Sonnet 5 are refused as advisors.

One more change fails no request but makes interfaces look dead: text between tool calls now comes back inside thinking blocks. If you stream that text to users, they see nothing until you set the display value.

The cost advantage depends on the effort dial

Artificial Analysis ran their own numbers and put it second on their Intelligence Index, behind Opus 5.5. They confirm the Terminal-Bench 4.0 jump: 64% against 14% for Sonnet 5.

The catch is in the same report. At max effort, Sonnet 5.5 spends roughly 193,000 output tokens per task – 60% more than Opus 5.5 at max, and around seven times GPT-6 Astra. That puts it off the intelligence-versus-cost frontier rather than on it. Anthropic’s 30% applies to everyday work, not to the top setting.

The distillation guard is the real change

Item three on that breaking-change list reads like housekeeping and isn’t. Sonnet 5.5’s thinking blocks only work in the account that produced them, or one linked to it. Send them from a different account and the API quietly drops them. The request succeeds, just without the reasoning.

Anthropic states the reason plainly: distillation through thousands of fake accounts. Sonnet 5.5 is the first Sonnet to ship with classifiers against reasoning extraction. That’s an admission about how easily capability leaks out of outputs – and the first time Anthropic’s answer to it makes life slightly worse for honest developers. Switch accounts mid-session in Claude Code and you will notice.

Sources

AnthropicClaudeModelsAPI