The interesting part of Anthropic’s Claude Sonnet 5.5 announcement is not the benchmark table. It’s the pricing setup: the token rates are identical to Sonnet 5 — $2 per million input tokens, $10 per million output tokens, $0.20 per million cache reads — but Anthropic reports it needs far fewer tokens to finish the same work, coming out up to 30% cheaper per task and 30%+ faster at the same time.
That’s a different kind of upgrade than a raw capability jump. It changes your routing math, not just your leaderboards.
The capability jump, briefly
The headline numbers are large. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 versus Sonnet 5’s 10.3%, and on GDPval-AA — a test of real-world work across occupations — it lands within two points of Opus 5.5 (1844 vs 1846). It’s also the first Sonnet model to beat Pokémon Red from screenshots alone, which Anthropic offers as evidence on long-horizon image understanding.
Anthropic is careful to draw the line clearly: on several evaluations Sonnet 5.5 at Max effort performs comparably to Opus 5.5, but in their own and external testers’ experience, Opus 5.5 stays clearly stronger on complex, open-ended work requiring sustained judgment. Sonnet 5.5 is positioned for well-scoped everyday tasks — bug fixes, polished documents, slides, spreadsheets — while Opus 5.5 keeps the heavy judgment calls.
Cost per task is the number that matters
The announcement’s most useful framing is score versus cost per task at each effort level. Two data points worth noting:
- At High effort on FrontierCode, Sonnet 5.5 scores 10 points above Sonnet 5 at the same setting, at about one fifteenth the cost per task.
- At Medium effort on Terminal-Bench, it far exceeds Sonnet 5’s best score for less than a tenth of the cost per task.
Early testers cited by Anthropic noticed a mechanism behind this: Sonnet 5.5 batches tool calls together more than Sonnet 5 did, producing fewer steps and lower costs per run. If your agent loops through tools one at a time, this efficiency shows up directly in your bill.
How this reshapes tiering
If you route between Sonnet and Opus today, the decision boundary moves. Anthropic suggests Sonnet 5.5 complements Opus 5.5 best at lower effort settings, where it costs less per task; at higher settings it can perform comparably at similar cost — meaning the premium for Opus is mostly justified on genuinely open-ended problems, not mid-complexity ones.
Epic Games’ COO Daniel Vogel said in early testing the model cleared a quality bar you’d expect from a higher tier, holding up on a system design audit and handling tens of thousands of lines of gameplay architecture code with less prescriptive prompting.
One caution: these cost-per-task figures come from Anthropic’s own benchmark runs. Your workload’s mix of input length, tool calls, and cache hit rate will move the number. If you already track per-task cost, rerun your own comparison before re-routing production traffic.
What’s still coming
Claude Haiku 5.5, aimed at high-volume cost-sensitive applications, arrives in the coming weeks — which matters if your stack already leans on cheap models. I ran the subagent cost math on the previous Haiku release, and the same question applies here: when the small model gets this much better, how much work actually still needs the expensive tier? With Sonnet 5.5, the honest answer is less than it was last month.
