Anthropic has released Claude Opus 5.5, the first model in its Claude 5.5 family. The headline claim is unusual for a frontier launch: it performs at roughly the level of Claude Fable 5.1 on most work, but costs 40% less to run than Opus 5 on typical workloads. For anyone budgeting agentic coding work, that’s the part worth reading twice.
The pricing actually changed shape
The per-token cuts are modest: $4 input and $20 output per million tokens, about 20% below Opus 5. The real drop is in cache reads, which Anthropic prices at $0.20 per million tokens — 60% less than Opus 5 — and which the company notes make up the majority of agentic and coding costs. Output is also more than 30% faster. Add a fast mode in Claude Code and the Claude Platform at up to 2.5x speed ($8/$40 per million tokens) and increased five-hour limits on Pro, Max, Team, and seat-based Enterprise plans.
If you’ve been running the cost math on agent fleets — like I did earlier when looking at what Sonnet 5.5 changed for model tiers — the cache-read number is the lever that matters. Long-running agents that re-read context every turn are exactly where a 60% cut compounds.
Where it wins: long, sprawling jobs
Anthropic’s own examples lean heavily on large-codebase work. One early tester audited and fixed a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and 2.5x the tokens. In an internal test translating HAProxy from C to Rust, Opus 5.5 passed nearly all of HAProxy’s regression tests, finished in 9.5 hours versus 12 for Fable 5.1, and cost 51% less.
Benchmark numbers point the same direction. On Terminal-Bench 4.0, Opus 5.5 at default effort matches GPT-6 Astra at roughly 40% of the cost; on FrontierCode it beats Astra’s top score at about a fifth of the cost per task; on CursorBench it outscores GPT-5.6 Sol by 11 points at about a third of the cost. GitHub’s testing reported among the fewest tokens and steps measured, solving more terminal tasks than Opus 5 in less than half the steps.
One honest caveat from Anthropic itself: at this capability level, benchmark margins have become a less reliable guide, and in its own use the gap between Opus 5.5 and Fable 5.1 is narrower than the scores suggest. Quote the ranges, not the point estimates, when you write your internal comparison doc.
Safety changes that affect how you build
Opus 5.5 scored best of any model on Anthropic’s automated behavioral audit, is described as less likely to take hard-to-reverse actions or act outside its boundaries, and is more resistant to prompt injection than Opus 5. It’s deployed with safeguards similar to Fable 5.1’s, based on comparable biology and cybersecurity capability. That last point matters if you route sensitive tasks: safeguard interventions count as failures in some third-party runs — Zapier’s AutomationBench results were reported without fallback models — so raw scores can understate what you’d see in practice.
What I’d do next
The obvious move is to re-run your own per-task cost comparison between Opus 5 and Opus 5.5 on a real workload, paying attention to tokens per task rather than tokens per call — Anthropic attributes the full 40% saving to the combination of cheaper tokens and fewer of them. Claude Sonnet 5.5 and Haiku 5.5 are coming in the following weeks, so if your workload sits below Opus tier, it may be worth waiting a few weeks before locking in a new default.
