If you run agents, most of your token spend probably doesn’t come from the smart model doing the hard reasoning. It comes from the cheap model doing the boring parts: summarizing, classifying, compacting context, running tool loops. That’s exactly the tier Anthropic is targeting with Claude Haiku 5.5, announced on October 7, 2026.
The headline numbers are straightforward. Anthropic’s announcement positions Haiku 5.5 as its cheapest, fastest small model, running on average around 75% less per task than Haiku 4.5. For anyone who already routes repetitive subagent work to a Haiku-class model, that’s a direct line-item cut — no architecture change required.
Where the model actually fits
Anthropic describes Haiku 5.5 as built for high-volume, cost-sensitive jobs: summaries, compaction, database queries, and classification. It also explicitly pairs with Opus 5.5 and Sonnet 5.5 as a subagent on coding work, and its speed makes it a candidate for latency-sensitive paths like live customer support and browser-driven agents.
That framing matters more than the benchmark table. The practical decision isn’t “is this model good?” — it’s “which of my agent’s steps can drop to this tier without the output quality dropping with them?” Classification and summarization are usually safe bets. Anything touching judgment calls or multi-step planning probably stays on the bigger model.
The effort setting changes the calculus
Haiku 5.5 is Anthropic’s first Haiku-class model with an adjustable effort setting, letting you trade intelligence for cost on the same model. Anthropic published accuracy-versus-cost curves on three benchmarks — OSWorld 2.1 (computer use), GDPval-AA v2.1 (professional knowledge work), and Humanity’s Last Exam (expert reasoning) — showing how performance shifts across settings.
For builders, this is a tuning knob rather than a fixed spec. If a workload sits near the boundary between tiers, dialing effort down might let the small model handle a task you’d otherwise route upward — or dialing up might let it cover something you’d previously split off. The tradeoff is now per-request instead of per-model.
A quick reality check on the benchmarks
The reported numbers are strong: on Anthropic’s published table, Haiku 5.5 scores 39.2% on Terminal-Bench 4.0 and 46.4% on FrontierCode 1.1, with large jumps over Haiku 4.5 on computer use and reasoning benchmarks. Treat these as vendor-reported results — Anthropic links its System Card for evaluation methodology, and your own evals are still the deciding vote. It’s the same discipline I described in the earlier piece on Anthropic splitting a frontier release into verified and safeguarded access: what a model does in a benchmark table and what it does in your pipeline are two different questions.
One piece of third-party signal: Aaron Vinh, a Staff Software Engineer at Asana, reported in early testing that Haiku 5.5 cut latency for task completions by over 30% and delivered up to 2.5x faster inference per agent turn in their AI Teammates evals.
The broader pricing shift
Two adjacent changes from the same announcement are easy to overlook. Sonnet 5.5 cache reads are being halved in price, which Anthropic says makes Sonnet about 20% cheaper on most agentic work — significant if your agents lean on long cached system prompts. And Claude Max and Team subscribers get a new monthly API credit aimed at agent and application building.
What I’d do with this
If you already have Haiku-class routing in place, the 75% cost drop makes the re-evaluation almost free: point your existing eval suite at Haiku 5.5, test both effort settings, and see how many steps can drop down a tier. The savings compound quietly in ways a flagship-model launch never will — and unlike most model upgrades, the risk is bounded to the cheap tasks anyway.
