AGENTIC COMMONSAI industry briefings

繁中EN

TAGAnthropicPUBLISHED 2026-10-07

All English articlesAnthropic

Running the Subagent Math on Anthropic's Latest Small Model

On this page6 sections
  1. Where the model actually fits
  2. The effort setting changes the calculus
  3. A quick reality check on the benchmarks
  4. The broader pricing shift
  5. What I'd do with this
  6. Sources

If you run agents, most of your token spend probably doesn’t come from the smart model doing the hard reasoning. It comes from the cheap model doing the boring parts: summarizing, classifying, compacting context, running tool loops. That’s exactly the tier Anthropic is targeting with Claude Haiku 5.5, announced on October 7, 2026.

The headline numbers are straightforward. Anthropic’s announcement positions Haiku 5.5 as its cheapest, fastest small model, running on average around 75% less per task than Haiku 4.5. For anyone who already routes repetitive subagent work to a Haiku-class model, that’s a direct line-item cut — no architecture change required.

Where the model actually fits

Anthropic describes Haiku 5.5 as built for high-volume, cost-sensitive jobs: summaries, compaction, database queries, and classification. It also explicitly pairs with Opus 5.5 and Sonnet 5.5 as a subagent on coding work, and its speed makes it a candidate for latency-sensitive paths like live customer support and browser-driven agents.

That framing matters more than the benchmark table. The practical decision isn’t “is this model good?” — it’s “which of my agent’s steps can drop to this tier without the output quality dropping with them?” Classification and summarization are usually safe bets. Anything touching judgment calls or multi-step planning probably stays on the bigger model.

The effort setting changes the calculus

Haiku 5.5 is Anthropic’s first Haiku-class model with an adjustable effort setting, letting you trade intelligence for cost on the same model. Anthropic published accuracy-versus-cost curves on three benchmarks — OSWorld 2.1 (computer use), GDPval-AA v2.1 (professional knowledge work), and Humanity’s Last Exam (expert reasoning) — showing how performance shifts across settings.

For builders, this is a tuning knob rather than a fixed spec. If a workload sits near the boundary between tiers, dialing effort down might let the small model handle a task you’d otherwise route upward — or dialing up might let it cover something you’d previously split off. The tradeoff is now per-request instead of per-model.

A quick reality check on the benchmarks

The reported numbers are strong: on Anthropic’s published table, Haiku 5.5 scores 39.2% on Terminal-Bench 4.0 and 46.4% on FrontierCode 1.1, with large jumps over Haiku 4.5 on computer use and reasoning benchmarks. Treat these as vendor-reported results — Anthropic links its System Card for evaluation methodology, and your own evals are still the deciding vote. It’s the same discipline I described in the earlier piece on Anthropic splitting a frontier release into verified and safeguarded access: what a model does in a benchmark table and what it does in your pipeline are two different questions.

One piece of third-party signal: Aaron Vinh, a Staff Software Engineer at Asana, reported in early testing that Haiku 5.5 cut latency for task completions by over 30% and delivered up to 2.5x faster inference per agent turn in their AI Teammates evals.

The broader pricing shift

Two adjacent changes from the same announcement are easy to overlook. Sonnet 5.5 cache reads are being halved in price, which Anthropic says makes Sonnet about 20% cheaper on most agentic work — significant if your agents lean on long cached system prompts. And Claude Max and Team subscribers get a new monthly API credit aimed at agent and application building.

What I’d do with this

If you already have Haiku-class routing in place, the 75% cost drop makes the re-evaluation almost free: point your existing eval suite at Haiku 5.5, test both effort settings, and see how many steps can drop down a tier. The savings compound quietly in ways a flagship-model launch never will — and unlike most model upgrades, the risk is bounded to the cheap tasks anyway.

Sources

AGENTIC COMMONSOperated by PHLEGON LABS
SHAREXEMAIL
Support us

Related reading

  1. Sonnet 5.5's Real Story Is the Cost Curve, Not the Benchmarks

    Claude Sonnet 5.5 cuts typical token costs about 30% and pairs tiered pricing with caching and batching discounts.

    Anthropic

  2. One Model, Two Doors: Anthropic Splits a Frontier Release into Verified and Safeguarded Access

    Claude Mythos 5.1 ships only through vetted-access programs, while Claude Fable 5.1 offers the same model with lighter-touch safeguards.

    Claude Mythos

  3. Claude Opus 4.8: Same Price, 41 Days Later, Hundreds of Subagents

    Anthropic ships Claude Opus 4.8 41 days after 4.7 at unchanged pricing, adding Dynamic Workflows with hundreds of subagents, effort control, and sharper uncertainty flagging.

    Anthropic