On September 1, 2026, Anthropic released Claude Fable 5.1 and Claude Mythos 5.1. The most interesting thing about this launch is the structure, not the scores: Anthropic states plainly that “Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards.” Fable 5.1 is generally available; Mythos 5.1 ships only through trusted access programs, with safeguards tailored for cybersecurity and life-science research.
One model, two shipping configurations — a strategy that echoes how OpenAI would tier Astra through Daybreak days later: the frontier capability is fixed; what changes is who gets it and under which guardrails.
Benchmark Numbers: Against Fable 5, Opus 5, and GPT-5.6 Sol
Anthropic calls Fable 5.1 “the world’s most advanced model for coding and knowledge work.” The comparison numbers (Fable 5.1 / Fable 5 / Opus 5 / GPT-5.6 Sol):
- Terminal-Bench 4.0: 55.8% (Mythos 5.1: 60.9%) / 42.0% / 52.3% / 37.3%
- Terminal-Bench-Science 0.1: 52.6% / 24.7% / 29.0% / 22.4%
- Humanity’s Last Exam: 60.9% without tools, 65.0% with tools
- GDPval-AA v2: 1853 / 1723 / 1824 / 1711
- OSWorld 2.0: 77.9% partial, 41.7% strict
- AutomationBench: 31.4% / 17.1% / 26.9% / 19.6%
- CursorBench 3.2.0: 73.4% / 70.5% / 70.0% / 67.2%
Mythos outscores Fable by 5 points on Terminal-Bench. Anthropic’s explanation: the gap reflects where the older, less precise cyber safeguards intervened, and with this safeguard update the gap is expected to shrink.
Pricing: List Prices Unchanged, Bills Go Down
Input stays at $10 per million tokens and output at $50, but two changes cut real costs: cache reads drop to $0.25 per million (75% cheaper), typical workloads run about 25% cheaper than Fable 5, and highly agentic workloads up to roughly 45% cheaper. Effort defaults to High in Claude Code and Medium in Claude Cowork and Claude.ai. Every’s CEO Dan Shipper put it as: “It’s friendly Fable. Fable-level intelligence, Opus-level price, Sonnet-speed.”
Guardrail Tuning: Fewer False Positives, Not a Blanket Loosening
The most product-relevant change is safeguard precision. On the cyber side, false-positive interventions drop by about 60% per Claude Code session; Fable 5.1 “can now be used to discover software vulnerabilities — though not to develop exploits for them,” while dual-use tasks such as penetration testing still redirect to Opus models. On the biology side, safeguards fire 85% less often on benign elementary biology and medical queries, though life-science R&D still routes to Opus unless you go through Mythos.
The safety disclosures are worth keeping on record: Mythos 5.1 shows greater bio capability than Mythos 5 yet stays below the next Responsible Scaling Policy risk tier, and its cyber capabilities — Anthropic’s strongest yet — sit within a lower Frontier Compliance Framework risk category. A new anti-distillation measure blocks new API accounts from editing prior context while preserving Claude’s thinking transcript.
Mythos 5.1 and Trusted Access: Who Gets It, and For What
Mythos 5.1 is currently limited to a set of US organizations, with expansion coordinated with the US government. Two channels exist: the Cyber Verification Program (CVP, reduced cyber safeguards for defensive security work) and the Life Sciences Verification Program (LSVP, developed in partnership with the US government, with the first participants enrolled). Claude Security’s codebase vulnerability scanning is now powered by Mythos 5.1. Enterprises also get Enterprise Frontier Safeguards (EFS), storing data on customer-controlled cloud infrastructure in a phased rollout starting fall 2026.
Three research cases illustrate what trusted access produces: Mythos 5.1 achieved binding affinities 10x higher than the best Adaptyv Bio competition entries on three protein targets (EGFR, Nipah G, 15-PGDH), with a near-50% hit rate across 12 targets against a typical 10–15%; Fable 5.1 rebuilt a Venus elevation map from Magellan radar data at 2–3 km resolution (previously 10–20 km), released under Creative Commons; and Mythos 5.1 sped up seven open-source genomics and protein models up to 2.5x on an H100, cutting estimated genome-wide GPU costs 30–60%.
Sources
AI-assisted summary compiled from the sources above, reviewed by a human before publishing.
