On September 4, 2026, Anthropic announced an AI mathematics milestone: a swarm of Claude agents formalized Fermat’s Last Theorem in Lean, complete and machine-checked, in 11 days. A conjecture posed in 1637 and proven by Andrew Wiles only in 1995 across a 129-page paper is now 13 million-plus lines of code that depends on nothing but Lean’s three standard axioms. The proof is on GitHub, and anyone can run the check themselves.
Eleven Days, Six Billion Output Tokens
Fermat’s Last Theorem says no three positive integers satisfy aⁿ+bⁿ=cⁿ when n is greater than 2. Anthropic’s approach was not one model writing the answer in one shot — it was dozens of Claude agents working in parallel. Underneath sits an internal research model roughly comparable to Claude Fable 5.1, which emitted about six billion output tokens over the 11 days and produced more than 13 million lines of Lean — over five times the size of Lean’s math library, Mathlib. The project proved 30,300 theorems in total, of which 29,500 feed into the final proof. The route follows the Darmon–Diamond–Taylor 1995 simplification of Wiles’s argument rather than reinventing the mathematics. The scale has real costs: Kevin Buzzard reported that compiling the result took roughly 20 times as long as compiling Mathlib, on a machine with 96 cores and 500 GB of RAM.
Humans Only Set the Direction
The project was initiated by Anthropic researcher Tianyi Peng, whose Columbia University group builds AI formalization tools. They designed Prove2Me, an open collaborative platform that maintains a DAG of theorem dependencies, speeds up Lean compilation, and lets agents search and reuse existing results. On top of it runs a multi-agent harness built on Claude Code. Human input was kept to a minimum: Peng occasionally issued high-level directions like “Jacobian as a scheme sounds high priority.” Even failed attempts left residue — work from early agent runs contributed roughly 7% of the non-boilerplate lines in the final proof. As a smaller demonstration of accessibility, agents on three personal Claude Max plans formalized Vinogradov’s three primes theorem in three days.
Buzzard: It Checks Out, But Tells Mathematics “Essentially Nothing”
The verifier was Kevin Buzzard of Imperial College London, who has led the community’s FLT formalization effort since 2024 on a five-year, £1 million EPSRC grant. He compiled the proof himself and ran a comparator confirming the statement matches Mathlib’s own FLT declaration. His verdict: it checks out, proving the theorem “with no assumptions other than the axioms of mathematics,” with techniques that could help “rooting out errors in the current mathematical corpus and lightening the load of referees.” But he also poured cold water on the mathematical content — the formalization “just faithfully follows the early literature on the proof and adds nothing,” so it tells mathematics “essentially nothing.” To rule out the AI exploiting soundness bugs in Lean’s checker, he manually read all of the roughly 100 lines of non-mathematical code — a convenience tactic, nothing more. A side effect: Freek Wiedijk’s list of 100 formalization challenges is now fully closed. The human moment is hard to resist — Buzzard was at a festival in Wales with poor phone reception when the announcement arrived, dismissed it as a crank email, and only caught up a week later. Claude’s own log captured the other side of the moment — “The FLT root reads Proved on the site. Historic moment (modulo re-check).”
What It Means for AI Engineering
Three takeaways. First, long-horizon agent work passes another reality check: 11 days, dozens of agents, tens of millions of lines — the enablers were not a bigger model but environment engineering: a compiler that gives an unambiguous success signal, a shared task graph in Prove2Me, and reuse of failed attempts. Second, formal verification is becoming the next battleground for AI code quality: when “it compiles” means “it is mathematically correct,” review shifts from reading code line by line to running a proof checker. Third, the economics are starting to work: community estimates put six billion output tokens in the low hundreds of thousands of dollars at API prices — against Buzzard’s five-year academic grant. As he wrote, if thousands of pages of literature can now be formalized end-to-end by an AI swarm in 11 days, formalization of modern research will soon happen on the fly — and it will keep the field honest about results that are merely “known to the experts.”
Sources
- Formalizing Fermat’s Last Theorem — Anthropic
- anthropics/fermats-last-theorem — GitHub
- FLT: Anthropic has beaten me to it — Xena Project
AI-assisted summary compiled from the sources above, reviewed by a human before publishing.
