An AI Swarm Proved Fermat’s Last Theorem in 11 Days. A Mathematician Checked It.

A five-year government grant. Dozens of AI agents. Thirteen million lines of computer code. Eleven days. That is what it took to formalize Fermat’s Last Theorem—not to prove it, but to verify it in a way no human reviewer could dispute.

Anthropic published the completed proof on Friday. The agents, using Claude as a swarm rather than a single assistant, produced 30,300 intermediate theorems and stitched them into a full, machine-checked proof of a conjecture Pierre de Fermat scribbled in a margin around 1637.

Kevin Buzzard of Imperial College London has led a separate, community-driven effort to formalize the same theorem since 2024. His project is funded by a £1m EPSRC grant spread across five years. When he heard about Anthropic’s result, he compiled their code, ran the standard checking tool, and wrote up his findings on his blog under the headline “Anthropic has beaten me to it.”

His verdict splits into two parts, and most coverage has focused on only the first.

On the mathematics itself, it changes nothing. Buzzard estimates the odds that Andrew Wiles‘s original 1995 proof is correct at 99.9%, with most of the number theory community sitting at certainty. The formalization, he noted, “just faithfully follows the early literature on the proof and adds nothing.” It is verification, not discovery.

On what it demonstrates, he draws a very different conclusion. An AI swarm just moved thousands of pages of mathematical literature from paper to verified code end to end in less than a fortnight. If that is now possible, he wrote, formalization of modern research will begin happening in parallel with publication rather than years after.

He added one observation no press release would carry. He was given £1m over five years. Anthropic took eleven days, and he wonders whether they spent more.

Cost Arithmetic

The arithmetic on that wondering is imprecise but instructive. Anthropic estimates the project consumed roughly six billion output tokens, produced by an internal research model it describes as roughly comparable to Fable 5.1. At Fable 5.1’s list price of $50 per million output tokens, six billion tokens runs to $300,000. That is an illustration, not a real invoice—the model was internal, and a company’s own inference costs less than list price. But it is the only public arithmetic available, and it sits against a multi-year government grant.

The first attempt failed. Anthropic is unusually candid about this, and the candor is the most useful part of the post. The agents made early progress, then lost track of the project’s state and stopped coordinating. Those failed runs still contributed roughly 7% of the non-boilerplate lines in the final proof.

What fixed it was not a better model. It was Prove2Me, an open platform built by Anthropic researcher Tianyi Peng with collaborators at Columbia University. The platform maintains a graph of theorem statements, so agents can see what to attempt next. It splits statements and proofs into separate files to speed compilation. And it keeps a plain-language description of each statement, so work can be found and reused.

With that coordination layer in place—and a multi-agent harness built on Claude Code—same models, same weights, same capability. The only variable was scaffolding. The job finished in under a fortnight.

Technical Verification

The proof Anthropic produced follows the 1995 exposition of the Wiles argument by Darmon, Diamond and Taylor. That is not the modern route Buzzard has been formalizing in Mathlib, the community’s mathematics library. It uses only Lean‘s three standard axioms, and a comparator confirmed the theorem it proves matches the statement in Mathlib. At 13 million lines, it is more than five times the size of Mathlib itself. Buzzard measured it at 13.4 million lines and says it takes nearly twenty times as long to compile on a 96-core machine.

He went looking for cheating. Lean has had soundness bugs discovered in recent years, and a sufficiently determined agent could in principle exploit one to prove anything. Buzzard checked. He asked an agent to flag every line of the repository that was not a definition or a proof. Roughly 100 lines came back, and he inspected them. They defined a convenience tactic. OpenAI’s models also reviewed Lean’s codebase and found no soundness issues in the version used. Buzzard’s own summary: a hack finishing the job is extremely unlikely. The code is plainly developing the mathematics the proof requires.

None of those 13 million lines can enter Mathlib as things stand. Buzzard, a Mathlib maintainer, says the library does not currently accept AI reviews. Human reviewers are also reluctant to tackle AI-generated code, because most AI-generated submissions are poor quality. The repository already has around 3,000 open pull requests, more than 600 active in the review queue.

The bottleneck has moved rather than vanished. The constraint shifted from writing proofs to reading them—which is precisely the problem Anthropic says formalization solves, arriving one layer up.

Buzzard still has a job. His grant committed him to submitting foundational number theory to Mathlib and to building a document that lets humans explore the modern proof. He doubts Anthropic will do the second.

Verification vs. Discovery

This is distinct from earlier AI maths claims. Timothy Groughs argued in August that the breakthroughs everyone quotes are counterexamples, not proofs. A complete machine-checked proof is a different object—and it is the thing that argument said had not yet happened. It is also not new mathematics, which is the distinction Buzzard is pressing.

OpenAI said in August that Astra had solved ten open problems. Claude found flaws in two cryptographic algorithms in July that expert review had missed. Those were claims about discovery. This is a claim about verification, and verification is the easier half.

One more number from the same post: three researchers on personal Claude subscriptions formalized Vinogradov’s Three Primes Theorem in three days.

Buzzard heard about the FLT result late. He was at a music festival with poor reception. An email arrived from a stranger, titled “End-to-end Lean formalization of Fermat’s Last Theorem,” and he initially wrote the sender off as a crank.

Leave a Comment