Claude Just Built the Longest Math Proof Ever—In 11…

Claude says its Claude AI just wrote the longest math proof ever made, and used it to formally prove Fermat’s Last Theorem. A problem that stumped mathematicians for 358 years. The machine finished in 11 days.

The proof runs 13 million lines of code—more than five times the size of Mathlib, the shared library mathematicians already use for this kind of work. That’s roughly equivalent to 160 novels of pure logical argument.

The work matters for a simple reason: checking that a major mathematical proof is correct can take years. Formalization—converting mathematical reasoning into a form computer proof assistants can verify—has been slow going. A 1908 German prize worth roughly $1 million to $2 million in today’s money, offered for the first valid proof of Fermat’s Last Theorem, drew 621 wrong submissions in its first year alone.

Mathematicians have been bad at policing this for a while.

The 358-Year Journey

Fermat’s Last Theorem says you can’t take three positive whole numbers, raise each one to a power higher than 2, and have the first two add up to the third. He scribbled that claim into the margin of a math book in 1637, adding that he had a truly marvelous proof that the margin was just too small to fit. Then he died.

Mathematicians spent the next 358 years trying to reconstruct whatever he thought he had.

The real proof didn’t show up until 1995, from British mathematician Andrew Wiles, and it came with a plot twist. Wiles announced his solution across three lectures in June 1993, only for a reviewer to find a hole in it later. He spent almost a year fixing it with a former student, Richard Taylor, nearly gave up, and finally published a corrected, 129-page proof in May 1995. It leaned on math that didn’t exist in Fermat’s lifetime, which is a big reason mathematicians now doubt Fermat’s own marvelous proof ever actually worked.

Kevin Buzzard, a mathematician at Imperial College London, kicked off a project in 2024 to translate Wiles’ proof into Lean, a language computers can check. It’s the kind of job that needs an army of volunteer mathematicians—the project’s own outline runs 86 pages, and its funding is locked in through 2029.

Claude finished the whole thing in 11 days.

A human-led project doing this exact same job has been running since 2024 and isn’t close to finished. Claude beat it to the finish line.

How Claude Actually Pulled It Off

Tianyi Peng, who builds AI formalization tools with a team at Columbia, decided to see how far Claude could get on its own. Dozens of Claude agents worked in parallel, writing definitions, proving small results, and stacking those into bigger ones, with almost no human input beyond the occasional nudge like “prioritize this theorem next.”

It didn’t go smoothly at first. Early on, the agents kept losing track of what they’d already proven and stopped collaborating. Those false starts still make up about 7% of the lines in the final proof.

What fixed it was a tool called Prove2Me, also built by Peng’s team, which gave every agent the same live to-do list of which smaller proofs still needed doing, so nobody duplicated work or wandered off. It also organized files so Lean could check everything faster, and kept plain-English notes on each result so agents could reuse each other’s work instead of reinventing the wheel.

By the time it was done, Claude had proven more than 30,000 supporting theorems and burned through billions of tokens, running on a research model Anthropic says is roughly comparable to Claude Sonnet 4, the version it later released to the public.

Why This Matters

Buzzard—whose own version of this project remains funded through 2029—reviewed Claude’s proof and gave it his blessing, saying it proves the theorem with no assumptions other than the axioms of mathematics.

This isn’t the same as Claude discovering brand-new math. Wiles already proved Fermat’s theorem three decades ago—Claude just built a machine-checkable receipt for it.

But that receipt matters. Mathematicians are increasingly swamped with unverified proofs, including AI-written ones, faster than humans can check them by hand. Formalized proofs eliminate the possibility of hidden errors slipping through.

That’s not a new problem. A computer-assisted proof of the Kepler conjecture took four years before a review panel would only commit to 99% certain. Grigori Perelman’s proof of the Poincaré conjecture took about as long to fully sink in.

If you don’t want to take Anthropic’s word for any of this, you don’t have to. The full 13-million-line proof is sitting on GitHub right now, free for any mathematician with enough free time to go pick apart, line by line.

Leave a Comment