Claude formalizes Fermat's Last Theorem in Lean 4
TECH

Claude formalizes Fermat's Last Theorem in Lean 4

42+
Signals

Strategic Overview

  • 01.
    Anthropic's Claude agents produced the first complete, end-to-end, computer-checked formalization of Fermat's Last Theorem in Lean 4, working largely autonomously over 11 days on top of Andrew Wiles' original proof - the result spans roughly 13 million lines of Lean code and about 30,300 intermediate theorems, more than 5x the size of Mathlib, Lean's main math library.
  • 02.
    The effort ran through Prove2Me, a collaborative platform designed by Tianyi Peng and collaborators at Columbia University that coordinates multiple Claude agents via a directed acyclic graph of theorem statements and dependencies, letting dozens of agents parallelize without duplicating work; Anthropic's internal research model generated about 6 billion output tokens with only limited high-level human guidance.
  • 03.
    As a secondary demonstration, the same approach formalized Vinogradov's Three Primes Theorem in about 3 days using only consumer-tier Claude Max plans, with no special internal model required.
  • 04.
    The proof underwent multiple independent verification passes - a from-scratch Lean 4.33.1 kernel compile across 60,475 modules and a separate nanoda (Rust) kernel re-check of over 1 million declarations - both returning zero errors, though about 7% of non-boilerplate lines are failed multi-agent attempts still retained in the build.

Deep Analysis

Formalization, Not New Mathematics: What Claude Actually Did

It is worth being precise about what happened here. Claude's agents did not discover a new proof of Fermat's Last Theorem - they translated Andrew Wiles' original 1993/1995 proof (completed with Richard Taylor) into a form that a computer can mechanically verify, checked against only Lean's three standard axioms. [1]That distinction matters because the hard mathematical insight was already accepted by the field three decades ago; what Claude's agents contributed was the enormous, tedious labor of encoding every logical step so a proof assistant can check it line by line. Even Tianyi Peng, who designed the coordination platform that made the run possible, hedged his confidence accordingly: "I'm 99% sure, but it's hard to be 100% certain about a proof this long." [1]

That hedge lines up with how the broader mathematics community appears to be reading the result: reaction ranged from astonishment at the sheer speed of the run to pointed scrutiny of how much of the achievement rests on human-authored Lean library structure Claude was able to reuse, with the dominant framing being formalization throughput rather than mathematical discovery. [2]That framing is reinforced by a 2025 arXiv paper that had only managed to formalize Fermat's Last Theorem in Lean for the narrower special case of regular primes - a reminder of how much scaffolding already existed for Claude's agents to build on rather than invent from scratch. [3]

Trust But Verify: How a 13-Million-Line Proof Gets Checked

A proof this size cannot be trusted just because Claude's agents say it compiles. Anthropic's team ran a from-scratch Lean 4.33.1 kernel compile across all 60,475 modules of the final build, and then had a second, structurally independent kernel - nanoda, written in Rust - re-check more than 1 million declarations from the ground up; both passes returned zero errors. [4]Running two differently-implemented kernels against the same claim is the standard defense against a bug in any single kernel silently rubber-stamping a false statement, and it is notable that roughly 7% of the non-boilerplate lines in the final build are not part of the proof at all but failed multi-agent attempts still retained in the build. [4]

That double-kernel design is exactly what came up when the result hit X and Reddit. Commenters drew a direct line to a well-known prior incident in which a bug in Lean's trusted kernel had been exploited to construct a fake formal "proof" of the Collatz conjecture, and the consensus in the discussion was that this FLT result is not vulnerable to the same class of exploit, precisely because it does not rely on a single kernel's say-so - the from-scratch recompile plus the independent nanoda cross-check would have to both fail identically for a bad theorem to slip through. The Lean project's own official account added a striking framing of its own, calling the result "the largest Lean proof ever constructed" - a claim that lines up with the roughly 13 million lines and 30,300 intermediate theorems Anthropic reported, more than five times the size of Mathlib itself.

Buzzard Scooped: What This Means for Human-Led Formalization and Mathlib

The most human part of this story belongs to Kevin Buzzard. The Imperial College mathematician has spent years running an EPSRC-funded, community-staffed effort to formalize Wiles' proof in Lean by hand, a project funded through 2029 and built around a route through Khare-Wintenberger and Kisin that Richard Taylor himself helped plan. [5]Reddit discussion pointed to Buzzard's own blog post on the news, reportedly carrying a title to the effect of him having been "scooped" - a mathematician who had committed years to painstakingly formalizing this exact theorem watching an AI system finish the job in eleven days, in the middle of his own still-running effort. It is a genuinely awkward moment for a field that has organized long-term human labor around exactly the kind of task Claude's agents just compressed.

The irony runs deeper than timing. Fermat's Last Theorem was, until this week, the last unformalized entry on Freek Wiedijk's well-known list of 100 landmark theorems that the formalization community uses as its benchmark of progress - the very list Buzzard's project was working toward closing out by hand. [6]Its closure by an AI system rather than the human project built to close it now forces an uncomfortable question for Lean's maintainers: whether to merge the roughly 29,500 new theorems Claude's agents produced, a body of work five times larger than Mathlib itself, into the shared library that the rest of the ecosystem builds on, and if so, how to review code nobody wrote by hand at a scale no human review process was designed for. [2]

The DAG Breakthrough: Why Multi-Agent Math Suddenly Works

What made eleven days possible was not a smarter model so much as a better way to split the work. Prove2Me, the coordination platform Tianyi Peng and collaborators built at Columbia, represents the proof as a directed acyclic graph of theorem statements and their dependencies, letting dozens of Claude agents parallelize across the graph without duplicating work or losing track of shared state. [2][4]Anthropic's internal research model generated roughly 6 billion output tokens against that structure with only limited high-level human guidance. [7]The same approach also formalized Vinogradov's Three Primes Theorem in about three days using nothing more than consumer-tier Claude Max plans - no special internal model required, just the same coordination trick applied to a smaller graph. [1]

That scaling story has a cost dimension worth being honest about. Commenters estimated the FLT run would cost roughly $300,000 at Anthropic's public list API pricing, though the real internal cost to Anthropic was likely far lower, in the rough range of $30,000 to $60,000. Either number is small next to a multi-year human formalization effort. Anthropic's push here follows its own earlier work applying Claude to Riemann zeta function research, and OpenAI has separately been running its own effort on Erdos problems - rival labs converging on formal mathematics as a proving ground for what frontier models can do unsupervised. [7][8]

Historical Context

1637
Stated the conjecture in the margin of a book, without providing a proof.
1908
A 100,000 German gold marks prize for a proof was established; it drew 621 incorrect submissions in its first year alone.
1993
Announced a proof of Fermat's Last Theorem that contained a critical gap.
1995
Published the corrected, 129-page proof of Fermat's Last Theorem after about a year of fixing the gap.
2024
Launched a multi-year, EPSRC-funded (EP/Y022904/1, through 2029) community effort to formalize Wiles' proof in Lean, following a route planned by Richard Taylor building on Khare-Wintenberger and Kisin.
2025
Published a formalization of Fermat's Last Theorem in Lean for the narrower special case of regular primes only.
2026-09-04
Announced Claude's complete formalization of Fermat's Last Theorem, produced via Prove2Me in 11 days, ahead of Buzzard's still-ongoing multi-year human-led project.

Power Map

Key Players
Subject

Claude formalizes Fermat's Last Theorem in Lean 4

AN

Anthropic

Developer and publisher of the formalization effort, provided the internal research model that drove the main run

TI

Tianyi Peng / Columbia University

Designed the Prove2Me DAG-based coordination platform used to run the agents

KE

Kevin Buzzard / Imperial College London

Leads a separate, pre-existing multi-year human-led Lean formalization of FLT, EPSRC-funded through 2029; reviewed and publicly validated Claude's proof

LE

Lean / Mathlib maintainers

Face an open question over whether to merge the roughly 29,500 new theorems, a body of work 5x larger than Mathlib itself, into the shared library

OP

OpenAI

Cited as a rival lab in the broader AI-for-math race, e.g. work on Erdos problems

Fact Check

8 cited
  1. [1] Formalizing Fermat's Last Theorem
  2. [2] Claude formalized Fermat's Last Theorem in 11 days
  3. [3] Formalizing Fermat's Last Theorem for regular primes
  4. [4] Claude formalizes Fermat's Last Theorem: 11 days of autonomous work
  5. [5] ImperialCollegeLondon/FLT
  6. [6] FLT use case - Lean Lang
  7. [7] Anthropic uses Claude to formalize proof of Fermat's Last Theorem
  8. [8] Riemann zeta research - Anthropic

Source Articles

Top 5

THE SIGNAL.

Analysts

Buzzard said the achievement - which Anthropic researchers say took only 11 days - "proves Fermat's Last Theorem with no assumptions other than the axioms of mathematics," and argued that "if the automatic formalization of FLT is possible now, then we have taken a big step towards automatic formalization of the modern mathematical literature." He added: "We see autoformalization of algebra, harmonic analysis, geometry and number theory, and we learn that AI autoformalization artefacts are now robust enough to be built upon."

Kevin Buzzard, Mathematician, Imperial College London
Validates the proof's soundness and calls it a landmark for automated formalization, while flagging what it means for the ecosystem

Peng said, "I'm 99% sure, but it's hard to be 100% certain about a proof this long," reflecting the reality that even independently kernel-checked proofs of this size invite continued scrutiny.

Tianyi Peng, Prove2Me designer, Columbia University
Confident but appropriately cautious given the proof's scale

Reaction ranged from astonishment at the speed of the run to scrutiny of how much human-authored Lean library structure Claude relied on, with the consensus framing being that this demonstrates formalization throughput rather than a new mathematical discovery.

Mathematician community reaction (per AI Weekly reporting)
Impressed by speed but careful to frame this as formalization throughput rather than new mathematics
The Crowd

@AnthropicAI has shared the first end-to-end, computer-checked proof of Fermat's Last Theorem: 13 million lines of Lean, 29,500 intermediate theorems. Their announcement calls it "the largest Lean proof ever constructed." See also Kevin Buzzard's blog post about the proof: https://t.co/RVGvA22wwY

@@leanprover1615

Respect where respect is due: Anthropic says Claude completed the first fully computer-checked proof of Fermat’s Last Theorem in 11 days. Andrew Wiles proved the theorem in 1995. Claude’s achievement was turning an existing proof into a form where a computer can check every

@@kimmonismus739

⚡️ INTERESTING: Anthropic says Claude produced the first complete computer-checked proof of Fermat’s Last Theorem in Lean after working largely autonomously for 11 days.

@@Cointelegraph300

Buzzard's formalization of Fermat's Last Theorem completed

@u/New-Committee-4052266
Broadcast
Claude AIがフェルマーの最終定理を11日で検算

Claude AIがフェルマーの最終定理を11日で検算

Fermat's Last Theorem Proof: What Claude Really Proved in Lean in 11 Days

Fermat's Last Theorem Proof: What Claude Really Proved in Lean in 11 Days

Claude's Formalizing Fermat's Last Theorem

Claude's Formalizing Fermat's Last Theorem