Science1 distinct publisher3 min readPublished
Anthropic says its Claude agents ran autonomously for 11 days to formalise Fermat's Last Theorem in 13 million lines of Lean, a result that rests on a finished human proof and 2 million lines of prior formalisation.
The Scientist · Science desk

product
Anthropic formalized Wiles' proof in 11 days by handing the review to a machine1 distinct publisher
product
Thomson Reuters spent $40M to make a $450K training run worth doing2 distinct publishers
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
build
Perf work stopped being a specialist queue item, and slow endpoints became a choice1 distinct publisher
Compiled by The ScientistSomething wrong?How this is made
Divide 13 million lines of Lean by 11 days and the run emitted about 1.18 million lines a day, close to 14 lines every second, without a pause [16]. Set the same file against the 100 pages of Wiles and Taylor's argument and each page of human prose expands to roughly 130,000 lines of code [17]. The 29,500 intermediate results average about 440 lines each [18]. None of those ratios measures mathematical difficulty. They measure how much a natural-language proof leaves implicit, and how cheap writing out the implicit part has become.
What the 11 days does not include is the substrate. The clock starts with an argument that was already finished and already repaired: Andrew Wiles worked on the problem for seven years in secret, announced in 1993, and then spent about a year with Richard Taylor fixing a flaw found in the proof [7]. It also starts with Mathlib, the central repository, already holding 2 million lines of formalised mathematics available to build on [5]. The agents were translating toward a target known to be true, with a specification in hand.
The run broke on coordination, not mathematics: several times the agents lost track of the project's state and stopped collaborating effectively, and progress came after Anthropic began using Prove2Me, a tool designed for human mathematical collaboration, to help agents track their work and pick next tasks [12]. Human experts also stepped in occasionally with high-level instructions to keep the work on track [11]. At this scale the binding constraint on a swarm of formalisers looks like coordination rather than proving.
Correctness is a different matter, and the mechanism is doing the work: a formalised proof sits in code that a machine grinds through step by step, which is what exposes a broken link [14]. Kevin Buzzard of Imperial College London drew the wider inference in his statement, saying the work shows autoformalisation of algebra, harmonic analysis, geometry and number theory, that such artefacts are now robust enough to be built upon, and that automatic formalisation of the modern mathematical literature is a big step nearer [10]. His is the only outside assessment in New Scientist's account, and it reached the public through Anthropic's own announcement [19].
So the demonstrated capability is narrower than the elapsed time suggests and still worth taking seriously: converting a settled, corrected human proof into machine-checkable form is now a days-long compute job, given a library beneath it and handlers watching. Buzzard's claim that these artefacts can be built upon is testable in the ordinary way, and the test is specific: the next formalisation project either imports the 13 million lines or re-derives them.
Ranked by verification strength, evidence, and original report placement.
Anthropic has created a formalised proof of Fermat's Last Theorem; a group of AI agents completed the task in 11 days, confirming that the human-found proof proposed in the 1990s is correct.
Anthropic said in a blog post that its Claude model worked continuously and autonomously for 11 days to write the proof, with numerous separate AI agents involved and different tasks, such as smaller chunks of the theorem, assigned to each.
Anthropic's formalisation runs to 13 million lines of Lean and covers roughly 29,500 intermediate theorems that were necessary stepping stones to the overall work.
The formalisation is over five times the current size of all previous work on Mathlib, which also makes it the largest Lean proof ever written.
There are already 2 million lines of formalised mathematics stored in a central repository called Mathlib.
Fermat's Last Theorem states that there are no whole numbers a, b and c satisfying a^n + b^n = c^n where n is a whole number greater than 2; it was proven in 1995 by Andrew Wiles.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 5, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One relay of a company blog post
The 11 days and the 13 million lines both originate in Anthropic's own announcement and reach readers through a single publisher. Lean code is checkable in a way most AI claims are not, which is what keeps this from scoring lower, but no independent check appears anywhere in the reporting: even Buzzard's endorsement arrived in a statement Anthropic published. The derived rates hold up arithmetically, which tests internal consistency rather than the underlying figures.
One run at one lab
What exists is a single completed job and one internal tooling decision: Anthropic reaching for Prove2Me, built for human mathematicians, when its agents stopped coordinating. The one downstream effect reported is subtractive, in that Buzzard's five-year Imperial project has been overtaken. Buzzard's remark about artefacts being robust enough to build upon reads as an invitation to future users, and nothing here shows the Lean output in third-party hands.
Confirmation presented as solution
"In short, the problem is solved" sits oddly on a result that verifies a proof Wiles completed in 1995 and that leans on 2 million lines of Mathlib written by other people. The 11-day headline also absorbs two things the same reporting supplies: human experts steering with high-level instructions, and agents that several times lost the project's state until an outside coordination tool was slotted in. The underlying achievement is substantial, so the overstatement is in the framing rather than the figures.
Claim and endorsement share a channel
Anthropic announced a capability result about its own model on its own blog, and the expert verdict that vouches for it was distributed in a statement Anthropic published. The specifics that would let anyone price or reproduce the run, compute, cost, and how many agents actually ran, are the specifics left out. That pattern of what is loud and what is quiet is consistent enough to read as marketing structure, built around mathematics whose validity is not in question.
Specific numbers, single origin
The reporting is precise, self-consistent and unusually candid about where the agents failed. That candor raises trust in the account as an account, though it still traces to one publisher relaying one company's blog post, a narrow base for numbers of this size. Until someone outside Anthropic compiles or reviews the Lean, the assessment cannot firm up much beyond this.