Product1 publisher3 min readPublished
Kevin Buzzard holds a five-year grant to do the same job by hand. He compiled Anthropic's 13 million lines himself and says the result teaches mathematicians nothing and anyone running agent fleets quite a lot.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
The check ran on a 96-core machine, and it was not quick. Buzzard measured the repository at 13.4 million lines, more than five times the size of Mathlib, and reports that it takes nearly twenty times as long to compile [17]. Twenty over five is four, so each line of this proof costs roughly four times what a line of the community library costs to verify [1]. That is the standing bill for anyone who wants to build on the artifact rather than admire it. It is also not the artifact Buzzard's grant was buying. The proof follows the 1995 exposition of Wiles by Darmon, Diamond and Taylor, not the modern route his project has been formalising [15]. The most instructive part of Anthropic's write-up is the run that failed. Agents made early progress, then lost track of the project's state and stopped collaborating, and those dead runs still supplied around 7% of the non-boilerplate lines in the finished proof [11]. The repair was not a stronger model. It was Prove2Me, an open platform built by Anthropic researcher Tianyi Peng with collaborators at Columbia University, which maintains a graph of theorem statements so an agent can see what to attempt next, splits statements and proofs into separate files to speed compilation, and keeps a plain-language description of each statement so finished work can be found and reused [12]. Wrap that in a multi-agent harness on Claude Code and the same models closed the job inside a fortnight [13]. The human steering that survives in Anthropic's account is fragments: "Jacobian as a scheme sounds high priority", and "Push Mazur to be done soon" [14]. Teams often assume the bottleneck is model capability and plan to wait for the next release. Anthropic's account undercuts that: the weights were identical, the run failed and then succeeded, and a shared state artifact was the only thing that changed. Then the money, which is where the comparison gets loose. Buzzard's project has £1m over five years, an average of £200,000 a year [7][2]. Anthropic says the run consumed about six billion output tokens from an internal research model it calls roughly comparable to Claude Fable 5.1 [8], and at the $50 per million output tokens Fable 5.1 lists, six billion is $300,000 [9], or about 2.2 cents per line of Lean [3]. No invoice was raised, the model was internal, and a lab's own inference costs it less than list price, so this is the only public arithmetic rather than a price [10]. Buzzard wonders in print whether Anthropic spent more than he was given [7]. That is a pound grant measured against a dollar illustration, and everyone quoting him is making the same comparison. On soundness: Lean has had bugs found in it recently, and an agent determined enough could in principle exploit one to prove anything [18]. After his own inspection, and OpenAI models' review of the Lean codebase finding no soundness issues in the version used, Buzzard's summary is that a hack finishing the job is extremely unlikely [20]. Whether this transfers to your own long-horizon agent work turns on two things. The first is whether anything can reject bad output before a human reads it; Lean compiles or it does not, and a comparator confirmed the theorem proved matches the statement in Mathlib [16]. The second is whether there is a durable artifact the agents read to pick their next task and write back to when they finish. With both in place, the result is 13 million lines and a mathematician willing to compile them. With a checker but no shared state, the result is the first run: real output that never assembled into a proof, contributing that 7% salvage.
Ranked by verification strength, evidence, and original report placement.
Anthropic published the proof on Friday. Dozens of Claude agents wrote 13 million lines of Lean code.
The agents proved 30,300 intermediate theorems and used 29,500 of them in a complete, computer-checked proof of the conjecture Pierre de Fermat scribbled in a margin around 1637.
Kevin Buzzard of Imperial College London has led the community effort to formalise Fermat's Last Theorem since 2024, funded by the EPSRC.
Buzzard compiled Anthropic's code himself and ran the standard checking tool over it, then wrote it up on his blog under the headline "Anthropic has beaten me to it".
On the mathematics, Buzzard says the formalisation changes nothing: he puts the odds that Wiles's proof is correct at 99.9%, says most of the number theory community sits at 100%, and wrote that the formalisation "just faithfully follows the early literature on the proof and adds nothing".
Buzzard was given £1m to run his project over five years; Anthropic took eleven days, and he wonders whether they spent more.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Independent check on the object, unverified account of the process
Every quantity about the run comes from Anthropic's own post: the six billion tokens, the 7% salvaged from the collapsed attempt, the description of Prove2Me. What lifts this above a relayed announcement is that Buzzard compiled the 13.4 million lines himself, ran the standard checker, confirmed the theorem matches Mathlib's statement and then went looking for an axiom hack. The mathematical object has been verified by someone outside Anthropic; how it was produced rests entirely on the lab's telling, and no second newsroom in our coverage has tested it.
A verified proof still waiting on Mathlib
The proof compiles, uses only Lean's three standard axioms and matches the community library's statement of the theorem, and Prove2Me is public, so there is a real artefact and a real tool. Uptake is a separate question. Buzzard, himself a Mathlib maintainer, says the library will not currently accept AI review, reviewers are wary because most AI-generated submissions are poor, and the queue already holds around 3,000 open pull requests with more than 600 active. So far the only user outside Anthropic is the mathematician who compiled it in order to doubt it.
The softest numbers travel furthest
$300,000 is arithmetic on list price for a model nobody was billed for, and this reporting says so before it says the number; it will still be quoted as what the proof cost. The eleven-days-against-five-years framing has the same weakness, since Buzzard's grant also committed him to submitting foundational number theory to Mathlib and building a document that lets humans explore the modern proof, which he doubts Anthropic will attempt. Against that, the durable finding is underplayed almost everywhere: identical agents failed, then succeeded once a theorem graph told them what to attempt next.
Anthropic's demo, audited by the mathematician it undercuts
Candour about the collapsed first run also serves Anthropic: it makes the platform the lab released the hero of the story. The counterweight is unusually good. The man who compiled and audited the code holds the five-year grant this result undercuts, and he still called the engineering half the important one.
The proof holds up; the process account is unverified
A machine-checked proof either compiles or it does not, and this one was compiled by a hostile-interest party who also hunted for an axiom exploit, which is about as firm as a single-outlet story gets. The process narrative is weaker: the token volume, the 7% salvage and the two-sentence human input are all Anthropic's own account, unverified, and Buzzard's forecast that modern research will be formalised on the fly has nothing behind it yet but this one eleven-day run.
invest
Claude formalized Fermat's Last Theorem in 11 days against a human project funded through 20292 publishers
build
The $559M-versus-$12.3B quarter matters more than the $65B run rate4 publishers
product
Anthropic formalized Wiles' proof in 11 days by handing the review to a machine1 publisher
leadership
Anthropic put its Fermat proof's correctness check inside the default build target3 publishers
Publishers with included, body-backed reporting in this cluster.
1 article · September 6, 2026