Product1 distinct publisher3 min readPublished
The formalization runs to 13 million lines of Lean, about 100,000 for every page of Wiles, and nobody is expected to read it, because a program grades every step and that is what makes the agent volume usable.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
Nobody will read 13 million lines of Lean, and the project does not require it. Divide that by the 129 pages of Wiles' original and you get about 100,800 lines of Lean per printed page [14]. Divide it by the 11 days and you get roughly 1.18 million lines a day [15]. Human review cannot run at that rate. It also does not matter here, because a formalized proof is checked automatically by computer, which is the reason mathematicians want proofs in Lean in the first place [9].
The phrase "AI does mathematics" hides the shape of the run. Those 29,500 intermediate theorems work out to about 441 lines of surviving Lean each [17], at roughly 460 tokens of model output per line that survived [16]. The decomposition is the mechanism itself. Lean is unforgiving in a specific way: arguments build on each other, so one bad line can invalidate everything after it [10]. Cutting the job into 29,500 separately graded pieces is how a failure stays local instead of fatal.
Step selection is the part a team can copy. Prove2Me is open source, it helps agents pick the next move in a long workflow, and Anthropic says it also lowers inference costs [8]. Picking the next move against something that will grade the attempt transfers to any place you already own an oracle: a compiler, a test suite, a schema validator, a reconciliation that has to balance. Compare that with the two other recent results in this vein. Anthropic reported new information about the Riemann zeta function a month earlier [12], and OpenAI used Astra last month to solve several Erdos problems [13]. Those are discovery claims that people have to judge. Formalization is the case where the grader is a program, and that is why the line count is allowed to get silly.
On the evidence: the 11 days, the token count and the theorem count come from Anthropic's own blog post as relayed by SiliconANGLE [1][4], with no independent reproduction reported [19]. Kevin Buzzard, whose work Claude drew on, says autoformalization artefacts are now robust enough to be built upon [11], which is a judgment about usability rather than a replication. There is no compute bill in the account and no count of human interventions behind the phrase "limited high-level input", and those are the two numbers you would need to price a run of your own.
So the tooling question comes down to whether a checker exists for the output and whether a human still has to sign the result, not to how capable the agent is. Draw the 2x2: one axis is whether an automatic checker exists for the output, the other is whether a human still has to sign the result. Checker, no signature required: more agents is close to free, and 13 million lines is a perfectly good answer. Checker plus a human signature: every extra token is review you have bought, and volume works against you. No checker at all: the output lands on your team as a tax, and the demo is quietly asking you to be the oracle. Anthropic's number is impressive because it sits in the first box. Most software teams live in the second, and the useful work is dragging specific tasks out of the third box into the first, one verifier at a time.
Ranked by verification strength, evidence, and original report placement.
Anthropic PBC said it used Claude to create a computer-verifiable version of the proof of Fermat's Last Theorem, detailing the project in a blog post published the day of the report.
Andrew Wiles developed the proof of Fermat's Last Theorem in 1995; it runs for 129 pages and took months of work to verify.
Anthropic's formalization comprises 13 million lines of Lean code, which SiliconANGLE reports makes it the largest-ever file of its kind.
Mathematicians expected formalizing Wiles' proof to take several years; according to Anthropic, its researchers completed the task in 11 days using an internal research model.
The model completed the task using only a limited amount of high-level input from humans, spinning up several dozen agents that generated 6 billion tokens of output.
The agents proved no fewer than 29,500 intermediate theorems while working on the proof.
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · September 4, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
leadership
Anthropic put its Fermat proof's correctness check inside the default build target2 distinct publishers
invest
OpenAI rates GPT-6 Astra capable of hacking hardened systems without human guidance1 distinct publisher
invest
Claude formalized Fermat's Last Theorem in 11 days against a human project funded through 20291 distinct publisher
invest
OpenAI ships a model it grades critical on its own cybersecurity threshold1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One outlet, one company blog post
Our record holds two items, but they trace back to one story and a single upstream document. Every figure in the piece — 11 days, six billion tokens, 29,500 lemmas — comes from Anthropic. What lifts this above bare assertion is the medium: a Lean file either compiles or it does not, so the claim is checkable in principle by anyone given the code. Nobody in this reporting is described as having done that, and the largest-file-ever ranking is asserted without a comparison set.
One artifact, one named user
Uptake, as distinct from announcement, comes to a single file and a single outside name. Kevin Buzzard says autoformalization output is now solid enough to build on, and his own work fed the run, which is the nearest thing here to third-party use. Prove2Me being open source means the tooling is at least available to others. No downstream project, library merge, or second team touching the 13 million lines appears anywhere in the reporting.
Superlatives ahead of the check
The overstatement sits in the framing rather than the arithmetic. 'Several years' compressed into 11 days is Anthropic's own comparison, the largest-file claim goes unsourced, and the model's capability tier is placed relative to two other models with no evaluation attached. Pushed the other way, one thing here is underplayed: the strict Lean checker is the reason dozens of agents and six billion tokens yield something usable, and only Buzzard's quote comes near saying so.
Vendor-scored result in a two-lab race
Anthropic is publishing a capability result about Anthropic, a month after its own Riemann zeta post and in the same window as OpenAI's Erdős claims, so the scoreboard is kept by the players on it. The write-up appeared the day of the blog post with no outside check sought, and SiliconANGLE's own solicitation of community and marketplace support closes the page. The figures may still be accurate, but nobody in the chain had reason to slow down and test them.
Coherent but uncorroborated
The account hangs together technically and the numbers are internally consistent, which is why this is not lower. But two identical copies of one story being consistent with each other just reflects duplication, not independent corroboration. A compile log from anyone outside Anthropic, or a second mathematician confirming the file builds, would move this a long way in either direction.