Invest1 distinct publisher3 min readPublished
Anthropic's own account has dozens of agents writing 13 million machine-checked lines in 11 days, about 7% of them dead ends, and only after a Columbia team built the shared to-do list that stopped them duplicating work.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
Thirteen million lines in eleven days works out to about 1.18 million lines a day, close to 49,000 an hour, sustained, with Lean rather than a person checking each one [1][2][13]. Divide those lines by the more than 30,000 supporting theorems Anthropic says the agents proved and the average result runs about 433 lines [10][15], which gives the shape of the job: thousands of small results stacked, rather than one long argument. Checking is the expensive half of mathematics, since finding a single broken link inside a hundred pages of dense argument can cost other mathematicians years [21], and a 1908 German prize for a valid proof of this same theorem drew 621 wrong submissions in its first year [20].
The disclosure's clearest number is also its least flattering. Roughly 7% of the lines in the finished proof are false starts from the early phase, when the agents kept losing track of what they had already proven and stopped collaborating [8], which is about 910,000 lines of dead work carried into the artifact [14]. Coordination broke down first, and Prove2Me fixed it: a tool from Tianyi Peng's Columbia team that handed every agent the same live to-do list, reorganized files so Lean could check them faster, and kept plain-English notes so one agent could reuse another's result [9][7]. The eleven days therefore sit on top of a human-built shared-state layer; the scarce input in this run was that coordination layer itself, not raw model access.
Which is where the capital-allocation reading gets thin. Anthropic's post is dated September 4, 2026 and puts completion in the previous month [11]; it quantifies consumption as billions of tokens and no dollars, on a research model it says is roughly comparable to the Claude Fable 5.1 version it later shipped [10]. The account does not say what share of Kevin Buzzard's 86-page outline the proof covers, nor whether it drew on the Imperial project's existing Lean library [19][4]. Buzzard's funding is committed through 2029 [4], and eleven days is about 0.6% of that window [17], a striking ratio measured in time, not dollars.
My read, with the counter-thesis beside it: the transferable capability is throughput on work where a checker returns a verdict at every step, and a verifier like Lean is the rarer asset here, not the agent. The counter is that Prove2Me is a to-do list and a file organizer rather than a new model [9], so if that scaffolding ports cheaply, any domain with a compiler or a test suite inherits the same curve. A third path runs through the review itself: Buzzard is the one named mathematician who has confirmed the proof reduces to math's most basic logical rules [3], and Wiles's 1993 announcement survived three lectures before a reviewer found the hole in it [5].
Note also the gap between the framing and the event. Decrypt's headline calls this a 350-year-old problem solved by AI [12] while its own text counts 358 years from Fermat's 1637 margin note [6] and credits Wiles with the proof, corrected with Richard Taylor and published across 129 pages in May 1995 [5]; what Claude did was translate that argument into a language a machine can check [1]. For scale on that translation, the formal version carries roughly 100,000 lines per published page of Wiles [16]. Until someone publishes the token bill, eleven days is a quote with the price torn off.
Ranked by verification strength, evidence, and original report placement.
Anthropic says its Claude AI produced the first fully computer-checked (formalized) proof of Fermat's Last Theorem in 11 days, largely on its own.
Anthropic's post announcing the result is dated September 4, 2026 and says that last month Claude completed the first formalized proof of Fermat's Last Theorem.
The proof runs to 13 million lines of code that a computer can check line by line, described as the longest math proof ever built.
Kevin Buzzard, the Imperial College London mathematician leading the human formalization project, reviewed Claude's proof and confirmed it holds up using nothing but math's most basic logical rules.
Buzzard began a project in 2024 to translate Wiles's proof into Lean; the project's outline runs 86 pages, its funding is locked in through 2029, and it is not close to finished.
Andrew Wiles announced his solution across three lectures in June 1993, a reviewer later found a hole in it, and he published a corrected 129-page proof in May 1995 after nearly a year of work with former student Richard Taylor.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 5, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
Eleven days of machine time turned Wiles's 100 pages into 13 million lines of Lean1 distinct publisher
product
Thomson Reuters spent $40M to make a $450K training run worth doing2 distinct publishers
product
Anthropic formalized Wiles' proof in 11 days by handing the review to a machine1 distinct publisher
invest
OpenAI rates GPT-6 Astra capable of hacking hardened systems without human guidance1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One lab's account, one public file
The 11 days, the 13 million lines, the 30,000-plus supporting theorems and the 7% dead ends all trace to Anthropic's own write-up as Decrypt relays it. Two things lift it above a vendor blog: Buzzard, who runs the rival human project, is quoted saying the proof needs nothing beyond the axioms, and the Lean file is public, where a checker rather than a referee decides. That check has not been run within this coverage.
Shipped and public, barely read
The artifact exists, is public, and the tooling that produced it ran at scale across dozens of agents. That is the extent of it: one named mathematician has looked at the proof, a second reviewer has not turned up, and neither Mathlib, the Imperial project, nor any other group has taken the 13 million lines up.
Headline a size class above the body
Decrypt's title puts a 350-year-old problem in the solved column; a few paragraphs down the same piece explains that Wiles solved it in 1995 and Claude wrote the machine-checkable receipt, and elsewhere counts 358 years rather than 350. The underlying result stands up, even though the framing around it borrows from a different, larger achievement than the one actually completed here.
Subject, sponsor and only source
Anthropic is the actor, the publisher of the result and the origin of every figure, and the post lands beside the public release of the model it says the research build resembles. Peng's Columbia team gets a showcase for Prove2Me out of the same run. Buzzard is the one voice with reason to push back, since his own project holds funding through 2029, and he signed off regardless.
Direction firm, magnitudes unaudited
Something large, machine-checked and public exists, with a credible reviewer's sign-off attached, so the direction is reasonably safe to trust, though the magnitudes rest on one interested account. The two questions that would size the work properly, how much of the 86-page blueprint the run covered and how much came from existing Lean libraries, are never put.