Anthropic published a 13-million-line Lean 4 proof of Fermat's Last Theorem that dozens of Claude agents wrote in 11 days. For teams weighing agents on correctness-critical code, the kernel checks every step, so what humans still review is the statement being proved.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+20
- Incentives55
- Confidence60
An interview with Lean's founder says an AI translated C zlib into Lean and proved compress-then-decompress round-trips in about a week. The remaining work is optimization without breaking the proof.
Reality
- Evidence25
- Adoption35
- Hype gap+40
- Incentives65
- Confidence30
Anthropic's own account has dozens of agents writing 13 million machine-checked lines in 11 days, about 7% of them dead ends, and only after a Columbia team built the shared to-do list that stopped them duplicating work.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence58
Kevin Buzzard holds a five-year grant to do the same job by hand. He compiled Anthropic's 13 million lines himself and says the result teaches mathematicians nothing and anyone running agent fleets quite a lot.
Reality
- Evidence75
- Adoption10
- Hype gap+20
- Incentives60
- Confidence70
Anthropic says its Claude agents ran autonomously for 11 days to formalise Fermat's Last Theorem in 13 million lines of Lean, a result that rests on a finished human proof and 2 million lines of prior formalisation.
Reality
- Evidence58
- Adoption32
- Hype gap+18
- Incentives70
- Confidence64
The build fails unless the theorem rests on exactly Lean's three standard axioms, and an independent kernel written in Rust re-checked more than a million declarations. The eleven-day timeline rests on Anthropic's own account.
Perspective Coverage
3 publishers
- Builder
- Builder 47%
- Operator
- Operator 17%
- Investor
- Investor 36%
Reality
- Evidence76
- Adoption22
- Hype gap+14
- Incentives72
- Confidence70
The Math-AI project turns plain-language mathematics into Lean 4 theorems, then files whatever compiles into a reusable library. The architecture is the claim, not the benchmark.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+14
- Incentives36
- Confidence44