Ramana Kumar used AI to exploit a different bug in each of two Lean kernels, passing off a false Collatz disproof as machine-checked in July. AI labs rely on Lean to vouch for their maths results, so those claims hold only as well as the checker does.
Reality
- Evidence50
- Adoption25
- Hype gap+15
- Incentives55
- Confidence55
OpenAI says about 10,000 AI agents produced a proposed finite-time singularity for the forced 3D Navier-Stokes equations in 88 hours. The construction fits one route the Clay rules allow and leaves unforced smoothness open, while the mathematicians whose forced-Euler work came first ask whether their Codex drafts reached the model.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence40
An interview with Lean's founder says an AI translated C zlib into Lean and proved compress-then-decompress round-trips in about a week. The remaining work is optimization without breaking the proof.
Reality
- Evidence25
- Adoption35
- Hype gap+40
- Incentives65
- Confidence30
Anthropic's own account has dozens of agents writing 13 million machine-checked lines in 11 days, about 7% of them dead ends, and only after a Columbia team built the shared to-do list that stopped them duplicating work.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence58
Kevin Buzzard holds a five-year grant to do the same job by hand. He compiled Anthropic's 13 million lines himself and says the result teaches mathematicians nothing and anyone running agent fleets quite a lot.
Reality
- Evidence75
- Adoption10
- Hype gap+20
- Incentives60
- Confidence70
Trail of Bits' agent-built Lean model of the Miden VM, backed by 95 machine-checked proofs, surfaced a way for a malicious prover to forge Falcon signatures. A dev.to account says the agents built tooling first, so that Miden's stack assembly could be reviewed at all.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives55
- Confidence40
A proof checker confirmed the steps of OpenAI's Navier-Stokes result within days. The argument now running through mathematics is over who explains the 166 pages and who gets the credit for them.
Perspective Coverage
7 publishers
- Builder
- Builder 46%
- Operator
- Operator 33%
- Investor
- Investor 21%
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+40
- Incentives70
- Confidence58
OpenAI says no specific user data was accessed to solve the problem, and that it cannot rule out that de-identified data derived from two mathematicians' use of its products improved its models. A procurement team has to read both sentences.
Perspective Coverage
5 publishers
- Builder
- Builder 36%
- Operator
- Operator 33%
- Investor
- Investor 31%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+45
- Incentives70
- Confidence60
OpenAI's write-up gives the token counts, the agent count and the verification time for its Navier-Stokes result. The verification time is the number that decides whether the method transfers to anyone else's workload.
Perspective Coverage
3 publishers
- Builder
- Builder 52%
- Operator
- Operator 33%
- Investor
- Investor 15%
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence55
With Navier-Stokes claimed and five Millennium problems left, the ordering of what falls next turns on which conjectures a single counterexample would settle. Five million dollars of Clay prize money is still unclaimed.
Perspective Coverage
3 publishers
- Builder
- Builder 37%
- Operator
- Operator 40%
- Investor
- Investor 23%
Reality
- Evidence55
- Adoption35
- Hype gap+35
- Incentives70
- Confidence50
OpenAI says its new nine-member mathematics advisory group will not advise on the pace of its internal research, and the group's own website says the decisions rest with the AI companies.
Reality
- Evidence45
- Adoption40
- Hype gap+35
- Incentives70
- Confidence45
Ten Claude Opus 5.5 agents produced the algorithm and a machine-checked proof of its runtime bound in about 15 hours. The theorem covers sparse directed graphs, and the margin over Dijkstra grows as the twelfth root of log n.
Reality
- Evidence52
- Adoption10
- Hype gap+20
- Incentives72
- Confidence55
Six months before the review of Miden's zero-knowledge VM began, Trail of Bits set its agents to writing developer tooling for the project's custom assembly language, then reviewed the code with it.
Reality
- Evidence46
- Adoption24
- Hype gap+12
- Incentives78
- Confidence55
A Caltech open letter says research mathematicians end up verifying and correcting what AI companies announce, without pay or credit. OpenAI's sponsorship was $1m in credits, and it took them back.
Reality
- Evidence62
- Adoption58
- Hype gap+14
- Incentives74
- Confidence60
Nvidia's Nemotron scored 30 of 42 at IMO 2026 writing proofs in ordinary prose, and the September 9 paper publishes the data, the code and the acceptance rule that decided which candidate proofs survived.
Reality
- Evidence36
- Adoption22
- Hype gap+12
- Incentives62
- Confidence40
Tristan Buckmaster says OpenAI accelerated onto Navier-Stokes after learning of his work with Anthropic's Levent Alpoge, and that the answer he got about his prompt history addressed lookup rather than training.
Reality
- Evidence40
- Adoption30
- Hype gap+45
- Incentives88
- Confidence50
The blue check marks in the editor gutter assume an honest author, and they stay blue when a dependency contains sorry. That is why the manual escalates to axiom listings and to re-checking .olean files.
Publishers:lean-lang.org
Reality
- Evidence80
- Adoption
- Insufficient
- Hype gap−10
- Incentives35
- Confidence70
The Navier-Stokes result cost more in compute than the Clay Institute's award, and the interesting line in OpenAI's blog post is the one conceding it cannot rule out that customers' de-identified prompts helped.
Reality
- Evidence32
- Adoption14
- Hype gap+34
- Incentives74
- Confidence44
OpenAI says the result resolves statement C of the official Millennium formulation, then says it will not claim the $1M prize. The smooth force applied to its fluid throughout is why that second sentence carries the weight.
Reality
- Evidence58
- Adoption20
- Hype gap+35
- Incentives72
- Confidence52
OpenAI says roughly 10,000 agents produced a Navier-Stokes singularity proof in about 88 hours, but the only proofs an outsider can run belong to two mathematicians who posted Lean formalisations with their preprints.
Reality
- Evidence42
- Adoption22
- Hype gap+45
- Incentives78
- Confidence52
Earlier coverage
- Eleven days of machine time turned Wiles's 100 pages into 13 million lines of Lean
Science · September 5, 2026 · 2 publishers
- Anthropic put its Fermat proof's correctness check inside the default build target
Leadership · September 5, 2026 · 3 publishers
- Requiring a reproduction path stops a review agent from quadrupling the feature
Build · September 6, 2026 · 1 publisher
- A $28 agent run swapped BCD for base-2^64 limbs and built its own oracle
Build · September 5, 2026 · 1 publisher
- Constructing the undo before the write admits 13.8% of write-capable MCP tools
Build · August 29, 2026 · 1 publisher
- Palomar registers Lean proofs against a commit, and refuses to referee them
Build · August 18, 2026 · 1 publisher
- Lanyon AI raises $10.6M to make proofs, not plausibility, the acceptance test for AI code
Build · August 17, 2026 · 1 publisher