Ramana Kumar used AI to exploit a different bug in each of two Lean kernels, passing off a false Collatz disproof as machine-checked in July. AI labs rely on Lean to vouch for their maths results, so those claims hold only as well as the checker does.
Reality
- Evidence50
- Adoption25
- Hype gap+15
- Incentives55
- Confidence55
OpenAI says about 10,000 AI agents produced a proposed finite-time singularity for the forced 3D Navier-Stokes equations in 88 hours. The construction fits one route the Clay rules allow and leaves unforced smoothness open, while the mathematicians whose forced-Euler work came first ask whether their Codex drafts reached the model.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence40
Vitalik Buterin says Ethereum's Hegotá upgrade, planned for 2027, is likely the last fork that someone who knew Ethereum in 2015 would recognize. The work he expects to follow has no dates yet, so ETH holders are pricing a long rebuild of the protocol off a single milestone.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives35
- Confidence55
Anthropic published a 13-million-line Lean 4 proof of Fermat's Last Theorem that dozens of Claude agents wrote in 11 days. For teams weighing agents on correctness-critical code, the kernel checks every step, so what humans still review is the statement being proved.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+20
- Incentives55
- Confidence60
An interview with Lean's founder says an AI translated C zlib into Lean and proved compress-then-decompress round-trips in about a week. The remaining work is optimization without breaking the proof.
Reality
- Evidence25
- Adoption35
- Hype gap+40
- Incentives65
- Confidence30
Anthropic's own account has dozens of agents writing 13 million machine-checked lines in 11 days, about 7% of them dead ends, and only after a Columbia team built the shared to-do list that stopped them duplicating work.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence58
Kevin Buzzard holds a five-year grant to do the same job by hand. He compiled Anthropic's 13 million lines himself and says the result teaches mathematicians nothing and anyone running agent fleets quite a lot.
Reality
- Evidence75
- Adoption10
- Hype gap+20
- Incentives60
- Confidence70
Trail of Bits' agent-built Lean model of the Miden VM, backed by 95 machine-checked proofs, surfaced a way for a malicious prover to forge Falcon signatures. A dev.to account says the agents built tooling first, so that Miden's stack assembly could be reviewed at all.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives55
- Confidence40
A proof checker confirmed the steps of OpenAI's Navier-Stokes result within days. The argument now running through mathematics is over who explains the 166 pages and who gets the credit for them.
Perspective Coverage
7 publishers
- Builder
- Builder 46%
- Operator
- Operator 33%
- Investor
- Investor 21%
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+40
- Incentives70
- Confidence58
OpenAI's write-up gives the token counts, the agent count and the verification time for its Navier-Stokes result. The verification time is the number that decides whether the method transfers to anyone else's workload.
Perspective Coverage
3 publishers
- Builder
- Builder 52%
- Operator
- Operator 33%
- Investor
- Investor 15%
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence55
With Navier-Stokes claimed and five Millennium problems left, the ordering of what falls next turns on which conjectures a single counterexample would settle. Five million dollars of Clay prize money is still unclaimed.
Perspective Coverage
3 publishers
- Builder
- Builder 37%
- Operator
- Operator 40%
- Investor
- Investor 23%
Reality
- Evidence55
- Adoption35
- Hype gap+35
- Incentives70
- Confidence50
The company says an internal model cleared more than 100 open problems after a month of training, and the advisory group it is backing at the Institute for Advanced Study has been handed the release schedule to work on.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+55
- Incentives70
- Confidence55
OpenAI says its new nine-member mathematics advisory group will not advise on the pace of its internal research, and the group's own website says the decisions rest with the AI companies.
Reality
- Evidence45
- Adoption40
- Hype gap+35
- Incentives70
- Confidence45
Ten Claude Opus 5.5 agents produced the algorithm and a machine-checked proof of its runtime bound in about 15 hours. The theorem covers sparse directed graphs, and the margin over Dijkstra grows as the twelfth root of log n.
Reality
- Evidence52
- Adoption10
- Hype gap+20
- Incentives72
- Confidence55
The model's answers wobble by a few hundredths on identical input, so the checkable properties live in a TLA+ spec and a Rust quorum rule that counts only votes whose margin clears a measured noise floor.
Reality
- Evidence55
- Adoption12
- Hype gap+12
- Incentives40
- Confidence45
A preprint from Ariel University reports laundering attack success of up to 68% against existing agent-memory defenses, and argues in machine-checked TLA+ that authority has to be bound to an item's origin at the moment it is written.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+35
- Incentives62
- Confidence47
Six months before the review of Miden's zero-knowledge VM began, Trail of Bits set its agents to writing developer tooling for the project's custom assembly language, then reviewed the code with it.
Reality
- Evidence46
- Adoption24
- Hype gap+12
- Incentives78
- Confidence55
The Institute for Responsible Superintelligence launches with cryptographers Shafi Goldwasser and Vinod Vaikuntanathan and a cofounder who left OpenAI, and its founding post concedes that deep learning itself has resisted theoretical analysis.
Reality
- Evidence38
- Adoption12
- Hype gap+18
- Incentives68
- Confidence45
The Clay Mathematics Institute has the claimed Navier-Stokes solution under review, and a summer of agent-found bugs in Lean's kernel shows how narrow the guarantee a Lean certificate actually gives.
Reality
- Evidence62
- Adoption60
- Hype gap+30
- Incentives68
- Confidence58
A LessWrong experiment had Claude Code verify the July counterexample to the Jacobian conjecture, then claimed the map had a typo. On byte-identical input the older checkpoint argued back and the newer one dropped it in all four runs.
Reality
- Evidence62
- Adoption15
- Hype gap+22
- Incentives28
- Confidence56