Meta's six Muse Spark math papers, out October 2, label which passages the model drafted and which the researchers wrote. They show a chat assistant writing search code and proof drafts for mathematicians who chose the problems.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives70
- Confidence45
Ramana Kumar used AI to exploit a different bug in each of two Lean kernels, passing off a false Collatz disproof as machine-checked in July. AI labs rely on Lean to vouch for their maths results, so those claims hold only as well as the checker does.
Reality
- Evidence50
- Adoption25
- Hype gap+15
- Incentives55
- Confidence55
Youness Lamzouri has independently confirmed Claude's result that at least 67.25% of Riemann zeta zeros lie on the critical line, against 40% known since 1989. The hypothesis itself is still open. The case shows a working split for AI in mathematics, in which a model finds a candidate result and a specialist rebuilds it.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence60
OpenAI says about 10,000 AI agents produced a proposed finite-time singularity for the forced 3D Navier-Stokes equations in 88 hours. The construction fits one route the Clay rules allow and leaves unforced smoothness open, while the mathematicians whose forced-Euler work came first ask whether their Codex drafts reached the model.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence40
Anthropic released an AI-generated proof of a percolation conjecture that Fields medallist Hugo Duminil-Copin had tried and failed to prove. What it shows about AI research depends on expert checking and on how much of the argument humans built first.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+30
- Incentives
- Insufficient
- Confidence40
Anthropic's own account has dozens of agents writing 13 million machine-checked lines in 11 days, about 7% of them dead ends, and only after a Columbia team built the shared to-do list that stopped them duplicating work.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence58
Kevin Buzzard holds a five-year grant to do the same job by hand. He compiled Anthropic's 13 million lines himself and says the result teaches mathematicians nothing and anyone running agent fleets quite a lot.
Reality
- Evidence75
- Adoption10
- Hype gap+20
- Incentives60
- Confidence70
Scientific American's account from the mathematicians' congress in Philadelphia has a Fields medalist quitting for OpenAI and a Millennium Prize problem falling to an opaque mix of models, with no verification trail reported for either.
Reality
- Evidence30
- Adoption35
- Hype gap+45
- Incentives65
- Confidence35
A proof checker confirmed the steps of OpenAI's Navier-Stokes result within days. The argument now running through mathematics is over who explains the 166 pages and who gets the credit for them.
Perspective Coverage
7 publishers
- Builder
- Builder 46%
- Operator
- Operator 33%
- Investor
- Investor 21%
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+40
- Incentives70
- Confidence58
OpenAI says no specific user data was accessed to solve the problem, and that it cannot rule out that de-identified data derived from two mathematicians' use of its products improved its models. A procurement team has to read both sentences.
Perspective Coverage
5 publishers
- Builder
- Builder 36%
- Operator
- Operator 33%
- Investor
- Investor 31%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+45
- Incentives70
- Confidence60
OpenAI's write-up gives the token counts, the agent count and the verification time for its Navier-Stokes result. The verification time is the number that decides whether the method transfers to anyone else's workload.
Perspective Coverage
3 publishers
- Builder
- Builder 52%
- Operator
- Operator 33%
- Investor
- Investor 15%
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence55
With Navier-Stokes claimed and five Millennium problems left, the ordering of what falls next turns on which conjectures a single counterexample would settle. Five million dollars of Clay prize money is still unclaimed.
Perspective Coverage
3 publishers
- Builder
- Builder 37%
- Operator
- Operator 40%
- Investor
- Investor 23%
Reality
- Evidence55
- Adoption35
- Hype gap+35
- Incentives70
- Confidence50
The company says an internal model cleared more than 100 open problems after a month of training, and the advisory group it is backing at the Institute for Advanced Study has been handed the release schedule to work on.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+55
- Incentives70
- Confidence55
The American Institute of Mathematics asked for problems where AI could search further than a person can, and Rachel Pries offered the one sporadic group with no known polynomial. Six mathematicians closed it in under three months.
Reality
- Evidence66
- Adoption22
- Hype gap+18
- Incentives35
- Confidence62
OpenAI said on August 1 that its then-unreleased Astra model had advanced sphere packing, one of ten claimed results. In a subject with proved optima in only four dimensions, verification is slow.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+45
- Incentives65
- Confidence45
The Clay Mathematics Institute has the claimed Navier-Stokes solution under review, and a summer of agent-found bugs in Lean's kernel shows how narrow the guarantee a Lean certificate actually gives.
Reality
- Evidence62
- Adoption60
- Hype gap+30
- Incentives68
- Confidence58
Levent Alpoge and Ava Howell used a simple prompt to an internal Claude and got elliptic curves of rank at least 30 and at least 31 in days, in a field where the last single-rank step took more than 18 years.
Reality
- Evidence27
- Adoption14
- Hype gap+40
- Incentives68
- Confidence45
Tristan Buckmaster says OpenAI accelerated onto Navier-Stokes after learning of his work with Anthropic's Levent Alpoge, and that the answer he got about his prompt history addressed lookup rather than training.
Reality
- Evidence40
- Adoption30
- Hype gap+45
- Incentives88
- Confidence50
OpenAI says tens of thousands of agents solved a 90-year-old problem carrying a $1 million prize. The solution still needs independent verification, and a rival pair of mathematicians has a claim on the credit.
Reality
- Evidence40
- Adoption45
- Hype gap+40
- Incentives80
- Confidence45
The Fields Medalists behind a new open letter accept that AI can speed up mathematics. Their objection is to solutions announced before anyone has had time to write them up properly or work out whose ideas they used.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+36
- Incentives62
- Confidence41