ProductReports disagree3 publishers2 min readPublished Updated
OpenAI's math proofs fall short of a Princeton panel's verification standards
OpenAI released 719 manuscripts of claimed solutions to hard math problems this week and included the model's chain of thought for only ten of them. The advisory group of mathematicians OpenAI consulted had set verification standards the lab's release misses on several counts.
The Product Desk

What happened
- The group's first recommendation to frontier labs was to stop testing advanced mathematical problems on proprietary models.
- OpenAI's release does the opposite, saying it evaluated its proprietary models on open research problems in mathematics.
- By TechCrunch's count, 42 percent of the released proofs had not been formalized, the check the group recommended for results people cannot follow.
- OpenAI also left out the machine-readable metadata the group requested to tie each proof's plain-language version to its formal one.
- Mathematicians at Cambridge and King's College London found at least two discrepancies between the natural-language proof and the Lean code for OpenAI's Navier-Stokes solution.
Why it matters
- cost The group suggested OpenAI help pay the human mathematicians who would make its solutions meaningful, work that otherwise falls on the rest of the field at its own expense.
- constraint Verifying any one of these proofs falls to a human who has to work out how the plain-language argument maps to the formal code, the correspondence OpenAI did not supply.
- exposure Because the documented gaps sit in OpenAI's answer to a million-dollar problem, one of its highest-profile claimed results is among those that cannot yet be taken at face value.
The standard the mathematicians emphasized most was the plainest one: a human should be able to understand the proof [4]. The group behind it, nine researchers at Princeton's Institute for Advanced Study, had published guidelines for frontier labs at the end of September [13]. OpenAI said it consulted the panel to avoid repeating an earlier controversy [12], and on that measure it fell short; it is not clear the lab is taking responsibility for ensuring human understanding follows [4]. The reasoning behind the proofs was mostly withheld, and the chain of thought came with about 1.4 percent of the manuscripts [3][17].
Terence Tao, who has criticized the lab's approach, put the worry plainly after the release. "Problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is 'solved', and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field," he wrote on social media [15].
The second gap is mechanical. A model writes a proof in ordinary language, then restates it in Lean, a programming language that checks a proof by compiling it as code [7]. If the restatement drifts from the prose, the Lean file can compile cleanly while proving something other than what the English claimed. The Cambridge and King's College authors say the gaps they found do not disprove either version, but they do question whether a model can be trusted to formalize its own work [10].
That risk is why the authors want the usual scrutiny applied. "Because of the phenomenon of mistranslations ... the NL proof by OpenAI and other autoformalised Lean proofs should not prima facie be trusted without the same peer review process and scrutiny that other proofs are subjected to," they conclude [8].
OpenAI did follow some of the group's requests, releasing results quickly and describing how the models reached them [9]. The group has not graded the release itself. "It is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully," it said [16]. Until that community has read these proofs and can stand behind them, a lab's claim to have solved a problem is a submission waiting for review.
What to watch
- Whether the Princeton advisory group publishes a fuller verdict on this specific release.
- Whether independent mathematicians can formalize the proofs OpenAI released without that step.
- Whether OpenAI adds chain-of-thought logs and the requested metadata to future proof releases.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+40
- Incentives55
- Confidence60
Perspective Coverage
3 publishers- Builder
- Builder 44%
- Operator
- Operator 38%
- Investor
- Investor 18%
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The advisory group's first request was 'to stop testing advanced mathematical problems on proprietary models.'
- [2]
OpenAI's release explicitly says that it is evaluating its proprietary models using open research problems in mathematics.
- [3]
Just ten of the 719 manuscripts OpenAI released included releases of the model's chain of thought.
- [4]
OpenAI fell short of the group's standards, particularly on the need for human understanding of a result, and it is still not clear the lab is taking 'responsibility for ensuring that human understanding will follow.'
- [5]
The advisory group suggested OpenAI should help fund the work of the human mathematicians who would be required to make the lab's solutions meaningful.
- [6]
A paper by mathematicians at the University of Cambridge and King's College London documents at least two discrepancies between the natural-language proof and the Lean code behind OpenAI's solution to a problem derived from the Navier-Stokes equations.
- [7]
When AI models solve math problems they first create a natural-language explanation, then try to express it in Lean, a programming language that in theory confirms the proof's accuracy by compiling it as code.
- [8]
Because of the phenomenon of mistranslations ... the NL proof by OpenAI and other autoformalised Lean proofs should not prima facie be trusted without the same peer review process and scrutiny that other proofs are subjected to.
ReportedSupportedSource: the authors of the 'lost in translation' paper, in their conclusion2 sources— create a free account to open themView cited source - [9]
OpenAI followed some of the group's principles, including releasing results as soon as possible and including information about how the models reached their conclusions.
- [10]
The discrepancies do not necessarily disprove either solution, but they raise questions about whether models can be relied on to formalize their own solutions without human involvement.
- [11]
The gaps the paper documented appear in a solution to a million-dollar problem ostensibly solved by OpenAI's models.
- [12]
OpenAI said it had consulted an advisory group of elite mathematicians before releasing the proofs, to avoid the controversy that came the last time one of its models solved a long-standing problem.
- [13]
The Advisory Group on Mathematics and Artificial Intelligence, hosted by Princeton University's Institute for Advanced Study, is made up of nine prominent researchers and released guidelines for frontier labs at the end of September.
- [14]
The advisory group asked OpenAI to 'include machine-readable metadata correlating the natural language and formal artifacts,' which the lab did not do with these releases.
- [15]
Problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is 'solved', and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field.
- [16]
It is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully.
ReportedSupportedSource: the Advisory Group on Mathematics and Artificial Intelligence, in a statementView cited source - [17]
Ten of the 719 manuscripts is about 1.4 percent.
- [18]
TechCrunch reported that 42% of the proofs OpenAI released had not undergone formalization, the step the mathematicians suggested for proofs people do not understand.
Sources
3 independent publishers whose own reporting we read for this story.
- techcrunch.comOpenAI’s math solutions aren’t meeting the field’s standards yet
2 articles · October 8, 2026
- techrepublic.comOpenAI Claims 370+ Mathematical Results, Researchers Raise Questions
1 article · October 9, 2026
- theverge.com‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest drop
1 article · October 9, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Formal VerificationFollow
- AI for MathematicsFollow
- Research credit and reproducibilityFollow