Skip to content

ProductReports disagree3 publishers2 min readPublished Updated

OpenAI's math proofs fall short of a Princeton panel's verification standards

OpenAI released 719 manuscripts of claimed solutions to hard math problems this week and included the model's chain of thought for only ten of them. The advisory group of mathematicians OpenAI consulted had set verification standards the lab's release misses on several counts.

The Product Desk

How we use AISend a correction

Illustration accompanying OpenAI's math proofs fall short of a Princeton panel's verification standards
Generated illustration

What happened

  • The group's first recommendation to frontier labs was to stop testing advanced mathematical problems on proprietary models.
  • OpenAI's release does the opposite, saying it evaluated its proprietary models on open research problems in mathematics.
  • By TechCrunch's count, 42 percent of the released proofs had not been formalized, the check the group recommended for results people cannot follow.
  • OpenAI also left out the machine-readable metadata the group requested to tie each proof's plain-language version to its formal one.
  • Mathematicians at Cambridge and King's College London found at least two discrepancies between the natural-language proof and the Lean code for OpenAI's Navier-Stokes solution.

Why it matters

  • cost The group suggested OpenAI help pay the human mathematicians who would make its solutions meaningful, work that otherwise falls on the rest of the field at its own expense.
  • constraint Verifying any one of these proofs falls to a human who has to work out how the plain-language argument maps to the formal code, the correspondence OpenAI did not supply.
  • exposure Because the documented gaps sit in OpenAI's answer to a million-dollar problem, one of its highest-profile claimed results is among those that cannot yet be taken at face value.

The standard the mathematicians emphasized most was the plainest one: a human should be able to understand the proof [4]. The group behind it, nine researchers at Princeton's Institute for Advanced Study, had published guidelines for frontier labs at the end of September [13]. OpenAI said it consulted the panel to avoid repeating an earlier controversy [12], and on that measure it fell short; it is not clear the lab is taking responsibility for ensuring human understanding follows [4]. The reasoning behind the proofs was mostly withheld, and the chain of thought came with about 1.4 percent of the manuscripts [3][17].

Terence Tao, who has criticized the lab's approach, put the worry plainly after the release. "Problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is 'solved', and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field," he wrote on social media [15].

The second gap is mechanical. A model writes a proof in ordinary language, then restates it in Lean, a programming language that checks a proof by compiling it as code [7]. If the restatement drifts from the prose, the Lean file can compile cleanly while proving something other than what the English claimed. The Cambridge and King's College authors say the gaps they found do not disprove either version, but they do question whether a model can be trusted to formalize its own work [10].

That risk is why the authors want the usual scrutiny applied. "Because of the phenomenon of mistranslations ... the NL proof by OpenAI and other autoformalised Lean proofs should not prima facie be trusted without the same peer review process and scrutiny that other proofs are subjected to," they conclude [8].

OpenAI did follow some of the group's requests, releasing results quickly and describing how the models reached them [9]. The group has not graded the release itself. "It is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully," it said [16]. Until that community has read these proofs and can stand behind them, a lab's claim to have solved a problem is a submission waiting for review.

What to watch

  • Whether the Princeton advisory group publishes a fuller verdict on this specific release.
  • Whether independent mathematicians can formalize the proofs OpenAI released without that step.
  • Whether OpenAI adds chain-of-thought logs and the requested metadata to future proof releases.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence58
Adoption
Insufficient
Hype gap+40
Incentives55
Confidence60

Perspective Coverage

3 publishers
Builder
Builder 44%
Operator
Operator 38%
Investor
Investor 18%
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The advisory group's first request was 'to stop testing advanced mathematical problems on proprietary models.'

  2. [2]

    OpenAI's release explicitly says that it is evaluating its proprietary models using open research problems in mathematics.

  3. [3]

    Just ten of the 719 manuscripts OpenAI released included releases of the model's chain of thought.

Sources

3 independent publishers whose own reporting we read for this story.

  1. techcrunch.com

    2 articles · October 8, 2026

    OpenAI’s math solutions aren’t meeting the field’s standards yet
  2. techrepublic.com

    1 article · October 9, 2026

    OpenAI Claims 370+ Mathematical Results, Researchers Raise Questions
  3. theverge.com

    1 article · October 9, 2026

    ‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest drop

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories