Skip to content

Build2 publishersAlso reported elsewhere2 min readPublished

One sign error pulled three of OpenAI's AI-generated math papers within a day of release

OpenAI withdrew three AI-generated math papers on October 7, one day after releasing 722 manuscripts, because of a sign error in an argument they shared. Only 300 of the 719 remaining results have Lean proofs a computer can check, so most of the collection still depends on human review.

The Engineer · Build desk

How we use AISend a correction

What happened

  • The error sat in a stabilization-trace cancellation argument, and OpenAI says it also broke the construction that two dependent papers relied on.
  • OpenAI says it gave its unreleased internal model about 4,000 problems, then grouped the results into manuscripts and families after judging their significance.
  • On September 29, an independent advisory group OpenAI consulted urged labs to stop testing advanced math problems on proprietary models and to fund work on human understanding.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost The remaining expense falls on whoever reviews the 419 results without Lean proofs, and that effort sits outside OpenAI's three-hour-per-result compute figure.
  • exposure Results built on shared constructions fail together, so one unchecked error in a reused argument can take its dependent papers down with it.
  • constraint Outside mathematicians can audit the outputs but cannot rerun the unreleased model, so verification has to work from the manuscripts and 10 reasoning summaries alone.
  • precedent A public, versioned repository with withdrawal notices and archived copies gives other labs a working template for retracting AI-generated results quickly.

Add the three withdrawals to the 14 revised papers and the 13 with reference updates, and the change history touches 30 manuscripts [19]. That is about 4 percent of the 722 OpenAI first published [20]. The reporting ties the sign error to the three withdrawals [5]. The other 27 entries are repaired proofs, corrected statements, clarified assumptions and reference updates [7].

In my view the repository handled the correction well. It is public and versioned, so the withdrawal shows up in the change history, and each withdrawn manuscript carries a notice linking to its archived version [6]. A reader who cited one of the three can still see what it said [6].

Lean is the machine check. It is a programming language for encoding proofs so a computer can verify them [10]. About 42 percent of the top-line results have a Lean formalization [9]. That leaves 419 of the 719 without one [18]. The README is candid about this: results sit at different verification stages, and unformalized work could contain issues [8].

Generation has a published cost. OpenAI says the average retained result used compute equal to roughly three hours of ChatGPT Pro thinking [11]. Across 719 results that comes to about 2,150 hours [21], in a unit of measure with one supplier. Runtimewire notes that the figure is OpenAI's own compute comparison, not a measure of human labor saved or mathematical value produced [11].

The reporting does not say how the sign error was found or how many reviewer hours the corrections took. On this evidence, verification is the unfinished part of the job [8]. The published figures measure generation only, so they cannot show that verification is also the larger cost [11].

The selection step adds a second layer to check. Runtimewire points out that OpenAI's process of assessing significance and grouping results describes how the catalog was assembled, and does not independently establish that each remaining result is correct, novel or useful [17].

The advisory group OpenAI consulted has no decision-making power at any AI company [13]. It warns that a person who prompts an AI can end up holding a mathematical argument they cannot follow, check or answer for, and it says the mathematical community must assess the work [16]. OpenAI's October 6 post says the company plans to fund workshops, conferences and special programs on understanding AI-generated results [15].

What to watch

  • Further withdrawals or revisions in the change history, especially among the 419 results that have no Lean formalization.
  • Whether the count of Lean-formalized results rises from 300, and how fast.
  • Whether OpenAI names or releases the model, which would let outside researchers reproduce results from the same system.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence58
Adoption
Insufficient
Hype gap+35
Incentives62
Confidence55
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    OpenAI withdrew three AI-generated mathematical manuscripts on October 7 after finding a sign error that invalidated an argument used across the papers, according to the repository's change history.

    ReportedSupportedSource: runtimewire.com, citing the repository change historyView cited source
  2. [2]

    The withdrawal came one day after OpenAI published a collection it initially described as 722 manuscripts in 372 families of results, produced by an unreleased internal model.

    ReportedSupportedView cited source
  3. [3]

    The OpenAI math repository now records 719 top-line results.

    ReportedSupportedView cited source

Sources

2 independent publishers whose own reporting we read for this story.

  1. cryptobriefing.com

    1 article · October 8, 2026

    OpenAI pulls three AI-generated math papers one day after release
  2. runtimewire.com

    1 article · October 8, 2026

    OpenAI withdraws three math papers after a sign error breaks linked arguments

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Entities

Loading related stories