Skip to content

Product1 publisher3 min readPublished

Buckmaster is using OpenAI's Codex to reconstruct how OpenAI's agents beat him to Navier-Stokes

A week and a half after accusing OpenAI of copying his approach to a problem carrying a $1 million bounty, the NYU mathematician is still running the company's coding agent on his own papers.

The Product Desk · Product desk

Photograph accompanying Buckmaster is using OpenAI's Codex to reconstruct how OpenAI's agents beat him to Navier-Stokes
Photo: abc.net.au

What happened

  • Tristan Buckmaster, a mathematician at New York University, says OpenAI used his work to rush ahead and beat him to a legendary problem carrying a $1 million bounty.
  • In the week and a half since he went public with that accusation, he has been using OpenAI's Codex agent to tidy up his research papers.
  • Andreas Thom, who spent two decades developing the geometric group theory techniques OpenAI said its Astra model used, got the company to amend a release claiming no progress had been made in a decade.
  • Cornell's Alex Townsend says that since August many colleagues have started asking what they need to know about the technology and how to set up subscriptions to higher-powered models.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • contradiction Buckmaster describes a market with little choice, yet the same report has him working the same problem with Anthropic's Claude next to an Anthropic researcher. The binding constraint is the grievance.
  • constraint Attribution is what peer review allocates, and Thom says he expects never to establish whether his own work fed the Astra result. Corrections to the record become a matter of the lab's word.
  • cost The correction burden lands on the aggrieved researchers: both amendments followed emails and public statements from the mathematicians, and the time spent on that is time not spent on mathematics.
  • decision Anyone weighing whether to drop a vendor over provenance now has a worked example where the alternative exists and does not answer the objection. The decision moves onto log retention and audit scope.

Tidying up a paper is clerical work. Working out which logical steps another lab's agents took to get from your own earlier workings to a finished proof is forensic work. Tristan Buckmaster is doing that second job with a tool built by the company he is accusing [3][4]. "Even if you don't agree with any of this, you're kind of stuck. With AI being so useful, it's hard to completely prevent oneself from using it," he told Wired [5]. He also said, "These companies have a monopoly, and there is not much choice" [6].

The monopoly line is harder to square with the rest of the report. Buckmaster worked on Navier-Stokes with Codex and with Anthropic's Claude, alongside the Anthropic researcher Levent Alpoge [7]. A substitute was already in his hands. Switching vendors settles who he pays, and leaves untouched the question he is asking about where the proof came from.

Buckmaster says OpenAI deployed tens of thousands of agents to reach the solution, but only after it learned the equation was close to being solved [8]. The company's response addressed one channel. It investigated and amended its announcement. The amended version says it "confirmed that Buckmaster's Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training" [9]. Two months back from September 8 starts on July 8, so the guarantee covers 62 days of prompt history [22].

Andreas Thom's complaint concerns a published paper from 2019, which no prompt log touches [15]. He told Wired he was amazed by the Astra announcement and wondered "how did they learn about it?" [14]. He asked the OpenAI researcher Mark Sellke whether his and a colleague's ChatGPT sessions on the problem had been fed into training data, and he says Sellke replied: "That did not happen" [16]. Thom has seen the statement about Buckmaster's prompts, says he does not trust it, and accepts he will probably never know whether his own work fed the result [17]. "AI really kills this entire idea that you could trace back who contributed what," Thom said. "That is probably over" [18].

This is one documented case. The Wired report names three mathematicians and describes continued use of OpenAI's tools after a complaint for one of them [24]. It does not include usage, retention or churn figures for the field. The grievances in it are about credit and traceability. Buckmaster called the practice of publishing solutions without fully crediting the human work behind them irresponsible and "childish", and pointed at the timing ahead of major IPOs [11]. Cornell's Alex Townsend, coauthor of a forthcoming book on the field's evolution, said: "If I want to make a contribution to mathematics, how do I do that as just a human nowadays when these trillion-dollar companies are in on the game?" [19]

Two questions sort a provenance objection before anyone pulls a tool. The first is whether a substitute exists that handles the task. The second is whether that substitute answers the objection. Buckmaster's case is a yes then a no. In that cell usage carries on however loud the complaint gets, and the thing left to negotiate is the record: how long prompt logs are kept, and how far back the vendor will audit them. For Buckmaster that reach was the two months before publication [9].

What to watch

  • Whether OpenAI extends its audit past the two-month prompt window or publishes the agent trace behind the Navier-Stokes proof.
  • Whether any of the three mathematicians named in the report moves off OpenAI's tools or gets written data-use terms instead.
  • Whether journals or funders begin asking for a provenance record when an AI-assisted proof is submitted for review.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories