Product1 publisher3 min readPublished Updated
OpenAI denies reading two mathematicians' Codex logs before racing them to Navier-Stokes
Tristan Buckmaster says OpenAI accelerated onto Navier-Stokes after learning of his work with Anthropic's Levent Alpoge, and that the answer he got about his prompt history addressed lookup rather than training.
The Product Desk · Product desk

What happened
- OpenAI said it had produced an AI-generated solution to the Navier-Stokes equation, one of the Clay Millennium problems, each of which carries a prize of $1 million.
- Sebastien Bubeck said OpenAI began training a new mathematics model on August 28 and, after hearing rumours of Anthropic's progress, put more than 1,000 agents on the problem for over 50 hours.
- Mark Chen, OpenAI's head of research, said the compute cost ran into the millions of dollars, considerably more than the company had spent on any previous mathematics problem.
- Tristan Buckmaster of NYU and Levent Alpoge of Anthropic posted advances in a related area on Monday, work they say they completed using several models including Claude and Codex.
- Buckmaster says OpenAI put several proposals to him, one of which had him announcing the Navier-Stokes solution as the work of an internal OpenAI model with Alpoge's name left off.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- exposure Anyone holding unpublished work in a vendor's agent history now has a worked example of how narrow a reassurance can be, since the answer Buckmaster reports getting covered lookup and left the training question open.
- cost A demonstration that costs more than the prize it targets sets the entry fee for this kind of race at a lab research budget, which prices out the university departments whose credit is being argued over.
- contradiction Neither party's statement can settle the dispute, because OpenAI's account of what made it accelerate and Buckmaster's account of when his work became known are compatible readings of one rumour.
- decision Research groups picking an agent for a live result now have to weigh whether the vendor is also a claimant on the same output, a question that did not arise while the tool was only writing code.
A mathematician mid-problem pastes a half-finished argument into a paid coding agent, because the agent handles Lean formalisation faster than he can by hand. Buckmaster and Alpoge say they used several models, including Claude and Codex, to reach the advances they posted [8]. Vendors tell themselves that unpublished, competitively sensitive work stays off their consumer surfaces. What users actually do is reach for whatever clears the next obstacle, and trust that the terms mean what they appear to mean.
Bubeck said neither the company's researchers nor its agents saw the pair's work before it went public, and that their prompt and proof were not used to prompt OpenAI models or direct its agents [12]. Buckmaster's account of asking about the Codex logs is that he was told the model "didn't look up user data", and that his questions about training went unanswered [10]. Lookup and training are different pipelines with different retention, so an answer about one does not answer the other.
What keeps either statement from closing the credit question is an employment fact. Bubeck says the accelerant was rumours that Anthropic was making progress toward Navier-Stokes [4]; Alpoge is a researcher at Anthropic [8]. The rumour OpenAI says it responded to and the work Buckmaster says it became aware of can be the same information travelling by two routes, and sincerity in a press briefing does not separate them.
More than 1,000 agents running for more than 50 hours is at least 50,000 agent-hours [5][16], and Mark Chen put the cost in the millions of dollars [7]. The Clay prize for the problem is $1 million [2], so the compute alone cost more than the prize pays out [17]. This was a demonstration, timed by a competitor's rumoured progress rather than by the mathematics.
The sourcing here carries a caveat: the proposals, including the one that would have carried Buckmaster's name on the announcement without Alpoge's, come from Buckmaster's own statement [11], and Wired reported that Buckmaster, Alpoge and Anthropic had not responded to its request for comment [15]. OpenAI says it recognises the pair's priority on unforced Euler [13], and its own mathematician, Ven Chandrasekaran, says the two solutions differ in nature [14].
For whoever decides on Monday which agent gets an unpublished result, the split worth making is between two things that usually get merged: whether the vendor's data commitment covers training on prompt content rather than only human access and retrieval, and whether the vendor is a plausible claimant for credit on that particular output. A vendor that only answers the retrieval half, and might publish the same result next month, is fine for boilerplate work; the unpublished work should go to whichever vendor puts the training answer in writing.
What to watch
- Whether OpenAI publishes a written commitment that Codex prompt content is excluded from training, not only from human access and retrieval.
- Whether Buckmaster, Alpoge or Anthropic speak on the record; Wired reported none had responded when it published.
- Whether outside mathematicians confirm the Lean-formalised solution and how they weigh it against the unforced Euler result.