Product1 publisher3 min readPublished
OpenAI now calls it impossible that a Codex user's prompts reached its model
Tristan Buckmaster used OpenAI's Codex on the Navier-Stokes problem, then asked whether his prompts had fed the system that beat him to a proof, and the company's answer to The Verge got firmer over time.
The Product Desk · Product desk
What happened
- OpenAI says roughly 10,000 agents, tens of millions of dollars of compute and 88 hours on an unreleased model produced a solution to the Navier-Stokes problem, one of the Millennium Prize problems.
- The Verge spoke with more than a dozen mathematicians about the company's work in the field, including Tristan Buckmaster and Andreas Thom, both at the center of the recent controversies.
- Until those more recent comments, according to The Verge, the company acknowledged it could not rule out that Buckmaster's Codex use had played a part.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- exposure Every Codex buyer now has a worked example of a customer asking where his inputs went and getting a moving answer, with no log he can read himself.
- constraint A result whose agent count, compute spend and runtime are all measured by the seller is not something a buyer can verify, however good the headline.
- decision Teams routing original research through a vendor's coding agent have to decide whether their contract says anything stronger than the vendor's public statement does.
- contradiction The fight reported here is about provenance and conduct, not correctness, so a buyer treating this as a story about whether the model is good enough is reading the wrong risk.
Buckmaster was doing what plenty of people do with a coding agent. He had been using OpenAI's Codex to work on the Navier-Stokes problem, one of the Millennium Prize problems [1][5]. Then OpenAI announced a solution, and he started asking whether his prompts had contributed to it [5]. According to The Verge, OpenAI denied that anyone or any agent had accessed his specific user data, and until its more recent comments the company acknowledged it could not rule out the possibility [8].
The firmer version came in a statement to The Verge. "We can say categorically that it is impossible for Dr. Buckmaster's Codex prompts over the last two months to have influenced the system in any way, including training," said OpenAI spokesperson Laurance Fauconnet [6]. Buckmaster is not persuaded. "Given their behavior up until this point, one should take such statements with great skepticism," he said [7].
The pitch is a benchmark line: roughly 10,000 agents, tens of millions of dollars of compute, 88 hours, one Millennium Prize problem [2]. The customer experience is a paying user of a coding tool asking a plain data-handling question and getting an answer that moved [8]. The dispute is over provenance and conduct; The Verge's account does not report anyone arguing that the proof itself is wrong, and the two sides broadly agree on the sequence of events [12][18].
The 88 hours works out to 3 days and 16 hours [14]. Take the lowest reading of "tens of millions", $10m, and the run cost about $113,600 an hour [15]. Only OpenAI can count those agents or read those logs, so every figure in the announcement is the vendor's own instrumentation describing the vendor's own claim [2].
Buckmaster's account of what he was offered: practically "unlimited compute" to finish his own work, plus sole authorship of OpenAI's paper announcing the breakthrough, on a path that excluded his collaborator Levent Alpoge, a researcher at Anthropic [9][4]. "All I had to do was throw Levent under the bus," Buckmaster told The Verge in a phone interview [10]. He said he rejected the offer and viewed it as a "bribe" [10]. Alpoge has described the work as a "personal collaboration" independent of his job at Anthropic [13].
Whoever bought Codex seats and told a research team their inputs stay put is the person who has to act on any of this. Two questions sort a lab's capability claim into four boxes: can you reproduce the result with your own inputs, and can you audit what inputs went in. A claim that passes both belongs in next quarter's plan. Pass the first and fail the second, and you have a useful tool whose marketing should not turn up in your own board deck. Fail both and it is a press release with a compute bill attached. On the record The Verge published, the Navier-Stokes run fails the second question, and OpenAI's statement about Buckmaster's prompts is the only account of what went in [8][6].
What to watch
- Whether OpenAI publishes the Navier-Stokes paper with an author list and a description of what went into the run.
- Whether Codex's data-handling terms are clarified or amended for research users after Buckmaster's complaint.
- Whether Anthropic says anything about Alpoge's collaboration, which The Verge reports was a problem for OpenAI.