Product1 publisher3 min readPublished
OpenAI gave two different answers about Buckmaster's Codex prompts in one week
Tristan Buckmaster told the ABC that the game is up for mathematicians. In the same week, OpenAI described what his Codex prompts did to its models in two incompatible ways, and a customer has no way to pick between them.
The Product Desk · Product desk
What happened
- OpenAI said in a blog post on Tuesday that it had solved the Navier-Stokes existence and smoothness problem, one of the Millennium Problems that carry a $1 million prize for a first correct solution.
- Tristan Buckmaster and Levent Alpoge had used several AI systems on a related problem, OpenAI's Codex among them, and Buckmaster came to suspect the company had used their Codex logs.
- By Buckmaster's account, OpenAI offered two ways to share credit, one of them publishing his results alongside the company's with Alpoge's name removed, and he refused both.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- contradiction Anyone trying to settle the provenance question by reading vendor statements gets two answers from the same company in one week, with no outside test that would pick between them.
- exposure The exposure sits with the customer: whether your unpublished inputs were used is knowable only if the vendor chooses to answer, and declining to answer is one of the available outcomes.
- decision Teams producing publishable work have to decide which drafts may enter a vendor tool before the work starts, because the alternative is negotiating authorship after the vendor has a post ready to publish.
- precedent With the CalTech letter, the remedy for a contested priority claim against an AI lab is a public statement from the field.
Paste an unpublished argument into a vendor's coding assistant and the only complete record of what you typed sits with the vendor. Tristan Buckmaster and Levent Alpoge were in that position while working on a problem related to Navier-Stokes, using several AI systems including OpenAI's Codex [6][7]. Alpoge works for Anthropic, though his collaboration with Buckmaster was solely as an independent researcher [10].
OpenAI has given two accounts of what those prompts did. Earlier in the week the company told reporters it could not rule out that Buckmaster and Alpoge's use of Codex had helped improve its models [15]. A spokesperson then told the ABC that OpenAI "can say categorically that it is impossible for Dr. Buckmaster's Codex prompts over the last two months to have influenced the system in any way, including training" [13]. Buckmaster has said he asked the company whether their work could have been used during model training and got no answer [14]. Gizmodo reported that neither OpenAI nor Buckmaster immediately replied to its requests for comment [16].
Both statements cover the same two months of prompts. Those prompts are described as impossible to have influenced the system and impossible to exclude as an influence [18].
Buckmaster's verdict is about worth. "I think it's pointless," he told the ABC. "Like, I think the game is up" [4]. Gizmodo reports him saying that AI has become so capable at reasoning through complex mathematics that the role of human mathematicians is forever transformed and diminished [5]. That account of the interview does not say AI-produced proofs have stopped being readable or checkable by people [19]. What is disputed in this episode is where the work came from and whose name goes on it. Gizmodo added that it is entirely possible Buckmaster is "getting a little ahead of his skis" [20].
A group of current and former CalTech mathematicians published an open letter on Wednesday saying that "[AI] companies appear guided by an unhealthy instinct to claim certain results before competitors at all costs, regardless of the collateral damage to mathematical understanding" [12].
For whoever has to approve an assistant for a team that produces original work, two questions separate cleanly. The first is contractual: will the vendor commit in writing, for the tier you actually bought, that your inputs are excluded from training and from internal evaluation of model capability? The second is procedural: if the vendor publishes something that resembles your unfinished work, who settles the credit? Buckmaster reached the second question after a blog post was already written, and he says he was offered two credit arrangements and turned both down [9].
The cheap control is a rule about inputs. Unpublished results stay out of the vendor tool until the paper is filed, and the cost of that rule lands on exactly the work where the tool helps most. A team that will not pay that cost is choosing to take the vendor's word for what happened to its prompts, and OpenAI's word this week came in two versions [18].
What to watch
- Whether OpenAI reconciles its two accounts by publishing which Codex logs entered training and internal capability evaluation.