Build1 publisher3 min readPublished
Buckmaster put every draft into Codex before OpenAI described the same proof path
The allegations against OpenAI include a demand to drop an Anthropic-affiliated co-author, but the detail that carries over to other projects is what a vendor's research staff may hold after you paste in your drafts.
The Engineer · Build desk

What happened
- Tristan Buckmaster and Levent Alpoge spent months on the Navier-Stokes equations using Anthropic's Claude and OpenAI's Codex running GPT-5.6 Sol, and say they had several breakthroughs by mid-August.
- Buckmaster emailed a mathematician at OpenAI on September 3 to say the work was a personal collaboration with no institutional agreements, and OpenAI answered the same day, offering compute and pressing for a call.
- On two Sunday calls, Sebastien Bubeck's side told him an internal model had produced a roughly 100-page proof for Navier-Stokes with forcing, initially described as needing very little human input.
- Buckmaster says every draft the pair wrote went into OpenAI Codex sessions throughout the project, and that the forced-Navier-Stokes route OpenAI described was one almost nobody else was working on.
- Bubeck twice asserted he wanted Alpoge removed from authorship because Alpoge works at Anthropic, and when Buckmaster refused, asked why he would ruin his career.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure Unpublished drafts left in a vendor's session store are reachable by that vendor's own research organisation on whatever terms the account carries, and the customer cannot audit that reach from outside.
- constraint Telling contamination apart from independent convergence needs retention records and training-data provenance, so nobody outside the vendor can settle the question from the finished proof.
- decision Collaborations that span competing labs now have a reason to fix authorship and affiliation in writing before a vendor with compute to offer takes an interest in the result.
- contradiction OpenAI's own August account of a low-compute Millennium Problems effort sits against the claimed 100-page proof weeks later, which means the capability claim cannot be weighed without a compute figure.
The two questions Buckmaster says he asked on those calls are about different systems, which is why one answer does not cover both. He says he was told the model did not look up user data, and that when he asked specifically about training he got no answer [8]. The first is a statement about inference time: no retrieval into a session store while the proof was being generated. Training is a question about what a retained transcript was eligible for weeks or months earlier. A model whose corpus included session data satisfies the first statement and still carries the drafts in its weights. A third path needs no training at all, only a human on an internal team with read access to transcripts. Buckmaster's account confirms only the first of these: no lookup at inference time. The other two, whether the drafts entered training and whether an internal team could simply read them, are not addressed by anything reported here.
The contamination reading depends on conditions that are not confirmed here: whether the sessions were retained under terms permitting training or internal inspection, whether that retained content reached whoever worked the problem, and whether the forced Navier-Stokes route was as rare as Buckmaster says it was [6]. Only the third is something a mathematician can assess from where he sits, and he asserts it. The other two live in retention logs and contract terms, and nothing in the report shows which account or which terms the Codex sessions ran under.
The timeline is checkable arithmetic rather than inference. Noam Brown said at the Astra launch in early August that OpenAI had been attempting the Millennium Problems, had not found a solution, and had not put much compute toward the effort [9]. Buckmaster's outreach was September 3, and the calls where a roughly 100-page proof of forced Navier-Stokes was described followed that same week [3][4]. That is about four to five weeks [10]. Either the effort scaled hard inside that window, which the later description of massive compute would fit [5], or the August statement understated what was already running. Either reading fits the timeline; whether session data was the source of the proof path OpenAI described is a separate question that this timeline does not settle.
Every element here is Buckmaster's, from his public statement as reported by the-decoder; both supplied source blocks are the same article [15]. OpenAI also replied the same day he wrote and offered compute resources [3], which is a brisk turnaround for an unsolicited note about somebody else's personal project. Of the allegations, the authorship demand is the most checkable, because it was made twice across two calls with other participants present [12], and because Buckmaster reports the rejoinder when he refused [13].
What is left for anyone else is unglamorous. If unpublished work goes into a vendor session, the retention and training terms on that account are the only thing between the draft and that vendor's research organisation, and Buckmaster's experience is that those terms do not get clarified retroactively on a phone call [8]. Dated notes, kept the way he kept them, are what make the question askable at all.
What to watch
- OpenAI publishing its forced Navier-Stokes result, or responding to Buckmaster's statement, either of which moves this off a single account.
- Formal verification and publication of the Buckmaster-Alpoge result for the Navier-Stokes variant, which is still unverified.
- Any disclosure of the retention and training terms the Codex sessions actually ran under.