Skip to content

Product1 publisher3 min readPublished

Andreas Thom wants OpenAI to prove his ChatGPT sessions stayed out of the training data

OpenAI says no specific user data was accessed to solve the problems it announced, while declining to rule out that de-identified usage data improved its models. For anyone holding unpublished work, that gap is the entire policy question.

The Product Desk · Product desk

Illustration accompanying Andreas Thom wants OpenAI to prove his ChatGPT sessions stayed out of the training data

What happened

  • Thom said his own doubts began after Tristan Buckmaster, a mathematics professor at New York University, publicly questioned whether OpenAI's models had benefited from his use of OpenAI's Codex.
  • In posts on Mastodon, mathematician Andreas Thom said interactions he and colleagues had with ChatGPT before OpenAI's announcement may have contributed to the company's success in his field.
  • One of the ten results OpenAI announced last month concerned non-sofic groups, Thom's specialty, and OpenAI acknowledged the result built heavily on earlier work by Thom and Gabor Kun.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • constraint The only party able to settle whether a user's chats shaped a result is the party being asked, which leaves the researcher with a suspicion and no instrument to test it.
  • exposure If stripping a name is the protection on offer, then the thing worth protecting in a proof sketch travels intact, and every in-progress argument typed into a consumer chat box is exposed on those terms.
  • decision Teams holding unpublished work now have to choose a surface on the basis of written commitments rather than interface defaults, because Thom's scenario is a supplier improving a model that then races the supplier to publication.
  • precedent A second named complainant within days makes provenance questions a standing feature of these announcements rather than one researcher's grievance, and sets the expectation that the burden of proof sits with the lab.

A researcher pastes a half-built argument into a chat window at the end of a long evening, just to see whether one step holds. In most departments that is not a violation of anything. It is how the tool is used, and it is why the sentence pair in OpenAI's Navier-Stokes announcement is the whole of the operator problem.

Two statements sit next to each other there. The narrow one: the researchers and agents did not see any of the users' work by any means until it was public, and no specific user data was accessed in order to solve the problem [9]. The wide one, in the same post: while unlikely, the company cannot rule out that de-identified data derived from usage of its products helped improve its models [10]. Those describe different machinery. The first is about retrieval while the work is being done; the second is about what went into the model beforehand. Denying the first does not touch the second [1]. Thom's account of his own email exchange with Sebastien Bubeck and Mark Sellke is that the reply he got covered only whether his conversations could be accessed directly, not whether they had entered the training pools, and he calls that dishonesty [6].

De-identification may remove a name, but it does not remove the intellectual content of a mathematical idea [11]. For a proof sketch, the identifying material is the argument itself. Stripping the author does nothing to the asset.

Then there is the evidentiary hole. Thom says researchers are not equipped to reverse-engineer the training pipeline and that only OpenAI holds the relevant data, so if the company wants to deny use, the burden is on it to disclose datasets and clarify the settings and terms governing data use [7][8]. He also says he was struck by OpenAI's detailed command of techniques that were neither the most obvious nor the most promising route at the time [16], which is exactly the kind of suspicion that cannot be settled from outside. The pattern around the announcement does not help: after criticism in mathematical circles for omitting recent work by Thom and Gabor Kun, OpenAI quietly amended its writeup [4]. OpenAI did not immediately respond to The Verge's request for comment [15].

So the practical grid for a team with work in progress has two axes, and neither is about model quality. First: will the vendor commit in writing, for the specific surface you are using, that inputs are excluded from training, with something you could hand to a lawyer. Second: is this work in a race, where being second to publish or to file destroys its value. If you have that commitment and there is no race, paste away. If you have it and you are racing, paste, but keep a timestamped local copy so provenance does not depend on anyone's logs. If you do not have it and there is no race, that is fine, but say so out loud rather than pretending it is a rule. If you do not have it and you are racing, the work does not go in the box.

Here is the criterion that settles this in ten seconds, not in a policy document nobody reads: if this exact text appeared in a stranger's published result in eight weeks, what could you show without the vendor's cooperation. Thom's position is that the answer is nothing, because only one party has the pipeline data [7]. That is why the argument is currently being conducted in Mastodon posts [1] rather than under a clause somebody negotiated in advance.

What to watch

  • Whether OpenAI answers the training-data question specifically, rather than the direct-access question, or publishes dataset documentation for the announced math results.
  • Whether any of the other announced results draw similar provenance challenges from authors whose earlier work was involved.
  • Whether universities and funders start demanding written no-training terms with audit rights for surfaces researchers already use daily.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories