Build1 publisherNot yet confirmed elsewhere3 min readPublished
Researchers curb unpublished uploads to AI models after two discovery disputes
OpenAI announced an agent-generated Navier-Stokes solution a day after Tristan Buckmaster and Levent Alpöge shared a year of work, Nature reported. No independent check shows OpenAI or Anthropic misused data, so researchers' practical exposure depends on account settings and on who publishes first.
The Engineer · Build desk
What happened
- After Buckmaster asked whether his uploaded research shaped the result, OpenAI said its investigation found his prompts from the prior two months could not have affected its system.
- Separately, Copenhagen doctoral researcher Mario Rodríguez Mestre said he had spent years studying the same CRISPR-like viral DNA patterns Anthropic announced, using Claude.
- Anthropic told The New York Times that its model was "not trained on any user transcripts," according to Nature.
- Nature reported that Sandra Laurentino now uses AI only for debugging, and that she swaps generic labels in for the variable names from her experiments.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure A lab on Anthropic's commercial tier keeps a training path open through any member who submits thumbs-down feedback, so confidentiality depends on individual clicks as well as the contract.
- decision Teams on an OpenAI business or API account have to confirm retention settings on their own, because the training default does not govern how long inputs are kept.
- constraint Data settings limit what a vendor receives. An AI system can still reach a similar finding faster and announce it first.
An upload can lead to three different uses, and the vendor policies handle each one separately. Inference is the step where the service processes a prompt to produce an answer [8]. The provider may also retain the conversation. Training goes a step further, because the data can shape later versions of the model [9]. By itself, an upload is no proof that the material was retained or used for training [10].
On business and commercial accounts, the training defaults are set correctly. OpenAI's business policy keeps business and API inputs and outputs out of training by default [11]. Retention is handled by separate controls for qualifying organizations [12]. Anthropic's commercial policy is the same by default for inputs and outputs. A user can still open them to training by granting permission or by sending feedback that carries the conversation [13]. The exception is the thumbs-down feedback button. A user who clicks it and submits feedback may send Anthropic the whole conversation connected to that feedback, not just the single prompt [14].
Consumer tiers are looser. ChatGPT users can opt out of training on new conversations, though OpenAI may still use information submitted through feedback [15]. Claude consumer conversations can be used for model improvement with permission, and flagged conversations may support safety enforcement or safeguards training [16]. These policies describe when data may be used. Whether either company reused unpublished research in these disputes is a question they leave unanswered [27].
So far, no independent verification shows misuse by either company [18]. Similar findings cannot reveal how a model reached an answer. Proving misuse would take evidence of what the model or the researchers could access while the work was underway [19]. OpenAI's reported finding covers Buckmaster's prompts from the previous two months [4], for work the pair had pursued for a year [1]. Its wider statement is that neither its researchers nor its agents saw the work before publication [5]. Anthropic's quoted line, that its model was "not trained on any user transcripts," addresses training alone [7]. The Neuron argues that vendor assurances should cover retention and access as explicitly as training [26]. The DNA result is also unsettled: researchers still do not know the biological function of the reported pattern [20].
Priority is a separate concern from data settings. The Neuron frames the second concern as whether researchers can still publish first when AI systems can reach similar findings much faster [21]. The workarounds Nature describes swap one cost for another. Samuel Mehr's lab prohibits uploads of protected information to commercial models and discourages their use during research [22]. Generic labels of the kind Laurentino uses cut what leaves the lab, but they can deprive an assistant of the context it needs [17][23]. Holding preliminary findings until publication gives colleagues fewer chances to review methods early. Labs with approved private systems can keep using AI during research [25][24].
In my context, I would put the control at the account. That means a business or API tier, the organization's retention settings confirmed, and a written lab rule for the feedback button. That keeps real variable names available to the assistant, while publication priority still has to be handled on its own [21].
What to watch
- Independent evidence of what OpenAI's systems or staff could access while Buckmaster and Alpöge were working, the standard The Neuron sets for showing misuse.
- Whether OpenAI or Anthropic extend their public assurances from training to retention and access.
- Whether universities publish rules on which materials researchers may upload to approved AI services, and when exceptions apply.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence45
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
On September 7, mathematicians Tristan Buckmaster and Levent Alpöge shared progress on Navier-Stokes after a year of work, including a solution to a simpler version of the problem.
- [2]
According to Nature, OpenAI announced an agent-generated Navier-Stokes solution the day after Buckmaster and Alpöge shared their progress.
- [3]
Buckmaster questioned whether research he had uploaded could have influenced OpenAI's result.
- [4]
OpenAI denied any influence; its investigation found that Buckmaster's prompts from the previous two months could not have affected the system through training or other means.
- [5]
OpenAI said neither its researchers nor its agents saw the pair's work before publication.
- [6]
After Anthropic announced findings on CRISPR-like repeated DNA in viruses made with Claude, Copenhagen doctoral researcher Mario Rodríguez Mestre said he had studied the same patterns for years with Claude's help and questioned whether his data had been reused for training.
- [7]
Nature cites Anthropic telling The New York Times that its model was "not trained on any user transcripts."
ReportedSupportedSource: Anthropic to The New York Times, cited by Nature, via The NeuronView cited source - [8]
When a researcher submits a prompt, the AI service processes it to produce an answer, a step known as inference.
- [9]
The provider may also retain the conversation, while training is a separate use in which data can influence future versions of a model.
- [10]
An upload alone does not show that the provider retained the material or used it for training.
- [11]
Under OpenAI's business policy, business and API inputs and outputs are not used for training by default.
- [12]
For OpenAI, retention has separate controls for qualifying organizations, so researchers handling confidential work need to check their account type and data settings.
- [13]
Anthropic's commercial policy says commercial inputs and outputs are not used for training by default; training may still occur if the user gives permission or submits feedback that includes the conversation.
- [14]
When a user clicks thumbs down and submits feedback, Anthropic may receive the conversation connected to that feedback, not just the single prompt the user originally entered.
- [15]
ChatGPT consumer users can opt out of having new conversations used for training, though the company may still use information users submit through feedback.
- [16]
For Claude consumer accounts, conversations can be used for model improvement with permission, while flagged conversations may support safety enforcement or safeguards training.
- [17]
Sandra Laurentino now limits her AI use to debugging and replaces experimental variable names with generic labels.
- [18]
So far, no independent verification shows that either company misused researchers' data.
- [19]
Similar findings cannot reveal how a model reached an answer; establishing misuse would require evidence showing what information the model or researchers could access while the work was underway.
- [20]
Researchers still do not know the biological function of the reported viral DNA pattern.
- [21]
The episode points to two concerns for researchers: whether unpublished work stays confidential and whether they can still publish first when AI systems can reach similar findings much faster.
- [22]
Samuel Mehr's laboratory prohibits uploads of protected information to commercial models and discourages their use during research.
- [23]
Removing labels can deprive an AI assistant of the context it needs to help.
- [24]
Laboratories with approved private systems can continue using AI during research.
- [25]
Researchers who hold preliminary findings until publication give colleagues fewer chances to review methods early.
- [26]
Vendor assurances should address retention and access as explicitly as training.
- [27]
The vendor policies explain when user data may be used, but they do not show that either company reused unpublished research in the reported disputes.
- [28]
Institutions could set clearer rules on which materials researchers may upload to approved AI services and when exceptions are allowed.
Sources
1 independent publisher whose own reporting we read for this story.
- theneuron.aiCould AI Beat Scientists to Their Own Research?
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- AI Training Data PrivacyFollow
- Research priority and creditFollow
- AI for Scientific DiscoveryFollow