Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

Researchers curb unpublished uploads to AI models after two discovery disputes

OpenAI announced an agent-generated Navier-Stokes solution a day after Tristan Buckmaster and Levent Alpöge shared a year of work, Nature reported. No independent check shows OpenAI or Anthropic misused data, so researchers' practical exposure depends on account settings and on who publishes first.

The Engineer · Build desk

How we use AISend a correction

What happened

  • After Buckmaster asked whether his uploaded research shaped the result, OpenAI said its investigation found his prompts from the prior two months could not have affected its system.
  • Separately, Copenhagen doctoral researcher Mario Rodríguez Mestre said he had spent years studying the same CRISPR-like viral DNA patterns Anthropic announced, using Claude.
  • Anthropic told The New York Times that its model was "not trained on any user transcripts," according to Nature.
  • Nature reported that Sandra Laurentino now uses AI only for debugging, and that she swaps generic labels in for the variable names from her experiments.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure A lab on Anthropic's commercial tier keeps a training path open through any member who submits thumbs-down feedback, so confidentiality depends on individual clicks as well as the contract.
  • decision Teams on an OpenAI business or API account have to confirm retention settings on their own, because the training default does not govern how long inputs are kept.
  • constraint Data settings limit what a vendor receives. An AI system can still reach a similar finding faster and announce it first.

An upload can lead to three different uses, and the vendor policies handle each one separately. Inference is the step where the service processes a prompt to produce an answer [8]. The provider may also retain the conversation. Training goes a step further, because the data can shape later versions of the model [9]. By itself, an upload is no proof that the material was retained or used for training [10].

On business and commercial accounts, the training defaults are set correctly. OpenAI's business policy keeps business and API inputs and outputs out of training by default [11]. Retention is handled by separate controls for qualifying organizations [12]. Anthropic's commercial policy is the same by default for inputs and outputs. A user can still open them to training by granting permission or by sending feedback that carries the conversation [13]. The exception is the thumbs-down feedback button. A user who clicks it and submits feedback may send Anthropic the whole conversation connected to that feedback, not just the single prompt [14].

Consumer tiers are looser. ChatGPT users can opt out of training on new conversations, though OpenAI may still use information submitted through feedback [15]. Claude consumer conversations can be used for model improvement with permission, and flagged conversations may support safety enforcement or safeguards training [16]. These policies describe when data may be used. Whether either company reused unpublished research in these disputes is a question they leave unanswered [27].

So far, no independent verification shows misuse by either company [18]. Similar findings cannot reveal how a model reached an answer. Proving misuse would take evidence of what the model or the researchers could access while the work was underway [19]. OpenAI's reported finding covers Buckmaster's prompts from the previous two months [4], for work the pair had pursued for a year [1]. Its wider statement is that neither its researchers nor its agents saw the work before publication [5]. Anthropic's quoted line, that its model was "not trained on any user transcripts," addresses training alone [7]. The Neuron argues that vendor assurances should cover retention and access as explicitly as training [26]. The DNA result is also unsettled: researchers still do not know the biological function of the reported pattern [20].

Priority is a separate concern from data settings. The Neuron frames the second concern as whether researchers can still publish first when AI systems can reach similar findings much faster [21]. The workarounds Nature describes swap one cost for another. Samuel Mehr's lab prohibits uploads of protected information to commercial models and discourages their use during research [22]. Generic labels of the kind Laurentino uses cut what leaves the lab, but they can deprive an assistant of the context it needs [17][23]. Holding preliminary findings until publication gives colleagues fewer chances to review methods early. Labs with approved private systems can keep using AI during research [25][24].

In my context, I would put the control at the account. That means a business or API tier, the organization's retention settings confirmed, and a written lab rule for the feedback button. That keeps real variable names available to the assistant, while publication priority still has to be handled on its own [21].

What to watch

  • Independent evidence of what OpenAI's systems or staff could access while Buckmaster and Alpöge were working, the standard The Neuron sets for showing misuse.
  • Whether OpenAI or Anthropic extend their public assurances from training to retention and access.
  • Whether universities publish rules on which materials researchers may upload to approved AI services, and when exceptions apply.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence40
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence45
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    On September 7, mathematicians Tristan Buckmaster and Levent Alpöge shared progress on Navier-Stokes after a year of work, including a solution to a simpler version of the problem.

    ReportedSupportedSource: The Neuron, citing NatureView cited source
  2. [2]

    According to Nature, OpenAI announced an agent-generated Navier-Stokes solution the day after Buckmaster and Alpöge shared their progress.

    ReportedSupportedSource: Nature, via The NeuronView cited source
  3. [3]

    Buckmaster questioned whether research he had uploaded could have influenced OpenAI's result.

    ReportedSupportedSource: The Neuron, citing NatureView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. theneuron.ai

    1 article · October 8, 2026

    Could AI Beat Scientists to Their Own Research?

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories