Skip to content

Build4 publishers3 min readPublished

A 30-day retention clause routes Nvidia's sensitive work to its own Nemotron models

Anthropic's June decision to keep Fable usage logs for 30 days has turned zero data retention into a procurement gate at Nvidia, Booz Allen and Palantir. The metadata channel stays open under the clause they are demanding.

The Engineer · Build desk

Photograph accompanying A 30-day retention clause routes Nvidia's sensitive work to its own Nemotron models
Photo: tomshardware.com

What happened

  • In June, Anthropic gave itself the right to retain usage logs from its flagship Fable model for 30 days, saying it needed them to defend against complex and novel attacks.
  • A large U.S. utility canceled its plan to test whether Fable could run its core power infrastructure after Anthropic would not agree to a nonrevocable zero-data-retention policy.
  • OpenAI in August began letting GPT-5.6 Cyber customers keep security logs on their own servers, and Anthropic is rolling out a similar program to select customers this fall.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision The retention term now decides model selection before capability does: the frontier model keeps the low-sensitivity tier of work, and in-house or open-source models take everything the buyer classifies as proprietary.
  • constraint An irrevocable ZDR clause deletes the log and leaves metadata, classifier outputs and usage telemetry in place, so the term these buyers are fighting for does not cover the extraction path Schulman describes.
  • precedent With both labs now conceding customer-held log storage to some accounts, customer-controlled logs become the term the next enterprise negotiation opens with.

Retention for misuse detection means a log store. The prompt and the completion land in it, an abuse-review process can reach them, and they age out on a timer [1]. That pipeline does not distinguish a jailbreak attempt from a security vendor's source tree. Bill Vass, Booz Allen's chief technology officer, said he worries a little that Fable might be learning from some of the firm's code [6]. Buyers are arguing over the word irrevocable. Palantir has pressed Anthropic for irrevocable zero-data-retention guarantees before it will make the models available through its own software [7], and CEO Alex Karp told a customer event that companies are tired of being "exploited" by AI labs [8]. The-decoder notes that Palantir's position also suits Palantir, which would prefer customers ran models through its platform instead of going to providers directly [9]. Zero data retention removes the log store and leaves the metadata channel open. Both labs collect metadata and technical usage data from enterprise customers who hold no-training contracts [12]. OpenAI states on its website that it runs business data through automated classifiers and security tools "to better understand how our services are used", and that the resulting classifications "do not contain any of the business data itself" [13]. C Spire, a telecoms company with no-training agreements at both labs, believes the technical usage data its contracts permit includes information about which applications the models are connected to [14]. OpenAI says it does not train on chain-of-thought data [15]; Anthropic says the usage data it collects is aggregated, anonymized and not used for training [27]. John Schulman, an OpenAI co-founder who worked briefly at Anthropic and is now at Thinking Machines, has described where the extraction actually happens [18]. Reinforcement learning tasks built from user traces carry a low risk of reproducing content and can still pull out customer IP; he puts explicit user feedback in reward-model training at the harmless end and "upload user's coding environment and commit history to turn into rl envs" at the invasive end [17]. "De-identification is weak," Schulman said, adding that users can be traced back "with just a small number of bits" [16]. Sarah Hooker, who previously worked at Cohere and Google DeepMind, points to "clever synthetic data techniques that can generate distributional equivalent data while preserving privacy" [20]. Her advice to firms in that position: "If you are a company with IP you have a limited window to build your own intelligence that leverages your IP" [19]. Across the named accounts the split is by sensitivity of the data: open-source projects still go to the frontier model; proprietary cybersecurity code, internal supply chain monitoring and core power infrastructure do not [29]. Justin Boitano, Nvidia's VP of enterprise AI, told The Information that the company believes zero data retention should be on by default [3]. Jensen Huang has said employees should use AI tokens worth half their annual salary every year [24]. The alternatives on the record all cost more than an API key. Northrop Grumman runs open-source models on its own air-gapped servers [21]. Novo Nordisk keeps using Anthropic's Claude for some tasks and bans proprietary data from reaching it [23]. Microsoft is pitching isolated cloud environments that send no data to outside AI companies; Tom's Hardware notes the approach is costly and that at least one customer is still weighing it [22]. All of these accounts trace to a single report by The Information. Palantir, Nvidia, Booz Allen, Anthropic and OpenAI did not immediately respond to Reuters' requests for comment [25].

What to watch

  • Whether Anthropic's customer-held log program becomes a default contract term or stays with select accounts, and whether the guarantee it offers is revocable.
  • Whether Palantir unblocks Fable in its own software, which would indicate Anthropic conceded an irrevocable zero-data-retention term.
  • Whether either lab publishes an enumerated list of the metadata and classifier outputs it retains under a no-training contract.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories