Skip to content

Security1 publisher2 min readPublished

NSA advisory tells AI providers to subtly degrade answers for suspected distillation accounts

The joint NSA, CISA and FBI advisory names six China-based companies and describes billions of tokens taken through paid APIs, cloud resellers and proxy transfer stations, which turns billing telemetry into a detection surface.

The Watch · Security desk

What happened

  • NSA, CISA and the FBI say DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI extracted billions of tokens across millions of requests from US frontier models since at least late 2024.
  • DeepSeek's campaigns targeted reasoning capabilities and domain-specific functions to train its R1 and V3 models, and Alibaba used the same approach to improve its Qwen family.
  • Requests arrived through native APIs, remote cloud providers and third-party aggregators that automatically strip user metadata, with gray-market proxies called transfer stations defeating geographic restrictions.
  • The agencies ask US AI companies to subtly alter responses for accounts suspected of malicious distillation, in order to reduce the value of what those campaigns collect.

Compiled by The WatchSomething wrong?How this is made

Why it matters

  • constraint Response tampering has a measured shelf life: the advisory says these campaigns run quality evaluation frameworks built to detect defensive countermeasures and fail over automatically between pathways when blocked.
  • decision Providers now choose whether to serve deliberately degraded output to paying accounts on suspicion alone, with no threshold, evidentiary bar or appeal path offered in the advisory.
  • exposure Cloud platforms and API aggregators are named both as the routing pathways and as the correlation partners, so the work lands on companies that own no frontier model of their own.
  • cost The remedy is continuous account-level monitoring funded by the provider, not a one-time fix, and it competes for the same trust-and-safety and platform engineering time as everything else.

Two of the three detection signals the advisory asks for sit in billing rather than security telemetry: subscription-to-usage ratios, and new accounts that reach maximum usage immediately [11]. The third is enterprise-scale throughput from an account that is not an enterprise [11]. In most model providers those numbers belong to finance and product analytics. A subscription-to-usage ratio is also the direct observable for the cost trick the advisory describes, since credentials shared across a team of developers consume far more than the seat count implies [20].

The access was bought. The advisory describes bulk procurement of premium subscriptions shared across developer teams as the cost-control method [8], with traffic deliberately spread across multiple providers, platforms and pathways so that no single operator sees the whole campaign [19]. Requests hit variants of Claude, GPT, Gemini and Grok [3], and the extraction techniques include pulling chain-of-thought reasoning out of the responses [9]. The advisory frames this as a terms-of-use violation rather than an intrusion [10], which is why the recommended controls are account-level and behavioural rather than network-level.

On the state question the advisory offers one clause: the extraction happened "likely with Chinese government awareness" [2]. Awareness is the claim. Direction, tasking and funding are not asserted, and the named actors are the companies themselves [2].

The agencies concede the technique. Distillation is described as legitimate and useful in AI research, with the distinguishing features here given as industrial scale, the targeting of restricted proprietary functionality, and evasion of geographic and account controls [15]. Six China-based companies are named against four US model families [17]. The payoff is stated as significantly shorter development timelines and reduced training expenditure [14], and the advisory attaches no figure to either [18].

For anyone who does not sell model access, resell inference or run an aggregator, none of the three actions is theirs to implement: they are addressed to U.S. AI companies [11][12][13]. The advisory asks for a monitoring practice instead of a patch, naming no vulnerability and setting no date. What it establishes is that a paying customer's usage pattern on your own API is now something three federal agencies expect you to detect and answer differently [1][12].

What to watch

  • Whether a follow-up release carries indicators: transfer-station domains, ASNs, aggregator names or account patterns providers can actually match on.
  • Whether any model provider writes response alteration for suspected distillation into its published terms, since it means serving degraded output to paying accounts.
  • Whether the cross-provider correlation gets a venue, an existing ISAC or a new channel joining model owners, clouds and API aggregators.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories