Security2 publishers3 min readPublished
NSA, CISA and FBI name six Chinese AI companies as industrial-scale distillers of US models
The conduct described is authorised API traffic at scale rather than an intrusion, the detection asked for is billing telemetry providers already hold, and on Beijing's role the advisory goes no further than likely government awareness.
The Watch · Security desk
What happened
- A joint NSA, CISA and FBI advisory names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI as extracting billions of tokens from Claude, GPT, Gemini and Grok since at least late 2024.
- By the advisory's own count, DeepSeek drew synthetic training data for its R1 and R3 models from twelve US model versions across Anthropic, OpenAI, Google and xAI.
- Moonshot AI is accused of distilling 18 US models to train Kimi-K2 and Kimi K3, using millions of queries aimed at agentic reasoning, coding, data analysis and computer vision.
- The evasion described spreads requests across accounts, models and platforms, routes them through cloud providers and third-party aggregators to obscure metadata, and uses proxies and gray markets to beat geographic limits.
Compiled by The WatchSomething wrong?How this is made
Why it matters
- contradiction CISA's release says the activity happened likely with Chinese government awareness, while CyberScoop reports the agencies called it tacitly encouraged but not directed by Beijing; the difference decides whether a provider is documenting a contract dispute with named customers or a state collection operation.
- cost The recommendation to subtly alter responses for suspected distillation accounts puts the error cost on whoever gets misclassified, because a quietly degraded answer looks like an ordinary answer to the customer paying for it.
- constraint Because the agencies concede that distillation and weight-sharing are routine research practice, any threshold strict enough to catch industrial volume also catches well-funded evaluation work, and the provider has to draw that line on contract terms rather than telemetry.
- precedent Three US agencies have now put six corporate names in a cyber advisory over terms-of-service violations, which makes account-level enforcement against those firms a documented federal position rather than a unilateral commercial call.
The mechanism is a paid API call and a saved answer. Nothing in the advisory describes an intrusion, a stolen weight file or a bypassed authentication control. Knowledge distillation trains a smaller model on a larger model's outputs, and CISA calls it a valid technique that can be misused to acquire a competitor's capabilities in less time and at less cost than building them [13]. The violation asserted is of terms of service [7].
That is why the recommended telemetry is ratio-based rather than volume-based. Requests spread across accounts, models and platforms, then routed through native APIs, remote cloud providers and third-party aggregators to break metadata correlation, defeat any per-account threshold [6]. So the advisory points instead at subscription-to-usage ratios, immediate maximum usage from new accounts, and enterprise-scale throughput on plans that do not justify it [8].
The counts are specific. DeepSeek's inputs break down as four versions of Claude, two of Gemini, five of ChatGPT and Grok 4 [3], which is twelve model versions feeding synthetic training data for R1 and R3 [4]. Moonshot AI is credited with 18 US models behind Kimi-K2 and Kimi K3, among them the system CyberScoop identifies as Fable 5, Anthropic's most advanced commercially available model [5]. Moonshot has been accused in public before: Michael Kratsios of the White House Office of Science and Technology Policy made the same claim about Fable 5 in June, and described the same guardrail-evasion apparatus [16].
On who ordered any of it, the two published accounts sit apart. CISA's release says the extraction happened "likely with Chinese government awareness" [11]. CyberScoop reports the practice is tacitly encouraged but not directed by political leaders in Beijing [12]. Neither is a tasking claim, and the advisory names six commercial companies rather than a service or a unit [2]. What the agencies do assert on the record is that the scale since 2024 makes distillation a critical part of China's AI industrial policy [20], and that the named activity is aggressive, malicious and targeted at industrial scale, as opposed to the weight-sharing and distillation that firms and open-source projects do in ordinary research [14].
The line between those two is contractual and quantitative, not a signature, and it lands on the same providers who are currently defending suits from artists, authors and media organisations over what went into their own training sets [15].
The third ask is the one no single provider can execute alone: correlate activity across model providers, cloud platforms and API aggregators to reveal distributed campaigns [10], supported by information sharing across the US government, industry and allied nations [17]. Nick Andersen, CISA's acting director, urged companies to take immediate steps [18]. Until a correlation channel exists, each provider sees its own slice of traffic the government measures in billions of tokens [2].
What to watch
- Whether any named provider publishes account terminations or token volumes tied to the six companies.
- Whether the cross-provider correlation channel CISA asks for gets an owner, or stays a recommendation.
- Whether a paying customer surfaces evidence of degraded responses after being misclassified as a distillation account.