Skip to content

Invest1 publisher3 min readPublished

Cohere's Gomez credits China's labs with capability they developed themselves

The US cybersecurity agency says industrial-scale distillation is the core of China's AI strategy. Aidan Gomez, who co-wrote the transformer paper, told CNBC that Chinese models now beat the best American ones on some benchmarks.

The Investor · Invest desk

Photograph accompanying Cohere's Gomez credits China's labs with capability they developed themselves
Photo: utoronto.ca

What happened

  • Cohere chief executive Aidan Gomez told CNBC that Chinese AI models are world class and that the lead held by the leading US labs is evaporating very quickly.
  • Gomez said some distillation is taking place but that it cannot explain all of China's progress on AI models, a view CNBC says runs counter to what US labs and Washington have been saying.
  • Sriram Krishnan, a former senior White House policy adviser on AI, said that even if Chinese labs are distilling, it is unclear how much it adds to their ability to advance models.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • contradiction Two American accounts of the same Chinese models imply opposite remedies: in one, controlling access to US models decides who leads; in the other, it postpones an outcome already set by capability built elsewhere.
  • decision Anthropic is staffing threat intelligence against an illicit market in access to Claude, and if distillation is not the main driver of Chinese gains, that spending buys detection and delay without closing any capability gap.
  • precedent Krishnan's wider definition, in which ChatGPT and Claude themselves came out of distilling human content, makes theft a harder foundation for policy because it puts American and Chinese labs under the same description.

Gomez's argument rests on a single logical step, and he put it as an absolute. "And you can't copy or distill to better. You can close the gap and reduce the gap by copying, but you can't outperform," he told CNBC [7]. The premise it needs is that Chinese models are already ahead somewhere, which he located on "some benchmarks, on some axis capabilities" [7].

He grants that the copying happens. Gomez said the talk about the Chinese copying, distilling and cheating was "definitely true to an extent" [6]. Of the four American voices in CNBC's account, two say distillation is the core of Chinese model development and two doubt it explains the gains. The one who runs a lab concedes some of it takes place [19].

So the disagreement is about a fraction, and no one CNBC quoted put a figure on it [20]. The Cybersecurity and Infrastructure Security Agency put that fraction at nearly all of it. It said this month that Chinese AI companies are conducting "systematic extraction of proprietary functionalities and capabilities of U.S. AI companies' models through industrial-scale knowledge distillation campaigns" that form "the core" of their AI development strategy [11]. Sriram Krishnan, until recently a senior White House policy adviser on AI, said even if Chinese labs are distilling, it is unclear how much that adds to their ability to advance models [14]. He also widened the term. ChatGPT and Claude "came out of distilling human content" [12]. "The idea of distilling has always been a core part of how computer science works," Krishnan told CNBC's "Squawk Box" [13].

Both sides here have a stake in the answer. The accusation comes from the company whose model is the alleged target. Anthropic has said on numerous occasions that Chinese labs' distillation amounted to theft [8], and its report this month named Alibaba, Moonshot and DeepSeek [10]. Its head of threat intelligence, Jacob Klein, told CNBC that "there's an entire illicit ecosystem to try to gain access to Claude and other models" [9]. Gomez runs a model developer of his own [3] and co-wrote the 2017 paper that underpins ChatGPT and Claude [4]. China's Ministry of Commerce has pushed back on the distillation claims [15].

Two readings survive the evidence. In the first, distillation cut the cost and the time of catching up, capability developed independently did the rest, and policing access to American models buys a delay of months. In the second, the benchmark wins are narrow enough that CISA's account holds, and access is the thing that decides the race. I lean to the first, because the second requires a distilled model to outrun the model it was distilled from, and nobody quoted claims that has happened. A benchmark breakdown showing the Chinese wins cluster on tasks where teacher outputs are cheap and plentiful would move me the other way.

What to watch

  • Whether Gomez or Cohere names the benchmarks on which Chinese models beat the best American ones, and whether anyone reproduces the result.
  • Whether Anthropic or CISA puts a figure on how much of Chinese model capability distillation accounts for.
  • Whether Alibaba, Moonshot or DeepSeek answer being named in Anthropic's report, and whether China's Ministry of Commerce goes beyond its current pushback.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories