Skip to content

Invest1 publisher3 min readPublished

Chinese models clear 75% of enterprise engineering tasks at a fifth of the US cost

American agencies say six Chinese labs bought bulk subscriptions to US models and trained on the outputs since 2024. Enterprise buyers are meanwhile paying a fifth as much for models that clear most of their engineering work.

The Investor · Invest desk

Illustration accompanying Chinese models clear 75% of enterprise engineering tasks at a fifth of the US cost

What happened

  • The FBI, NSA and CISA said Tuesday that DeepSeek, Moonshot and four other Chinese companies extracted "capabilities worth billions" since 2024 by buying bulk subscriptions to US models and training on the outputs.
  • China's foreign affairs ministry called the accusations groundless and said the country's AI development is a result of high-level scientific and technological self-reliance.
  • Larridin cofounder Ameya Kanitkar said Chinese models including GLM 5.2 and Kimi 2.6 and 2.7 handle around 75% of engineering tasks reasonably well in tracked enterprise workflows, at a fifth of the US cost.
  • A White House report puts 74% of the world's compute in the United States, where hyperscalers are pouring billions into data centers to generate more of it.
  • Hugging Face reported that Chinese open-source models accounted for 41% of total downloads last year, a larger share than US models took.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • constraint One in five business leaders in McKinsey's survey already say costs such as buying tokens limit their use of AI, so a fifth-of-the-price option decides more than which vendor signs: it decides how much of a roadmap gets built at all.
  • capability Because R1 downloads from Hugging Face and can be run on Amazon Web Services, a buyer who will not send data to a China-based company can still take the cheaper weights and host them itself.
  • exposure The channel the officials describe is the American labs' own paid subscription tier. The seller's commercial product sits at the centre of the alleged transfer.

Larridin's two figures do not compound into an 80% saving. If the quarter of engineering tasks the Chinese models do not handle well has to be re-run on an American frontier model at full price, the buyer pays a fifth for three-quarters of the work and full freight for the rest. The blended bill is 45% of an all-US stack [11][1]. The saving is 55%. It has room to degrade, too: the mixed stack stops being cheaper only when the hit rate falls from 75% to 20% [2].

Ameya Kanitkar is describing what buyers pay. The reporting does not say whether the fifth reflects cheaper computation or a price set below the cost of serving. Brendan Burke, the semiconductors and supply chain analyst at Futurum Group, told Fortune that "Chinese labs found algorithms that reduce the complexity of those calculations by an order of magnitude, and then achieve better results because they're able to summarize the most relevant tokens" [7]. He was talking about attention, the mechanism Google researchers introduced in a 2017 paper, which gets more computationally expensive as the context window lengthens [6].

Burke traced that to scarcity. US restrictions on Nvidia's best chips pushed Chinese labs toward domestic alternatives such as Huawei's [9], and "because they had less compute to work with, they found that computationally efficient method instead of just throwing more compute at an inefficient technique, as U.S. labs initially did," he said [8]. He called US frontier labs "token hogs" whose systems are designed to be exploratory and to test base models [16].

Training on a rival's outputs, if the agencies have it right, moves capability into a model. It does not make that model's attention calculations cheaper per token for the customer running it, and what Larridin measures is a running cost [11]. The agencies said the practice let DeepSeek understate its $5.6 million training cost [3].

The Stanford figure and the Larridin figure measure different distances. Stanford put Anthropic's top model ahead of DeepSeek's by 2.7% earlier this year [5]. Larridin's tracked workflows show a 25-point shortfall in tasks handled reasonably well, about nine times the benchmark gap [4].

In my view the workflow number is the one to follow. It is measured on repeated work and can be re-measured as versions ship; the benchmark gap is one snapshot. Kanitkar marks the limit himself: "Frontier U.S. models still have an advantage on the most complex tasks, but Chinese open-weight models are becoming more than capable enough for the majority of everyday enterprise engineering work," he said [12]. A task mix weighted toward that complex tail is what breaks the 45%: at a 60% miss rate, the blended bill is 80% of an all-US stack [5].

What to watch

  • Whether the FBI, NSA and CISA identify the four Chinese companies cited alongside DeepSeek and Moonshot.
  • Whether Chinese model pricing holds at a fifth of US levels once the labs disclose anything about serving economics.
  • Whether Larridin's 75% hit rate moves as GLM and Kimi ship new versions.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories