Invest1 publisher3 min readPublished
China's robot data centres are outrunning their buyers
Twenty-two provincial innovation centres opened within 14 months, and Wuhan alone logs 24,000 training entries a day. But Beijing's Shijingshan centre dropped its robot partner over weak data sales, so subsidy is still paying the bill.
The Investor · Invest desk
What happened
- China had 22 humanoid robot innovation centres at provincial level or above as of the first half of this year, comprehensive hubs for the robotics industry.
- At least 90 lower-tier data collection centres are operating, under construction or being planned across the country.
- The Wuhan centre alone generates 24,000 data entries a day to train robots, with the government paying subsidies for the output.
- Beijing's Shijingshan training centre terminated its contract with partner firm RealMan, citing data quality that fell short of expectations and low sales revenue.
- Scale AI, a US data-labeling firm, puts China's share of commercially available robot AI data at about 90%, with production costs 60% below American levels.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- constraint External data sales were meant to fund the centres and have not, so the size of the training corpus is set each year by what provincial budgets will carry, not by what buyers will pay.
- exposure A robot maker in one of these partnerships books a one-off hardware sale and then depends on data revenue that its government partner can cancel, as RealMan's contract shows.
- decision A lab that rules out Chinese physical-interaction data is choosing to pay roughly 2.5 times as much on Scale AI's cost gap.
- contradiction Global Times treats robot blunders as valuable data while Chen Tao describes a shortage of failure data and Shijingshan complained about quality, so the daily entry counts may not be the asset they appear to be.
Most of these centres are joint operations. A local government partners with a robot maker, buys that maker's robots, and produces training data with them, and the plan is that selling the data to outside buyers covers operating costs [8]. Purchase demand has not been sufficient [8]. Wuhan's daily output annualises to about 8.76 million entries a year [16], and the report does not say what an entry sells for.
The strongest case that the spending pays comes from Scale AI's cost gap, turned round: a buyer who declines Chinese data pays roughly 2.5 times as much for the American equivalent [17]. The estimate comes from a US data-labelling firm that sells into the same market [10]. Guo Ping, chairman of Huawei's supervisory board, named computing infrastructure, data and talent as the three pillars of AI competition and said "China lags the United States in computing infrastructure" [15].
Robot models cannot be trained on internet text and images the way a chatbot is; the material has to be physical interaction data, collected piece by piece in the field [14]. Chen Tao, who heads Fudan University's deep learning research institute, said the "biggest problem with current robot vision-language-action (VLA) models is that there is plenty of success data but a shortage of failure data" [11]. State-run Global Times said of the robot Olympics that closed on Aug. 26 that "spectators see the blunders, but researchers obtain valuable data" [5]. An entry count measures volume. Whether those entries hold recovered failures in a form a model can learn from is a separate question, and the Shijingshan cancellation turned on it [9].
The build rate is the part with no precedent in the earlier campaigns. Twenty-two centres within 14 months of the October 2023 road map from the Ministry of Industry and Information Technology is one opening about every 19 days [6][18]. In electric vehicles and semiconductors, two to three years passed between a policy announcement and a first centre [7]. Now in the third year, the local-government missteps of the battery and solar industries are being repeated in robotics [13].
The corpus is real. Its price is still unsettled. In my view the Shijingshan contract is the first public mark on this asset, and it came in below what the seller expected [9]. Two other readings deserve room. If a VLA model does break through on this material the subsidy will look cheap in hindsight, and Samsung Securities, reporting after a visit to a Chinese robotics exhibition, concluded that "the real bottleneck remains in the cerebrum" [12]. Or the centres keep accumulating the kind of data the labs already hold, in which case a province ends up owning a warehouse of successful tea-pouring. What would prove this wrong is a centre covering its operating costs out of external data sales.
What to watch
- Any provincial centre reporting that external data sales cover its operating costs, and at what price per entry.
- Further partner terminations on the Shijingshan pattern, or renewals, at the other 21 provincial-level centres.
- The year-end reckoning on June's mandate of 100-plus application scenarios and deployment capacity at 10,000 sites.