Skip to content

Invest1 publisher2 min readPublished

Beijing is writing the data standards for humanoid robots before the products ship

The National Data Administration's rules will cover how embodied-AI data is collected, labeled, stored and shared, and they sit on top of a February framework, a June industry benchmark and an end-of-2026 commercialisation target.

The Investor · Invest desk

Photograph accompanying Beijing is writing the data standards for humanoid robots before the products ship
Photo: globaltimes.cn

What happened

  • China's National Data Administration is drafting formal standards for how embodied-AI data is collected, labeled, stored and shared, covering the datasets used to train machines that act in the physical world.
  • The agency will also strengthen data-governance planning guidance for provincial and municipal authorities, and support companies that want to increase their investment in data resources.
  • MIIT's humanoid robotics standardisation committee published China's first national standards framework on February 28, 2026, covering six areas including data lifecycle management.
  • An industry benchmarking standard for embodied AI took effect on June 1, 2026, giving companies a yardstick against which to measure their systems.
  • Measures the National Development and Reform Commission released in late August 2026 set up dedicated training grounds and pilot bases for testing embodied-AI systems in controlled real-world settings.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • decision Any firm already logging teleoperation data has to choose between keeping its house format and re-cutting its archive to a schema still in draft, and the re-cutting bill scales with how much it has recorded.
  • exposure If early alignment eases access to state testing facilities and pilot programmes, as Crypto Briefing expects, a robot maker's test slots depend on its data paperwork.
  • contradiction The regulator's stated purpose is to help firms pool data and to lower the entry cost for companies that cannot fund a corpus of their own, so the biggest private dataset is the asset the policy works against.

A million hours is the number the sector quotes. Crypto Briefing reports that companies in embodied AI are building more than 1 million hours of proprietary or open datasets [6]. At eight recorded hours a day that is 125,000 operator-days, and at 250 working days a year about 500 person-years of someone standing in front of a robot [12]. That is a real bill, and it is one a well-funded competitor can also pay.

What the National Data Administration is drafting sits underneath the hours: how they are collected, labeled, stored and shared [2]. A labeling convention decides whether one company's recorded hours can train another company's gripper at all. There is no scraped-text equivalent in this branch of AI: recorded physical interaction is the entire supply [11].

A firm shipping a warehouse robot in China now answers to three rule-writers at once: the ministry's humanoid committee on the framework [4], the planning commission on the test sites [8], and the data administration on the data [2]. Beijing is targeting notable advances in embodied-AI commercialisation by the end of 2026, and the scaffolding is meant to speed deployment up [9]. From the June benchmark taking effect to that deadline is seven months [13]. The NDA has not said when the draft will be published, or whether it will bind anyone [14].

The part with money attached is access. Crypto Briefing expects that companies aligning early with the emerging standards will find it easier to reach government-backed testing facilities, data-sharing arrangements and pilot programmes [10]. Storage is cheap, and whoever runs a state training ground decides who gets a slot on it [8].

I think the durable advantage is in who gets allocated a slot. The counter-thesis is that the schema gets written around the collection stacks of the firms already at scale, in which case compliance is cheap for them and expensive for late entrants, and the pooling amounts to little because everyone contributes their dull hours and keeps the hard ones. The falsifier is a pilot base open to all comers on published criteria alongside shared datasets that stay thin; that combination puts the advantage back with whoever recorded the most hours.

What to watch

  • Publication of the NDA draft, and whether it sets a compliance date or ties data format to pilot-base admission.
  • Whether the NDRC's training grounds and pilot bases publish admission criteria that reference data collection or labeling practice.
  • Whether MIIT's six-area framework, including its data lifecycle management section, is amended to point at the NDA's rules.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories