Product1 publisher3 min readPublished
Humanoid robotics finally has a rubric that punishes the off-camera human
China's World Humanoid Robot Games scored household and hotel runs out of 300 points, docking collisions and drops and halving credit when a person stepped in. The penalties are the interesting part.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The scenario-based events are part of the second World Humanoid Robot Games (WHRG) and test robots in hotel, library and household tasks in less standardised environments.
- The scoring system was designed to penalise mistakes that could matter in a real household: robots lost points for collisions, dropped objects and crossing boundaries.
- If a human intervened during a run, the attempt was treated as remote-controlled and the scoring weight was reduced to 0.5, according to the report.
- Robots were scored out of 300 points for three tasks: tidying a living room, folding clothes in a bedroom, and washing and hanging laundry in a bathroom and on a balcony.
- A household run with a human intervention is capped at an effective 150 points out of 300.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
The second World Humanoid Robot Games has pushed part of its programme out of the arena and into a hotel, a library and a staged household, and it scores those runs with a rubric that deducts points for collisions, dropped objects and crossing boundaries [1][2]. If a human intervenes during a run, the attempt is treated as remote-controlled and the scoring weight is cut to 0.5 [3].
The multiplier is the load-bearing design decision here. A demonstration that says nothing about human help lets a joysticked run and an autonomous one share the same number, which is how most humanoid footage arrives on an operator's desk. Halving the score for intervention prices the assistance instead of hiding it.
The household course was worth 300 points across three tasks: tidying a living room, folding clothes in a bedroom, and washing and hanging laundry in a bathroom and on a balcony [4]. On that arithmetic, a run in which a person steps in once cannot exceed 150 points however neat the folding [5]. Organisers also inserted an unannounced parcel delivery, which required the robot to respond to a voice command, bring a package indoors, and then resume the chores it had abandoned [6]. Resuming is the expensive half of that instruction, because it tests whether the machine kept any state while it was interrupted.
The hotel event began Sunday at the Beijing Continental Grand Hotel, with 30 minutes to move guests' luggage to assigned rooms, replenish towels, slippers and bottled water, remove used linen and make beds [7]. Finishing inside the limit was not sufficient for a high score: neatness was assessed, and autonomous operation carried a higher weighting than teleoperation [8]. The building supplies the difficulty, in the form of unstructured space, deformable soft goods, loaded trolleys that are hard to balance, and long task sequences that expose instability [9]. Analysts pointed to narrow corridors and the job of getting luggage through them [10]. A library event the same day gave robots 30 minutes to collect returned books, shelve them and correct misplaced volumes, which the WHRG called a test of precision, visual recognition and reasoning [11].
Around the scenario work sit six events in real-world settings and eight dexterous-hand micro-manipulation contests, including picking up beans and driving screws [12]. The games run 22 to 26 August at Beijing's National Speed Skating Oval and drew 666 teams and 2,056 robots from 16 countries, roughly three machines per team [13][14]. X-Humanoid entered TianYi 2.0, a wheeled robot that adjusts its torso between 130 and 165 centimetres and runs more than two hours per charge, according to Global Times [15]. The company said its robots use the Huisi Kaiwu platform, and drew the distinction plainly: conventional events chase "higher and faster", scenario events aim at "better at working" [16]. AGILINK, spun off from AgiBot in January, told Global Times that winning the dexterous-hand contest shows a product has passed a "demonstration-grade" threshold for real-world use [17]. Demonstration grade is a marketing threshold, not a procurement one.
Experts quoted in the coverage argue that moving off standardised courses tests mechanical components, algorithms and scenario adaptation together, and that autonomy matters because teleoperation still occupies a second person [18]. That is the cost line a hotel operator would actually model.
What to watch: whether per-run scorecards are published with intervention counts and a split between autonomous and teleoperated attempts. Without those, the 0.5 weight is a stated rule rather than an observed result, and the scenario events remain a better-designed demo instead of a benchmark anyone can audit.