Product1 publisher3 min readPublished
Three of twelve: the humanoid pass rate that matters more than the robot count
Beijing's humanoid games ran a firefighting mission at a working fire brigade. Three of the 12 teams that ran on Sunday finished all three tasks inside the 30-minute limit.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Twenty-three humanoid robot teams competed in a simulated firefighting and rescue mission in Beijing on Sunday, as part of the second World Humanoid Robot Games.
- The emergency management event was held at a real fire brigade rather than a controlled indoor mock-up, requiring robots to navigate realistic scenarios and perform practical tasks.
- Each robot was given 30 minutes to complete three tasks: identify hazardous materials, shut off valves, and extinguish a fire.
- Robots had to identify two randomly placed simulated hazardous substances and report their types through images, locate and close three randomly selected valves, then find a fire extinguisher and spray it until the fire was put out.
- Human firefighters helped stage the exercise by lighting the fire and monitoring whether robots handled the extinguishers correctly.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
Humanoid robot teams at the second World Humanoid Robot Games in Beijing were sent into a simulated firefighting and rescue mission on Sunday, staged at a real fire brigade rather than a controlled indoor mock-up [1][2]. Of the 12 teams that competed that day, three completed the full challenge [6] - a 25 percent completion rate [7], and the only figure from this event worth carrying forward.
The task design is what makes it usable as a yardstick. Each robot had 30 minutes to do three things: identify hazardous materials, shut off valves, and put out a fire [3]. Specifically, that meant identifying two randomly placed simulated hazardous substances and reporting their types through images, locating and closing three randomly selected valves, then finding a fire extinguisher and spraying until the fire was out [4]. Human firefighters lit the fire and watched whether the robots handled extinguishers correctly [5]. Randomised placement, a hard clock, and a human judging tool use are the three ingredients that a staged demo video normally omits.
The environment did the work that mock rooms cannot. Rain, changing lighting and outdoor conditions degraded visual recognition and manipulation [10]. UniX AI's robot finished inside the limit but moved considerably slower than human firefighters and needed two attempts to line its hand up with the extinguisher target [11]. That is a legible failure mode: not a crash, just a retry that would have cost a run with less slack in it.
Two caveats sit on top of the pass rate. First, the denominator is unclear. The same account describes 23 teams competing in the mission and also reports that 12 teams competed on Sunday [1][6]; against 23, the completion rate would be about 13 percent [9]. Anyone quoting a number should say which denominator they are using. Second, autonomy was not uniform. Some teams used VR headsets to remotely control robots shortly before or during their runs, with operators receiving first-person camera views and guiding movements in real time [19]. A completion rate that mixes autonomous and teleoperated runs measures the stack, not the policy, and the two should not be scored together.
The hardware notes point the same way. Different tasks demanded both high-precision joints for hands and high-torque actuators for strength and locomotion [22]; some entrants used omnidirectional wheels rather than feet, with multi-jointed hands for finer manipulation [18]. Several teams ran China-developed joint modules with servo torque spanning low-force precision to high-force movement [20]. Wheels instead of feet is a reasonable engineering choice, and also a reminder that "humanoid" is a marketing category, not a spec.
The scale figures are the part most likely to be repeated and the part that tells you least. The games are slated to bring 2,056 robots from 666 teams and 16 countries [12], roughly three robots per team [24], across 30 competitive and 21 scenario-based events, up from 26 at the inaugural edition [15], with team numbers up 138 percent and robot participation quadrupled [16]. New events include long jump, weightlifting, tug-of-war and table tennis [13]. The organisers, the Beijing Municipal Government and China Media Group, stretched the schedule from three days to five [14].
Watch for three things. Whether future runs publish the autonomy level per attempt, whether the per-task failure breakdown is released rather than just the finish count, and whether the failure data experts say these events generate [21] shows up as a year-on-year change in the completion rate on the same mission. Repeating this fixture with the same three tasks would be worth more than adding a fourth.