Skip to content

Product1 publisher3 min readPublished

Vision plus touch gets a robot hand to 85% success across five tasks and 25 objects

Humanoid locomotion is rehearsed and it films well. The manipulation figure in the published record is 85% success on five tasks with 25 objects. A household or warehouse pilot has to clear that bar on its own objects.

The Product Desk · Product desk

Illustration accompanying Vision plus touch gets a robot hand to 85% success across five tasks and 25 objects

What happened

  • A humanoid robot can run, jump, dance and perform a backflip while still struggling to fold a shirt, pick up a wet sponge or turn a key, according to an account published by Interesting Engineering.
  • A 2026 Science Robotics study from Zhejiang University combined visual and tactile sensing with reinforcement learning and online imitation learning to reach an 85% success rate across five complex tasks involving 25 objects.
  • Unitree's H1 became known for a standing backflip and its G1 for increasingly athletic movement, while Chinese firms including X Square Robot have moved on to testing litter pickup and flower arranging.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost Manipulation data cannot be scraped the way image sets were, so whoever wants the robot folding their shirts pays for the demonstrations, including the synchronized force, torque and tactile rig that records them.
  • exposure Failure lands on whatever the robot is holding, and at one attempt in seven the operator absorbs the breakage and the human recovery time.
  • constraint A better hand comes with the rest of the platform: bimanual coordination, stable navigation and reachability have to be specified together.

Eighty-five percent looks like a passing grade until you count the grasps a chore requires. At that rate, roughly one attempt in seven fails [1]. A shirt fold runs through a sequence: the robot has to locate the fabric, understand its changing shape, choose a grasp point, pull without losing the garment, reposition its hands, account for wrinkles, and place the result [16]. Five contact events at 85 percent each come to 44 percent end to end (0.85^5 = 0.44) [2]. A pilot budget inherits that arithmetic. The Zhejiang tasks were scored individually, so the compounding is a calculation on the reported figure rather than a measured result.

During a grasp, vision, the sensor that scales cheaply, is asking the wrong question. A camera can tell a robot that it is holding an egg, but it cannot directly tell the robot how close it is to crushing it, according to the account published by Interesting Engineering [4]. Motor count does not fix that either. A hand with many degrees of freedom still needs to know where its fingers are, where the object is, how hard it is touching, whether the object is slipping, and how much force is safe [5].

Teams pricing a humanoid tend to treat it as a locomotion platform plus an integration project, because locomotion is what the reel shows [10]. The part that decides whether the deployment works is contact, and researchers at Ohio State University describe physical contact as difficult to model and control, particularly when objects are soft, fragile or irregular [8]. Work on household robotics has named bimanual coordination, stable navigation and sufficient reachability as three requirements for whole-body manipulation, so the hand has to be specified alongside the rest of the platform [9].

The factory version of this problem was solved by removing the variance. An assembly arm gets a fixed trajectory, predictable lighting, known object dimensions and specialized tooling [6]. A home supplies none of it: the shirt is crumpled differently every time, the glass is partly filled, the sponge changes shape when squeezed, and the plastic bag has no fixed geometry at all [7].

The vendors are already moving from locomotion demos to chore testing. Reuters titled a report "After running and dancing, Chinese robot firms target household chores" [11], and X Square Robot, among others, has been testing humanoids on chores including picking up litter and arranging flowers [12]. Those systems have to be trained [15], and a single useful demonstration may need synchronized joint positions, camera feeds, force, torque and tactile readings [13]. Nothing of that shape exists at internet scale the way image collections do [14].

So the grid for anyone evaluating a humanoid has two axes. One is whether your objects hold their shape. The other is whether their position and lighting repeat. Rigid and repeatable is the quadrant industrial arms already own [6]. Deformable and variable is the quadrant where a backflip predicts nothing, because the backflip is a rehearsed sequence and the sponge is not [2]. The test that tells you something costs a morning: the pilot's own object mix in front of the machine, a count of consecutive successes, and the failures priced at what the broken item plus the human recovery actually costs. Published breadth so far is 25 objects across five tasks, an average of five objects per task [3].

What to watch

  • Whether X Square Robot or its peers publish per-attempt success rates on chore sequences instead of demonstration video.
  • Whether a follow-up to the Zhejiang University result reports success beyond 25 objects or in homes the robot has not rehearsed in.
  • Whether humanoid vendors start quoting grasp success rates and object counts in specifications the way locomotion is quoted.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories