Build1 publisher3 min readPublished
Simple AI's robot-free dataset matches teleoperation at more than ten times the demonstrations
The 2,000-hour HiFi-UMI-2K release swaps collection capital for volume. Three policy backbones land within 3.1 points of teleoperated baselines, and the three reported deltas sum to zero.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Xiaofei Li, founder and CEO of Simple AI, released a 2,000-hour robotics dataset on August 20 that eliminates the need for a target robot or robot teleoperation during data collection.
- The HiFi-UMI technical report, first posted on July 28, reports approximate performance parity with conventional robot teleoperation across three policy architectures and four tabletop task groups.
- Simple AI used more than 10 times as many demonstrations than conventional robot teleoperation in the HiFi-UMI comparison.
- Robot teleoperation produces useful data because every recorded action reflects a specific machine's geometry, sensors and physical limits, but it ties each hour of collection to a robot arm, a control rig and an operator, making expansion across homes, hotels, warehouses and other varied environments slow and capital-intensive.
- The HiFi-UMI system records people performing two-handed manipulation tasks with portable sensor-equipped grippers, then converts those demonstrations into trajectories for training real robot arms.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Xiaofei Li, founder and chief executive of Simple AI, released a 2,000-hour robotics dataset on August 20 collected without a target robot or robot teleoperation anywhere in the loop [1]. The accompanying technical report, first posted on July 28, claims approximate performance parity with conventional robot teleoperation across three policy architectures and four tabletop task groups [2], and the price of that parity is stated in the same materials: more than ten times as many demonstrations [3].
That ratio is the actual offer. Teleoperation data is expensive because every recorded action already encodes a specific machine's geometry, sensors and limits, which also ties each hour of collection to an arm, a control rig and an operator and makes expansion across homes, hotels and warehouses slow and capital-intensive [4]. Simple AI's HiFi-UMI system instead records people performing two-handed tasks with portable sensor-equipped grippers, then converts the demonstrations into trajectories for real arms [5]. On the reported evidence, each robot-free demonstration is worth under a tenth of a teleoperated one at the margin [6]. Whether that is a good trade depends entirely on whether your bottleneck is hardware-hours or storage and processing.
The approach descends from the Universal Manipulation Interface, a robot-free human-demonstration system introduced in a 2024 paper by researchers affiliated with Stanford University, Columbia University and Toyota Research Institute [7]. Simple AI's work targets the leftover fidelity problems: trajectory accuracy, synchronization, camera coverage and the relative position of the two grippers [8]. The build puts a stereo camera and inertial sensors on the operator's head and two non-parallel fisheye cameras on each hand module, with a shared hardware trigger aligning them [9]. Simple AI reports roughly 3 millimeters of local end-effector accuracy, sensor synchronization under 40 microseconds and about 200 degrees of visual coverage per hand [10].
The unglamorous part is the rejection step. Each trajectory is reconstructed and replayed in simulation, with reported pass rates of about 98% at both stages [11], which compounds to roughly 96% surviving if the gates run in sequence [12]. Cheap collection stops being cheap the moment engineers have to inspect episodes by hand.
HiFi-UMI-2K holds 2,000 hours and more than 482,000 episodes across over 110 scenes, per its dataset card, with synchronized multi-view video, bimanual trajectories, gripper states, language annotations, subtask boundaries and quality-control metadata, under CC BY 4.0 [13]. That works out to about 15 seconds per episode [14]. Simple AI says the release is less than a tenth of a processed corpus of more than 20,000 hours and 4.32 million episodes across over 480 scenes [15]; by hours the open share is 10%, by episode count about 11% [16].
On the numbers themselves: against policies post-trained on teleoperation data, the three HiFi-UMI-only backbones produced aggregate success-rate differences of -2.5, +3.1 and -0.6 percentage points [17]. Those three values average exactly zero across a 5.6-point spread [18], which is a defensible reading of parity and also a small sample. On a precision-insertion task the strongest HiFi-UMI-only policy hit 85% success, and the report notes the teleoperation baseline had an apparent environmental advantage there [19].
Watch whether outside teams reproduce parity using only the released 2,000 hours, since the internal result was presumably tuned against the full corpus [15][2]. Watch the demonstration ratio: if it drops below ten with better fidelity, the argument gets stronger; if it climbs, the storage and compute bill eats the hardware savings [3][6].