Build1 publisher3 min readPublished
World Labs bets robot progress is a data problem, and moves the test budget into simulation
Its R2S2R engine turns one real demo into thousands of variations, and says sim-only policies ran unattended on real robots. Every number here is the company's own.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- World Labs, the startup founded by AI pioneer Fei-Fei Li, unveiled a simulation engine that trains robot control systems entirely in virtual environments.
- The company's "Real-to-Sim-to-Real" (R2S2R) engine turns real-world robot tasks into simulations for training and evaluating control models, cutting out expensive tests on actual hardware.
- World Labs says the main bottleneck in robot deployment is not model architecture but the sheer volume of experience a robot needs to operate reliably; real-world data is expensive and hard to control, and even online videos do not systematically cover the full range of objects, physical conditions and failure states.
- The technology comes from SceniX, a startup World Labs acquired in July.
- The engine captures robots, sensors, the environment and task demos, then rebuilds them as an interactive virtual world that does not just look like the original but behaves the same way physically, by combining generative world models with task-oriented robot simulation.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
World Labs, the startup founded by Fei-Fei Li, has unveiled a Real-to-Sim-to-Real engine that rebuilds a real robot task as an interactive simulation and trains control models entirely inside it [1][2]. The pitch is a budget argument rather than an architecture one: the company says the binding constraint on deployment is the sheer volume of experience a robot needs, not model design, because real-world data is expensive and hard to control and online video does not systematically cover objects, physical conditions and failure states [3].
The engine, which comes from SceniX, a startup World Labs acquired in July [4], captures the robot, its sensors, the environment and a task demonstration, then rebuilds them as a virtual world meant to behave like the original physically and not merely look like it, combining generative world models with task-oriented robot simulation [5]. From one real task it generates thousands of variations by changing lighting, object position and count, the surrounding environment, physical properties such as friction, and camera angle [6]. To check fidelity, World Labs runs the same action sequence in simulation and reality and compares observations, object movements and outcomes [7]. Demonstrated tasks span rigid, movable and deformable objects: cable routing, inserting an elastic cable end into a hole, two-handed box packing [8].
Models train in simulation and then transfer to real robots [9]. One test platform was ALOHA, the open-source Stanford dual-arm rig operated by puppeteering two smaller control arms, whose blueprints are public and which costs a fraction of commercial systems [10]. According to World Labs, the models each ran for one hour on four further platforms with no human intervention [11], on tasks including wrapping a power cord around a refrigerator with both hands, repositioning test tubes, and separating markers or pencils from a dense jumble [12]. That is at least five platforms [13]. Note the distance between the framing and the figure: the announcement says models run reliably for hours on real hardware [14]; the detail is one unattended hour per platform [11].
The more consequential claim is about evaluation. World Labs argues robot development lags language models because judging a control model has mostly required real hardware [15], and that a simulator need not reproduce real success rates so long as it answers the same questions: where a model fails, which version is better, and whether improvements carry over [16]. In a two-handed cube handoff on ALOHA, the company says the simulation reproduced both the borderline cases where the robot barely caught the cube by an edge and the matching failed attempts [17], and that model rankings in simulation and reality stayed largely the same across model types including GR00T N1.6 and pi-0.5, across training stages, and for both known and unseen cube positions [18]. Each checkpoint was scored on 2,000 simulated runs and 100 real ones, a 20-to-1 ratio [19][20].
If those rankings hold outside the demo, hardware time shifts from screening to confirmation: weak checkpoints die in simulation and the rig is reserved for finalists [21]. World Labs also says a world reconstructed once can be reused for new models and different robots, since the system is not tied to a particular control model or robot type [22], and points to autonomous driving, where some Level 3 and Level 4 systems train on a mix of real and simulated data [23]. Its stated long-term goal is to scale the worlds in which robots learn [24].
All of this is the vendor reporting on its own simulator [25]. Watch whether the rank correlation survives outside controlled setups: the open question, as reported, is how well the results transfer to more complex environments, other robot types and less controlled everyday situations [26]. Watch the 100-run real-world sample, which is the thin end of the comparison [19], and watch whether a reused world actually trains a second robot [22].