Skip to content

Build1 publisher3 min readPublished

A 4B world model on the robot: Cosmos 3 Edge posts 22.9% in closed loop

NVIDIA's tutorial post-trains a 4B model into a Franka manipulation policy that runs on Jetson Thor with about 0.6 seconds of latency slack per cycle. Closed-loop success is 22.9%.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying A 4B world model on the robot: Cosmos 3 Edge posts 22.9% in closed loop
Generated illustration

What happened

  • Cosmos 3 Edge is a 4B omni-model, including a 2B NVIDIA Nemotron-based reasoner, in the Cosmos 3 family.
  • Cosmos 3 Edge was pretrained on the same physical-world data as Cosmos 3 Nano and Cosmos 3 Super, starting with the same grounding in how objects move and interact.
  • The model is small enough to run on-device on NVIDIA Jetson Thor.
  • The tutorial covers post-training Cosmos 3 Edge to predict robot actions, serving the policy on Jetson Thor, running inference inside a receding-horizon control loop, and evaluating behaviour in closed-loop simulation.
  • Each step is reproducible from the open cosmos-framework repo, and the released checkpoint is available on HuggingFace.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

NVIDIA has published a walkthrough for post-training Cosmos 3 Edge, a 4B-parameter omni-model with a 2B Nemotron-based reasoner, into a robot manipulation policy that fits and runs on a Jetson Thor module [1][3][7]. The consequence that matters is architectural rather than benchmark-shaped: the inference loop moves off a data-center GPU and onto the arm, and the whole pipeline is reproducible from the open cosmos-framework repo with the checkpoint published on HuggingFace [5][7].

The two constraints NVIDIA names are the honest ones: the model plus runtime state has to fit in the robot's memory, and the full inference pipeline has to be fast enough for the required control frequency [6]. Memory is handled by size. Cosmos 3 Edge was pretrained on the same physical-world data as Cosmos 3 Nano and Super, so it inherits the same grounding in contact and motion at a footprint that fits on-device [2][3].

Latency is handled by chunking rather than speed. On a Jetson AGX Thor T5000, running at 640x540 and 15 Hz, the DROID policy produces an action chunk in about 1.53 seconds, and each chunk covers roughly 2.13 seconds of robot motion [8][9]. That is 0.60 seconds of slack per cycle, or about 32 action steps generated per inference [1][3]. Compute consumes roughly 72 percent of the window it is buying, a headroom factor of about 1.39 [2]. The arm keeps moving because the next chunk lands before the current one runs out, and the policy replans after each inference cycle rather than after every observation [9][10]. That is a real-time system with a thin margin, not a comfortable one: anything that inflates inference time by 40 percent closes the gap.

The scoreboard is where the marketing usually goes quiet, and to NVIDIA's credit the number is stated plainly. On closed-loop RoboLab tasks, the post-trained policy reaches 22.9 percent success [11], which means it fails roughly 77 percent of attempts [4]. That is a backbone demonstration, not a deployable manipulator.

The training data is nvidia/Cosmos3-DROID: 76,000 successful teleoperated trajectories, about 350 hours across 86 tasks and 564 scenes, collected on a Franka Panda arm with a Robotiq gripper [12]. Averaged out, that is about 16.6 seconds per trajectory, which tells you these are short, filtered episodes [5]. The set ships in LeRobotDataset v3.0 format at 640x360, and preparation is three stages: drop idle and non-task frames, keep the successful demonstrations, then apply random cropping, rescaling and colour jitter during training [13]. Note the resolution reported at inference, 640x540, is not the resolution the dataset is packaged at [6].

One asymmetry to keep in view: inference is on-device, but post-training is not. NVIDIA's validated training hardware is a DGX Station with a GB200 or GB300 Grace Blackwell superchip, on CUDA 13.0 and the NGC 26.06-py3 container [14]. Reproducing the loop means owning or renting that.

Watch the port cost. For a DROID-like Franka the primary change is a dataset path, but any other embodiment needs its own experiment configuration covering action space, dimensionality, camera layout and normalization [15]. Cosmos 3 lists dual-arm Franka, UR, WidowX 250 and LeRobot SO101 among supported embodiments [16], so the near-term test is whether someone outside NVIDIA gets a non-Franka policy to that same 1.53-second chunk budget, and whether the 22.9 percent moves.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories