Skip to content

Build2 publishers2 min readPublished

Reka's 19B Rho-1 folds video generation and robot actions into one shared context

Reka released a research preview of Rho-1, a 19B-parameter model that generates text, images and video and emits robot actions in one network. Its robot results so far come from LIBERO simulation tasks, and physical control remains unproven.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Reka's 19B Rho-1 folds video generation and robot actions into one shared context
Generated illustration

What happened

  • Inside its transformer blocks, Rho-1 runs two expert streams, one for language and visual understanding and one for image and video generation, sharing attention and context.
  • Reka says it trained the model from scratch on 320 H100 GPUs over roughly three months.
  • Native video output in the preview is capped at 672 by 384 pixels.
  • Reka announced a $110 million investment backed by NVIDIA and Snowflake in July 2025.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Reka rates object grounding across video as unreliable, and a manipulation policy has to track the object it acts on, so putting Rho-1 on hardware has to wait for that fix.
  • cost A lab trying to reproduce the approach from scratch would need roughly 690,000 H100-hours of compute before it could compare results with Reka's.
  • decision With no public pricing or self-serve download, a robotics team that wants to test Rho-1 on its own tasks has to start by contacting Reka directly.

The shared state is the strongest part of the engineering claim. In the familiar design, an agent links a chain of specialist models [6]. In Rho-1, Reka says, each turn reads from and writes to one shared state, with no second model or tool call in the sequence [7]. The demo conversation starts with a generated lighthouse image. It then locates an object in the scene, animates it, changes the weather in the video and answers a question about the edit [7]. Carrying one scene through all of those turns is clean design, and it is the part of the October 5th preview I would test first [1].

The robotics case reuses the same weights. The model predicts future camera frames and emits joint actions [3]. Reka's bet is that a system that predicts what it will see and produces the action can cut the handoffs between a world model, a planner and a robot policy [5]. Robot training data is scarce, so Reka built an inverse dynamics model that pulls control signals out of ordinary internet video, according to The Decoder [8]. Those action labels are inferred from footage. I'd want to see how far the labeller's errors carry into the joint actions the model emits. The announcement does not show Rho-1 controlling a physical robot [3]. Simulation is the right place to start. Nothing in LIBERO breaks when it falls over.

The speed figures come from Reka's own tests [13]. The base model generates video at a median 0.79 times real time, and a stream starts in roughly six seconds [9]. The Decoder describes the output as continuous real-time video that takes new instructions without restarting [19]. A median tells you about the middle run, and a live stream stalls on the slow ones. The distilled variant cuts the denoising path from 99 steps to eight and produced a 5.3-second clip in about one second [10]. That is about 12 times fewer steps [17], and output at roughly five seconds of video per second of compute [18]. For either figure to transfer, a buyer's prompts, clip lengths and serving hardware would have to match Reka's test setup.

Reka published its failure modes with the preview, and that is good practice. Long video rollouts drift in structure, and targeted edits remain brittle across prompts [11]. The company attributes its limitations largely to training scale and data, and expects both to improve as it scales them [20]. RuntimeWire noted that the preview has not established that scaling will improve them [21].

What to watch

  • A Rho-1 demonstration on a physical robot, with action-rate and latency figures reported alongside the video ones.
  • An independent measurement of the 0.79-times-real-time median and the distilled one-second clip.
  • Whether a larger training run lifts the 672 by 384 output cap and fixes cross-video object grounding, as Reka expects from scaling.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories