Reka releases a 19B model for video generation and robot control

CEO Dani Yogatama's team says Rho-1 was trained on 320 H100 GPUs over three months; the research preview still has short-horizon and resolution limits.

By · Published

Primary source: X

Why it matters

Rho-1 tests whether one shared model can replace parts of the multimodal pipeline used for generation, simulation and robot control. Its current preview also makes the gap between an architectural demonstration and reliable physical deployment clear.

Reka (@RekaAILabs): Today, we are releasing a research preview of Rho-1, our 19B omni model that understands and generates text, images, video and robot actions (19-post thread) — Today, we are releasing a research preview of Rho-1, our…

Reka released a research preview of Rho-1 on October 5th, a 19-billion-parameter model that the company says can understand and generate text, images and video, then produce robot actions within one network. Reka's announcement on X pitches a different architecture from the familiar chain of specialist models linked by an agent: Rho-1 keeps modalities in a shared context as a conversation unfolds.

https://x.com/RekaAILabs/status/2107118937490006363

Video from the original post on X.

Yogatama's bet, Reka's co-founder and CEO, is that one model can handle a loop of perception, generation and action without passing a task between separate systems. Yogatama previously spent six years as a senior staff research scientist at Google DeepMind and is an associate professor of computer science at the University of Southern California, according to his biography.

The Rho-1 research post demonstrates a conversation that begins with a generated lighthouse image, locates an object in the scene, animates it, changes the weather in the video and answers a question about the edit. Reka says each turn reads from and writes to one shared state, with no second model or tool call in the sequence. The company describes the architecture as two expert streams inside transformer blocks: one for language and visual understanding, another for generating images and video, with shared attention and context.

For robotics, Reka says the same model predicts future camera frames and emits joint actions. Its demonstration uses LIBERO simulation tasks. That is evidence of a research direction, not a field test: the announcement does not show Rho-1 controlling a physical robot. Reka's thesis is that a system that predicts what it will see and produces the action to take can reduce the handoffs between a world model, a planner and a robot policy.

The speed figures are company-reported. Reka says the base model generates video at a median 0.79 times real time, with a stream starting in roughly six seconds. A distilled variant, which reduces the denoising path from 99 steps to eight, produced a 5.3-second clip in about one second in the company's test. Reka also reports training the model from scratch on 320 H100 GPUs over roughly three months. These figures describe the company's own tests; the launch materials do not provide an independent evaluation of the claims.

Reka also lists constraints that put the preview's scope in focus. Long video rollouts can drift in structure, object grounding across video is still unreliable, and targeted edits remain brittle across prompts. Native video output is capped at 672 by 384 pixels. The company attributes the limitations largely to training scale and data, and expects both to improve as it scales them. The preview has not established that scaling will improve them.

Rho-1 enters a field where large labs are also building interactive world models. Google DeepMind describes Genie 3 as a real-time model for generating navigable environments from text. Reka is aiming to combine that kind of simulation with language reasoning and robot action in the same model. The distinction is architectural; the available demonstrations do not establish that Rho-1 matches systems built specifically for interactive worlds or physical robot control.

The launch continues Reka's effort to build multimodal AI products alongside its research. In July 2025, Reka announced a $110 million investment backed by NVIDIA and Snowflake. Its 2023 funding announcement named DST Global Partners as lead investor, Radical Ventures as a founding investor, and Snowflake Ventures among strategic participants. Those financings frame Rho-1 as research from a venture-backed company with existing commercial platforms, rather than a standalone academic release.

Reka says Rho-1 is available as a research preview and invites teams working on embodied robotics, interactive simulation and vision-action systems to contact it. The launch post does not offer public pricing or a self-serve download. For now, the company's case rests on a technical proposition and demonstrations: one model can carry visual state forward, change it, reason about it and act on it. The stated failure modes show how much work remains before that proposition becomes a dependable product.

Why it matters: Rho-1 tests whether one shared model can replace parts of the multimodal pipeline used for generation, simulation and robot control. Its current preview also makes the gap between an architectural demonstration and reliable physical deployment clear.

Reader comments

Conversation for this story loads after sign-in.