SDSignal Desk

Reka Releases Rho-1: A 19B Omni-Reasoning Model That Understands, Generates Video and Outputs Robot Actions in One

Oct 5, 2026, 11:45 PM · MarkTechPost

Image: MarkTechPost

Reka's Rho-1 folds text, images, video and robot control into one 19-billion-parameter model, an ambitious bet against agent pipelines that's still only a research preview.

Why it matters

Most multimodal AI today is a relay race: a central model plans, then hands off to specialists for images, video or object detection. Each handoff adds delay, and each specialist sees only a slice of the task. Reka's Rho-1, a 19B model trained from scratch, tries to do all of it in one network — understanding and generating text, images and video, reasoning across them, and emitting robot actions.

In a demo Reka describes, the model draws a lighthouse, puts a box around it, animates it, turns the scene into a snowstorm and explains the change, in five turns with no tool calls and no second model. It's a research preview with no public weights, API or pricing.

From the desk

We find the architecture idea genuinely compelling. Two expert streams — one for understanding, one for generating images and video — share attention and a single KV cache. That means the generator works from everything the model already knows in the conversation, and things like bounding boxes come out as coordinate tokens rather than from a separate detector. If it works as described, it's a cleaner design than stitching specialists together.

The speed claims are worth noting and worth discounting. Reka reports video at about 0.79x real-time, a first clip in roughly seven seconds against an illustrative 13.8 seconds for a multi-agent pipeline, and a distilled variant that cuts denoising from 99 steps to 8 and returns a 5.3-second clip in about a second. These are vendor-run tests, and that 13.8-second comparison is explicitly illustrative. We'd wait for outside testing.

The robotics piece is the long game. Generating future frames and actions from the same internal state, and using a separate model to infer control signals from raw video, is a way around the scarcity of human teleoperation data. That's the bottleneck holding back a lot of robot learning, so it's a smart place to push.

The constraints are real too: video is capped at 672 by 384, and the model was trained on 320 H100s over three months — modest by frontier standards. And one model that both sees and acts raises the safety bar. When video generation and robot commands share a brain, errors in one can flow straight into the other. I'm watching for how Reka tests that before anything ships beyond a preview.

Context

Rho-1 uses discrete tokens for text and high-level commands and continuous tokens for image latents, video frames, robot actions and proprioception. Training combines next-token prediction with flow matching. A robotics demo in the LIBERO simulation shows it emitting seven action channels.

Who feels it

Agent and app builders
A single omni model could replace brittle multi-model pipelines, if the quality holds up outside demos.
Robotics researchers
Pairing world modeling with action output, plus learning from raw video, targets the data shortage in robot training.
Creative tools
Interactive, steerable video that updates mid-stream points toward new editing workflows, at limited resolution for now.

What to watch

  1. Whether Reka opens weights, an API or pricing beyond the research preview
  2. Independent benchmarks of Rho-1's video quality and latency
  3. Results on physical robots, not just simulation

Read the original

Continue at the source.

MarkTechPost