Reka Releases Rho-1: A 19B Omni-Reasoning Model That Understands, Generates Video and Outputs Robot Actions in One
Oct 5, 2026, 11:45 PM · MarkTechPost

Reka's Rho-1 folds text, images, video and robot control into one 19-billion-parameter model, an ambitious bet against agent pipelines that's still only a research preview.
Why it matters
Most multimodal AI today is a relay race: a central model plans, then hands off to specialists for images, video or object detection. Each handoff adds delay, and each specialist sees only a slice of the task. Reka's Rho-1, a 19B model trained from scratch, tries to do all of it in one network — understanding and generating text, images and video, reasoning across them, and emitting robot actions.
In a demo Reka describes, the model draws a lighthouse, puts a box around it, animates it, turns the scene into a snowstorm and explains the change, in five turns with no tool calls and no second model. It's a research preview with no public weights, API or pricing.
From the desk
We find the architecture idea genuinely compelling. Two expert streams — one for understanding, one for generating images and video — share attention and a single KV cache. That means the generator works from everything the model already knows in the conversation, and things like bounding boxes come out as coordinate tokens rather than from a separate detector. If it works as described, it's a cleaner design than stitching specialists together.
The speed claims are worth noting and worth discounting. Reka reports video at about 0.79x real-time, a first clip in roughly seven seconds against an illustrative 13.8 seconds for a multi-agent pipeline, and a distilled variant that cuts denoising from 99 steps to 8 and returns a 5.3-second clip in about a second. These are vendor-run tests, and that 13.8-second comparison is explicitly illustrative. We'd wait for outside testing.
The robotics piece is the long game. Generating future frames and actions from the same internal state, and using a separate model to infer control signals from raw video, is a way around the scarcity of human teleoperation data. That's the bottleneck holding back a lot of robot learning, so it's a smart place to push.
The constraints are real too: video is capped at 672 by 384, and the model was trained on 320 H100s over three months — modest by frontier standards. And one model that both sees and acts raises the safety bar. When video generation and robot commands share a brain, errors in one can flow straight into the other. I'm watching for how Reka tests that before anything ships beyond a preview.
Context
Rho-1 uses discrete tokens for text and high-level commands and continuous tokens for image latents, video frames, robot actions and proprioception. Training combines next-token prediction with flow matching. A robotics demo in the LIBERO simulation shows it emitting seven action channels.
Who feels it
- Agent and app builders
- A single omni model could replace brittle multi-model pipelines, if the quality holds up outside demos.
- Robotics researchers
- Pairing world modeling with action output, plus learning from raw video, targets the data shortage in robot training.
- Creative tools
- Interactive, steerable video that updates mid-stream points toward new editing workflows, at limited resolution for now.
What to watch
- Whether Reka opens weights, an API or pricing beyond the research preview
- Independent benchmarks of Rho-1's video quality and latency
- Results on physical robots, not just simulation