SDSignal Desk

Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video

Sep 10, 2026, 9:30 AM · NVIDIA Blog

Image: NVIDIA Blog

Skild’s S1 foundation model learns long-horizon robot tasks from one video demo — no weight update — and NVIDIA’s physical-AI stack is underneath the training loop.

Why it matters

NVIDIA’s blog says Skild AI’s new S1 robot foundation model is designed to learn previously unseen, long-horizon tasks from a single video demonstration. The operator records the task; S1 maps intent, objects, and sequence into actions for the robot in front of it — without updating weights or running task-specific post-training. That is in-context learning for manipulators.

The timing is commercial, not just lab. Skild says it hit a $100 million annual revenue run rate 10 months after its first commercial deployment, with more than 60 deployment partnerships across manufacturing, logistics, inspection, security, food prep, and more. Separately, Skild, NVIDIA, and Foxconn are putting the Skild Brain on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell systems.

From the desk

We’re watching robotics try to exit the retrain-every-SKU trap. Fixed industrial arms die when the line changes. S1’s bet is that a short video prompt is enough to compose skills for tasks lasting up to about 10 minutes — plant potting, pancake making, pour-over coffee, kit assembly — including work that wasn’t in the pretraining set. In one plant-potting test, Skild says the team went from recording the demo to autonomous execution on hardware in 11 minutes.

The numbers they publish are directional and vendor-sourced. On new multistep tasks, S1 succeeded about 66% of the time at each step versus 9% for a similar AI system — a more than sevenfold gap in their framing. Skild also estimates one short video can be as useful as roughly 380 hands-on training examples, which a person might take 50–100 hours to collect. We’re treating those as Skild’s internal comparisons, not an industry benchmark.

Our read: in-context learning for robots is the useful part if it holds on the factory floor, not just in demos. A 66% per-step success rate on unfamiliar work is still a lot of intervention for high-precision assembly — and the Blackwell busbar-and-screws workflow they show with Foxconn is exactly where contact, sequence tracking, and recovery matter. The harm if this scales without honest failure modes is operators trusting a “one video” story while the robot quietly drops parts. NVIDIA’s stack — Cosmos, Omniverse, Isaac Sim, Isaac Lab, TensorRT — is the industrial substrate; Skild is the brain claiming it can adapt without a new training run. I’m watching whether customer lines keep that 11-minute path when the scene is messier than the promo.

Context

Skild built S1 and ran the research on NVIDIA AI infrastructure. Cosmos helps diversify and describe video; Omniverse and Isaac Sim provide simulation; Isaac Lab with the Newton physics engine supports reinforcement learning; Nsight and TensorRT cover training bottlenecks and inference. Skild and NVIDIA are also jointly developing GPU-accelerated simulation solvers for contact and grip, planned for Newton.

Who feels it

Manufacturing operators
A video prompt beats a multi-week retrain cycle if recovery and supervision are real. Plan for the miss rate, not the demo reel.
Robotics startups
Skild’s $100M run-rate claim and 60-plus partnerships raise the bar for what “commercial physical AI” means in a pitch deck.
NVIDIA ecosystem buyers
This is another proof that Cosmos + Isaac + Omniverse is the default training path NVIDIA wants robot companies on.

What to watch

  1. Whether Foxconn’s Blackwell assembly deployment publishes independent yield and intervention rates.
  2. How often S1 needs a second video or human takeover once tasks leave the lab set.
  3. When the joint Skild–NVIDIA contact solvers land in Newton for other developers.

Read the original

Continue at the source.

NVIDIA Blog

Companies: NVIDIA