The new S1 foundation model enables robots to execute complex physical movements by observing a human demonstration rather than relying on language instructions. Skild AI's S1 can learn multi-step processes lasting up to 10 minutes—including making pour-over coffee, assembling kits, or flipping pancakes—from a single video prompt without fine-tuning or post-training. In one performance test, a robot began autonomously potting a plant 11 minutes after the human demonstration was recorded.

S1 achieved a 66% success rate on unseen tasks in internal benchmarks, compared to 9% for a language-prompted Vision-Language-Action (VLA) model using the same 100K hours of pretraining. Skild estimates that one video prompt provides performance equivalent to roughly 380 task-specific post-training demonstrations. The model was trained on Nvidia AI infrastructure and is currently being deployed with a limited group of industrial partners ahead of a wider rollout in the coming months.

Sign in to suggest edits

Key sources

  1. SOURCE@skildai“It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning”x.com
  2. SUPPORT@wallstengine“S1 reached a 66% success rate on unseen tasks versus 9% for a language-prompted VLA at the same 100K hours of pretraining”x.com
  3. SUPPORT@rohanpaul_ai“change the economics of robot learning completely”x.com
  4. SUPPORT@chris_j_paxton“this to me does feel like the true gpt moment for robotics”x.com
Markdown