---
format: "aidr-story-markdown/v1"
id: "042faabcce4839affc9f89ba6a93e1f177c965229bc42ae9a1091b349a0af0e0"
canonical_url: "https://aidr.today/042faabc?lang=en"
title: "Skild AI S1 Model Learns 10 Minute Robot Tasks From One Video, Replacing 380 Examples"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-25T18:52:21.000Z"
category: "Research"
topics: ["robotics","reasoning","vision-language-action"]
source_urls: ["https://huggingnews.com/ai/skild-ai-s1-model-learns-10-minute-robot-tasks-from-one-video-replacing-ce34c26a","https://x.com/SkildAI/status/2092300842900865389","https://x.com/wallstengine/status/2092304576917721440","https://x.com/rohanpaul_ai/status/2092312459382235407","https://x.com/chris_j_paxton/status/2092304651995529392"]
summary: "The new S1 foundation model enables robots to execute complex physical movements by observing a human demonstration rather than relying on language instructions. Skild AI's S1 can learn multi-step processes lasting up to 10 minutes—including making pour-over coffee, assembling kits, or flipping pancakes—from a single video prompt without fine-tuning or post-training. In one performance test, a robot began autonomously potting a plant 11 minutes after the human demonstration was recorded. S1 achieved a 66% success rate on unseen tasks in internal benchmarks, compared to 9% for a language-prompted Vision-Language-Action (VLA) model using the same 100K hours of pretraining. Skild estimates that one video prompt provides performance equivalent to roughly 380 task-specific post-training demonstrations. The model was trained on Nvidia AI infrastructure and is currently being deployed with a limited group of industrial partners ahead of a wider rollout in the coming months."
---

# Skild AI S1 Model Learns 10 Minute Robot Tasks From One Video, Replacing 380 Examples

> [Open the canonical story](<https://aidr.today/042faabc?lang=en>)

**Published:** 2026-08-25T18:52:21.000Z
**Category:** Research
**Topics:** robotics, reasoning, vision\-language\-action

## Summary

The new S1 foundation model enables robots to execute complex physical movements by observing a human demonstration rather than relying on language instructions\. Skild AI's S1 can learn multi\-step processes lasting up to 10 minutes—including making pour\-over coffee, assembling kits, or flipping pancakes—from a single video prompt without fine\-tuning or post\-training\. In one performance test, a robot began autonomously potting a plant 11 minutes after the human demonstration was recorded\. S1 achieved a 66% success rate on unseen tasks in internal benchmarks, compared to 9% for a language\-prompted Vision\-Language\-Action \(VLA\) model using the same 100K hours of pretraining\. Skild estimates that one video prompt provides performance equivalent to roughly 380 task\-specific post\-training demonstrations\. The model was trained on Nvidia AI infrastructure and is currently being deployed with a limited group of industrial partners ahead of a wider rollout in the coming months\.

## Sources

- [Story source](<https://huggingnews.com/ai/skild-ai-s1-model-learns-10-minute-robot-tasks-from-one-video-replacing-ce34c26a>)
- [Story source](<https://x.com/SkildAI/status/2092300842900865389>)
- [Supporting source](<https://x.com/wallstengine/status/2092304576917721440>)
- [Supporting source](<https://x.com/rohanpaul_ai/status/2092312459382235407>)
- [Supporting source](<https://x.com/chris_j_paxton/status/2092304651995529392>)

