August 25, 2026 — Skild AI released S1, a robotics foundation model that treats a video demonstration as the program. Show a 10-minute unseen task — pour-over coffee, kit assembly, plant potting — and S1 is supposed to execute in real time without task-specific fine-tuning.
Skild founder Deepak Pathak framed it as in-context learning for manipulation graduating from last year's locomotion ICL work. The claim that matters for builders: ~7× performance on unseen long-horizon tasks versus language-only prompting in Skild's reported scaling curves.
TL;DR — what people are asking
| Question | Answer |
|---|---|
| Announced? | August 25, 2026 — Skild blog |
| Input prompt? | One video of human doing the task |
| Horizon? | Up to ~10 minutes, dozens of steps |
| Fine-tune needed? | No per Skild — in-context only |
| Unseen tasks? | Yes — not in pre-training mix (claimed) |
| Completion rate? | 60–80% early real-world (vendor-reported) |
| Weights? | Not open at announcement |
What Skild demonstrated
Skild's release video highlights compositionality and novel test-time behaviors — digging soil for repotting, pressing a coffee filter, flipping pancakes — driven by a single visual demo.
Compared to concurrent ICL manipulation work Skild cites (Generalist AI, Jiang et al.), Skild argues competitors mostly cover short horizons or in-distribution tasks. S1's pitch is out-of-distribution length — the hard part for warehouse and home robots alike.
Skild Brain pre-training: trillions of simulated physics episodes plus millions of human action videos. S1 maps a new video onto that prior — similar in spirit to how LLMs map few-shot examples, but in torque and contact space.
What this means for what you build or pay
Data economics: If video-ICL holds, the cost per new skill shifts from weeks of teleop collection to one good demonstration — relevant for factories retooling lines frequently.
Eval skepticism: No public leaderboard drop — run your embodiment, your lighting, your grippers before budgeting deployment. World models guide covers why sim-to-real gap still kills demos.
Stack placement: Pair simulation from NVIDIA Cosmos with policy models like S1; keep Jetson edge compute plans separate — S1 does not solve deployment hardware.
S1 vs language-conditioned robotics
| Approach | Prompt | Unseen 10-min tasks | Data per skill |
|---|---|---|---|
| Language-only VLA | Text goal | Weak (per Skild) | Medium |
| S1 video ICL | Video demo | Claimed strong | One demo |
| Classical BC + FT | Dataset | Strong in-distribution | Large |
Honest limitations
- Vendor-reported metrics only — wait for third-party replication.
- 60–80% is not factory-ready without safety cages and human oversight.
- Embodiment coverage unclear — humanoid vs arm vs mobile manipulator support not fully documented in GA post.
- Closed platform — no weights for academic ablations yet.
- Sequoia-backed startup risk — distinguish capability research from shipped product SKUs.
Related on explainx.ai
- Figure Index crowdsourced robot dataset (same launch day)
- NVIDIA Siggraph 2026 — Cosmos and physical AI
- NVIDIA Cosmos 3 open physical AI guide
- What are world models?
- Matic cues Jetson Orin Nano robot AI
- Xynova Prima1 dexterous hand WRC
- Xiaomi U0 world foundation model
- Gemini robotics whole-body intelligence
- ViMax agentic video — demo to storyboard
Skild S1 capabilities per Skild AI's August 25, 2026 announcement — verify deployment terms with Skild before production robotics integration.
