explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR — what people are asking
  • What Skild demonstrated
  • What this means for what you build or pay
  • S1 vs language-conditioned robotics
  • Honest limitations
  • Related on explainx.ai
← Back to blog

explainx / blog

Skild AI S1: Robot Tasks From One Video Demonstration

Skild AI launched S1 Aug 25, 2026 — a robotics foundation model that runs up to 10-minute unseen manipulation tasks from a single in-context video prompt, no task-specific fine-tuning. 60–80% completion in early real-world tests.

Aug 26, 2026·3 min read·Yash Thakker
Skild AIRoboticsPhysical AIWorld ModelsNVIDIA Cosmos
go deep
Skild AI S1: Robot Tasks From One Video Demonstration

August 25, 2026 — Skild AI released S1, a robotics foundation model that treats a video demonstration as the program. Show a 10-minute unseen task — pour-over coffee, kit assembly, plant potting — and S1 is supposed to execute in real time without task-specific fine-tuning.

Skild founder Deepak Pathak framed it as in-context learning for manipulation graduating from last year's locomotion ICL work. The claim that matters for builders: ~7× performance on unseen long-horizon tasks versus language-only prompting in Skild's reported scaling curves.

TL;DR — what people are asking

table · 2 cols
QuestionAnswer
Announced?August 25, 2026 — Skild blog
Input prompt?One video of human doing the task
Horizon?Up to ~10 minutes, dozens of steps
Fine-tune needed?No per Skild — in-context only
Unseen tasks?Yes — not in pre-training mix (claimed)
Completion rate?60–80% early real-world (vendor-reported)
Weights?Not open at announcement
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

What Skild demonstrated

Skild's release video highlights compositionality and novel test-time behaviors — digging soil for repotting, pressing a coffee filter, flipping pancakes — driven by a single visual demo.

Compared to concurrent ICL manipulation work Skild cites (Generalist AI, Jiang et al.), Skild argues competitors mostly cover short horizons or in-distribution tasks. S1's pitch is out-of-distribution length — the hard part for warehouse and home robots alike.

Skild Brain pre-training: trillions of simulated physics episodes plus millions of human action videos. S1 maps a new video onto that prior — similar in spirit to how LLMs map few-shot examples, but in torque and contact space.

What this means for what you build or pay

Data economics: If video-ICL holds, the cost per new skill shifts from weeks of teleop collection to one good demonstration — relevant for factories retooling lines frequently.

Eval skepticism: No public leaderboard drop — run your embodiment, your lighting, your grippers before budgeting deployment. World models guide covers why sim-to-real gap still kills demos.

Stack placement: Pair simulation from NVIDIA Cosmos with policy models like S1; keep Jetson edge compute plans separate — S1 does not solve deployment hardware.

S1 vs language-conditioned robotics

table · 4 cols
ApproachPromptUnseen 10-min tasksData per skill
Language-only VLAText goalWeak (per Skild)Medium
S1 video ICLVideo demoClaimed strongOne demo
Classical BC + FTDatasetStrong in-distributionLarge

Honest limitations

  • Vendor-reported metrics only — wait for third-party replication.
  • 60–80% is not factory-ready without safety cages and human oversight.
  • Embodiment coverage unclear — humanoid vs arm vs mobile manipulator support not fully documented in GA post.
  • Closed platform — no weights for academic ablations yet.
  • Sequoia-backed startup risk — distinguish capability research from shipped product SKUs.

Related on explainx.ai

  • Figure Index crowdsourced robot dataset (same launch day)
  • NVIDIA Siggraph 2026 — Cosmos and physical AI
  • NVIDIA Cosmos 3 open physical AI guide
  • What are world models?
  • Matic cues Jetson Orin Nano robot AI
  • Xynova Prima1 dexterous hand WRC
  • Xiaomi U0 world foundation model
  • Gemini robotics whole-body intelligence
  • ViMax agentic video — demo to storyboard

Skild S1 capabilities per Skild AI's August 25, 2026 announcement — verify deployment terms with Skild before production robotics integration.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Aug 11, 2026

DYNA-2 World-Action Model and the Robotics Scaling Law Claim

Dyna Robotics unveiled DYNA-2 on August 10, 2026, a "world-action model" pre-trained on over 1,000,000 hours of egocentric human video with no robot data at all. explainx.ai unpacks what a world-action model is versus a VLA, what a scaling law actually claims, and what the published exponents do and don't prove.

Jul 14, 2026

Xiaomi-Robotics-U0 — 38B World Model That Boosts π₀.₅ OOD Success to 63%

Xiaomi's 38B autoregressive world foundation model keeps general T2I and editing in the training mix while learning multi-view robot scene synthesis. Synthetic embodied transfer data nearly doubles π₀.₅ out-of-distribution success on real manipulation tasks. explainx.ai breaks down tasks, architecture, and limits.

Aug 26, 2026

Figure Index: Crowdsourced Robot Training Dataset Goes Public

Figure AI came out of stealth on August 25, 2026 with Index — a global app that pays people to record everyday tasks and feeds the footage into Helix, Figure's humanoid AI. explainx.ai maps the 16M-upload pipeline, the $1B data bet, and how Index compares to China's robot academies and Skild's one-video learning stack.