explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR — what people are asking
  • Why Index exists — generalization is a data problem
  • The Index pipeline — consumer app, industrial QA
  • Index vs other data strategies
  • What this means for what you build or pay
  • Robots as a service — Figure's stated endgame
  • Honest limitations
  • What Index was actually built to teach
  • Related on explainx.ai
← Back to blog

explainx / blog

Figure Index: Crowdsourced Robot Training Dataset Goes Public

Figure AI, Physical AI, Humanoid Robots, Robotics, Data Collection

Figure launched Index on Aug 25, 2026: a crowdsourced physical-AI dataset of 16M+ videos from 108 countries, with $1B committed to scale Helix.

Aug 26, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
Figure Index: Crowdsourced Robot Training Dataset Goes Public

August 25, 2026 — Figure AI came out of stealth with Index, a crowdsourced app that pays people worldwide to record everyday physical tasks and feeds that video into Helix, the company's humanoid AI stack. The launch tweet crossed roughly 270K views within a day — a signal that the robotics industry's bottleneck has shifted from hardware demos to data economics.

Figure's thesis is blunt: "The data needed to scale a truly general purpose robot doesn't exist on the internet — it has to come from the real world." Index is Figure's bet that 16 million uploads from 108 countries beats anything a vendor catalog or simulation farm can synthesize.

XSource postOpen on X ↗
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR — what people are asking

table · 2 cols
QuestionAnswer
Announced?August 25, 2026 — Figure blog
What is it?Crowdsourced physical-AI dataset + consumer app (Google Play, App Store)
Stealth stats?264K downloads, 108 countries, 16M+ videos, ~44K weekly active Creators
Upload rate?30 minutes of video per second (~4.9 years of human work/day)
Paid to Creators?$15M to date
Next 12 months?$1B+ committed on data and compute; path to 100× scale
Endgame?Robots as a service — chore-doing humanoids for homes and businesses
Feeds what?Helix VLA policy on Figure 03 humanoids

Why Index exists — generalization is a data problem

Figure frames Index as the missing layer for general-purpose humanoids. Industrial pilots — like the Figure 03 deployment at BMW — prove structured environments. The home is the opposite: cluttered layouts, unfamiliar objects, and tasks nobody pre-labels in a sim.

Helix already demonstrated collaborative home skills in the Helix-02 bedroom tidy demo. Index is how Figure plans to scale the training distribution beyond a few staged rooms.

Per Figure's published metrics, every 1,000 hours of Index data averages:

  • 373 unique tasks
  • 1,146 unique manipulated objects
  • 116 unique environments

That density matters because robotics foundation models — whether Figure's Helix, Skild AI S1's video in-context learning (also announced August 25), or NVIDIA Cosmos physical-AI world models — all hit the same wall: long-tail physical variation that lab teleop cannot cover.

The Index pipeline — consumer app, industrial QA

Figure tried buying data from vendors first. The blog states vendors failed on throughput, diversity, and quality — so Figure rebuilt infrastructure around consumer-app constraints: 24/7 availability, continuous compute, real-time Creator feedback.

The five-stage pipeline:

  1. Filtering — automated technical, visual, and semantic quality screens
  2. Fraud review — human analysts audit user-level evasion attempts
  3. Deduplication — embedding similarity thresholds discard near-duplicates
  4. Rebalancing — task quotas and embedding clusters preserve diversity beyond labels
  5. Annotation — hierarchical text captions on every accepted episode

This is closer to a payments + moderation platform than a robotics lab notebook — and that is the point. Figure needs Creator retention (hence $15M paid out) as much as it needs joint-angle logs.

Index vs other data strategies

table · 4 cols
StrategyWho captures dataScale leverExample
Figure IndexPaid global Creators on phonesCrowdsourced video, 108 countriesAug 25, 2026 launch
China robot academiesCentralized humanoid hardwareState-backed training fieldsShanghai + Hangzhou schools
Brookfield / Go-BigCorporate real-estate partnershipsResidential unit accessFigure's prior pretraining initiative
Sim + world modelsSynthetic rolloutsCompute, not humansNVIDIA Cosmos / Edge MCP
One-shot video ICLSingle demo per taskModel architectureSkild S1

None of these cancel the others. The YC Paper Club robotics session named sim-to-real gap and embodiment drift as persistent blockers — Index attacks the distribution side; Cosmos-class sim attacks throughput; Skild S1 attacks sample efficiency at deploy time.

What this means for what you build or pay

If you train embodied models: Index validates that paid human video at internet scale is now a first-class strategy, not a research side project. Expect more Creator-marketplace launches and tighter data-licensing terms — Figure's pipeline is Figure-exclusive.

If you deploy humanoids: The $1B commitment is a capex signal. Figure is buying its way to home-generalization before competitors lock up the same Creator supply. Watch whether Index data actually moves Helix success rates on unstructured tasks — Figure says internal generalization results are validating the thesis but has not published benchmarks yet.

If you compete on hardware: World Humanoid Robot Games and China's training academies show the hardware race is parallel. Index shows the data race may be won by whoever runs the best global payout app — not whoever has the flashiest demo clip.

Privacy and labor angle: Creators upload real homes, workplaces, and faces. Figure's fraud and dedup stack implies they know quality gaming is inevitable at $15M paid. Builders evaluating similar pipelines should budget for moderation, consent, and regional compliance — not just GPU hours.

Robots as a service — Figure's stated endgame

Figure closes the announcement with a service vision: "Today, you have people coming to help clean your house; eventually, a robot will do everything for you."

That maps to the same home timeline Brett Adcock has cited elsewhere — and to Index's dual mode: record your own chores, or book a Creator who comes to your home or business. The Creator marketplace is training data collection disguised as gig work — a pattern that scales faster than Brookfield-style real-estate partnerships alone.

Whether "robots as a service" arrives on Figure's timeline depends on Helix converting phone video into torque commands on Figure 03 reliably — the gap between world models and closed-loop control is still where most demos die.

Honest limitations

  • Vendor-reported metrics only — 16M uploads and 108-country diversity are Figure's numbers; independent audits are not public.
  • No open dataset release — Index data feeds Helix exclusively; researchers cannot download the corpus.
  • Human video ≠ robot embodiment — phone-camera egocentric footage still requires a transfer bridge; Skild's one-video ICL and Figure's teleop logs solve different slices of that gap.
  • Quality vs quantity tension — 30 minutes/second ingestion is impressive throughput; generalization claims await published evals.
  • Creator economics may not scale linearly — $15M for 16M videos implies roughly $0.94/video average; 100× scale may need very different unit economics or automation.

What Index was actually built to teach

The headline numbers describe the corpus. The more consequential detail is what Figure says it got out of it, because it reframes Index from a labelling operation into a training-method bet.

Figure has stated it trained Helix using "100% egocentric human video data, collected passively as people do behaviors in real Brookfield homes," to translate human navigation strategies into robot control — with no robot demonstrations whatsoever.

Every clause there is doing work:

  • 100% egocentric human video. Not robot teleoperation logs, not simulation, not motion capture. Phone-perspective footage of ordinary people.
  • Collected passively. Contributors were doing the task, not performing it for a camera rig. Passive collection is what makes the volume possible; a staged shoot cannot reach 16 million clips.
  • In real homes. The Brookfield partnership, covering more than 100,000 residential units, is what supplies genuinely messy environments rather than lab mockups. Lab kitchens are tidy in ways real kitchens never are.
  • Navigation strategies, not manipulation. This is the surprise. Helix was publicly an upper-body model, described as the first VLA to control wrists, fingers, torso and head at high rate. Navigation is a legs-and-whole-body problem.
  • No robot demonstrations. The transfer happens without paired robot data at all.

If that holds up, it inverts the economics of the whole field. Robot demonstration data is the expensive input: it needs hardware, operators, and real time, and it scales linearly with money spent. Human video already exists, is generated continuously by billions of people, and costs roughly a dollar a clip to acquire through a gig app.

That is why the $1 billion commitment is better read as a land grab than as a training budget. Whoever assembles the largest corpus of passively-collected egocentric human video owns the input that everyone else has to buy.

The unresolved question

None of this proves the transfer works at the quality bar a product needs. Human video gives you what a task looks like from the inside; it does not give you the robot's own proprioception, its actual joint limits, its mass distribution, or what recovering from its own particular failure modes feels like. Bridging that embodiment gap is the open research problem, and Figure asserting it solved it is not the same as showing it.

Update — September 17, 2026: Figure teased an undisclosed AI breakthrough ahead of a demo. Observers have reported its robots walking outdoors near company headquarters since August 22. Index's design — human egocentric video, no robot demonstrations — is the strongest clue to what the breakthrough is. See what the evidence points to.

Update — September 18, 2026: The breakthrough shipped: Helix 2.5 generalized zero-shot across 30 unseen homes, with an ablation showing Index pretraining alone took zero-shot success from 9% to 56%.

Related on explainx.ai

  • OpenAI confirms it will build its own humanoid robot
  • Figure AI: robots outnumber humans milestone
  • Figure Helix-02 collaborative bedroom tidy
  • Skild AI S1 — robot tasks from one video
  • China humanoid robot training academies
  • NVIDIA Cosmos 3 physical AI guide
  • World Humanoid Robot Games 2026
  • YC Paper Club — why robotics still isn't solved
  • What are world models?

Index stats and pipeline details per Figure AI's August 25, 2026 announcement — verify Creator payout terms and data usage in the app before contributing.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 1, 2026

Figure Melted F.02 in a Finnish Arc Furnace — For Real

On September 30, 2026 Figure published F.02 Decommission: robots trained in San Jose leaped into a 75-ton electric arc furnace in Imatra, Finland. The point was IP destruction, not a new Helix SKU. US and Mexican foundries refused lithium-ion batteries; the film is not generated footage.

Sep 18, 2026

Figure Helix 2.5: Zero-Shot Autonomy Across 30 Unseen Homes

On September 18, 2026, Figure released Helix 2.5 and rented 30 Bay Area homes to prove a specific, falsifiable claim: a single foundation model, pretrained on human video, can walk into a home it has never seen and tidy a room, fold towels, or make a bed with no fine-tuning in that home. The headline ablation is the real story — Index pretraining alone took zero-shot success from 9% to 56%.

Sep 15, 2026

OpenArm: A $6,500 Open-Source Humanoid Arm for Physical AI Research

OpenArm is a fully open-source, 7-degree-of-freedom humanoid arm from Enactic built for physical AI research — teleoperation, imitation learning, and contact-rich manipulation — with a complete stack of hardware, ROS2, Isaac Lab, MuJoCo, and dataset repos. A full bimanual system starts at $6,500.