At 6:23 AM on September 17, 2026, Brett Adcock posted eleven words: "We've had an AI breakthrough at Figure and will be showcasing this tomorrow." By the time this was written the post had 444,800 views and the story was trending with over 220 posts.
The post says nothing. That is the honest starting point, and it is why most coverage of it is worthless. But three verifiable things published in the weeks before it narrow the space considerably, and they all point the same direction.
This article separates what is known from what is inferred, and gives you a checklist for judging the demo when it lands.
TL;DR
| Question | Answer |
|---|---|
| What was announced? | Nothing technical. A one-line teaser and a promise of a demo |
| Who said it? | Brett Adcock, Figure CEO, September 17, 2026, 6:23 AM |
| What is the strongest clue? | Figure robots reported walking outdoors near its HQ since August 22 |
| What did Figure say Helix learned? | To translate human navigation strategies into robot control from 100% human video, no robot demos |
| Best guess at the breakthrough | Locomotion or navigation generalising from egocentric human video to unstructured outdoor environments |
| Confidence? | Inference, not fact. Figure has disclosed nothing |
| What else happened this week? | UBTECH opened a 10,000-unit-per-year humanoid plant on September 12 |
| Should you update your beliefs now? | No. Wait for the demo, then apply the checklist below |
What is actually known
Three things are on the record, all published before the teaser.
1. Index exists and is enormous. Figure launched Index publicly on August 25, 2026. It pays people to record everyday tasks on camera and feeds the footage into the models running its humanoids. Reported numbers: 264,000 downloads across 108 countries, more than 44,000 weekly active users, and over 16 million uploaded videos. Figure describes the ingest rate as 30 minutes of video every second, which works out to roughly 4.9 years of human activity per day. The company has committed to spending over $1 billion on data and compute across the next twelve months.
2. Project Go-Big paired that with real homes. Figure's pretraining effort came out of roughly four months of stealth work and a partnership with Brookfield, which owns more than 100,000 residential units worldwide. This matters because it solves the single hardest problem in home robotics data: getting egocentric footage of ordinary tasks in real, messy, non-lab houses.
3. The training claim is unusually specific. Figure has stated it used "100% egocentric human video data, collected passively as people do behaviors in real Brookfield homes" to train Helix to translate human navigation strategies into robot control, with no robot demonstrations whatsoever.
Read that third point slowly, because it is the load-bearing one.
Why the outdoor sightings matter
On August 22, 2026, roughly four weeks before the teaser, The Humanoid Hub reported seeing Figure robots walking outdoors around the company's headquarters:
The timing cuts both ways, and it is worth being precise about it. Four weeks is long enough that this is not a leak of tomorrow's demo. But it is also evidence that outdoor walking has been in routine testing near a public road for a month, which is not where you put a capability you are still unsure of.
That is a small observation with large implications, for one reason: Helix has historically been an upper-body model. Figure described it as the first vision-language-action model to enable high-rate control of the entire upper body, listing wrists, fingers, torso and head. Legs were conspicuously not on that list. Locomotion on humanoids has conventionally been handled by separate controllers, often reinforcement-learned in simulation, running underneath the manipulation stack.
Put the three facts together:
- Helix was trained to turn human navigation strategies into robot control
- That training used only human video, no robot demonstrations
- Robots are now walking outdoors, in the least controlled environment a humanoid can be put in
The coherent story is that locomotion and navigation have joined the same learned model as manipulation, trained from human egocentric video rather than robot teleoperation. If true, that is a genuine milestone, because it breaks the data bottleneck that has capped humanoid progress: robot demonstration data is slow and expensive to collect, human video is effectively free and already exists at planetary scale.
To be explicit: this is inference. Figure has disclosed nothing. It could equally be a dexterity result, a speed result, a cost result, or a much narrower engineering win dressed in big language. Treat the above as the highest-probability hypothesis, not a leak.
Why outdoors is genuinely hard
It is worth being precise about why a robot walking down a sidewalk is a different problem from a robot walking across a warehouse floor.
Indoor demos quietly assume flat, uniform floors, controlled and constant lighting, known layouts that can be mapped in advance, static or slow obstacles, and no weather. Every one of those assumptions fails outside. Sunlight blows out cameras and moves shadows across the ground. Pavement is cambered, cracked, and littered. Wind applies continuous disturbance to a tall, top-heavy body. Pedestrians and vehicles move unpredictably and fast.
This is precisely the territory Yann LeCun's argument about Moravec's paradox covers: the things humans find effortless are the things machines find hardest, and walking over uneven ground while not falling is the canonical example. It is also what Gemini Robotics 2's whole-body intelligence work has been chasing from a different angle, and what Mistral's Robostral navigation model targets in the embodied-navigation niche.
The week's other story: throughput, not capability
While Figure was teasing, UBTECH was shipping. On September 12, 2026, the Chinese manufacturer began production at a Liuzhou plant it describes as the world's first facility purpose-built for 10,000-unit annual humanoid capacity.
| Detail | Figure (Sept 17) | UBTECH Liuzhou (Sept 12) |
|---|---|---|
| Nature of news | Undisclosed capability teaser | Operational manufacturing plant |
| Concrete output | None yet | One robot roughly every 10 minutes |
| Scale claim | 16M training videos | 10,000 robots per year |
| Products | Figure 03 | Walker S and Cruzr series |
| Verifiable today | No | Yes |
The Liuzhou site covers about 14,000 square metres with a 13.8-metre building height, was built jointly with Siemens Digital Industries Software using digital twin and manufacturing-operations tooling, and includes a 65-square-metre automated warehouse storing 112 humanoids. Cruzr robots and autonomous logistics equipment run parts of the line, which is where the "robots building robots" headline comes from.
These are not the same race. Figure is competing on model capability and data scale. UBTECH is competing on manufacturing throughput and unit economics. A company can win one decisively and lose the other. Anyone framing this as a single leaderboard is selling something.
What people are asking
"Why announce a breakthrough without saying what it is?" Because the teaser is the product for a day. Figure is reportedly valued around $39 billion, already runs more robots than humans at its own facility, and is competing for engineers, partners and capital against OpenAI's humanoid hardware push and well-funded Chinese incumbents. A same-day teaser buys a news cycle the demo alone would not.
"Is the skepticism in the replies fair?" Yes, and Figure has earned some of it by omission rather than dishonesty. The replies are worth reading as a signal: one asks the company to "clarify what you mean by break through," another says "This or we're not interested," and one commenter joked he read it as "we've had an AI breakdown." The word breakthrough has been spent so freely in robotics that it now carries close to zero information. NVIDIA's account replied with an eyes-and-arm emoji, which tells you the industry is watching and also tells you nothing.
"Does human-video training actually work, or is it a data-labelling story?" It is the central bet of the field right now, and it is not unique to Figure. The wager is that egocentric human video contains enough structure about how bodies move through environments to bootstrap robot policies without paired robot data. If Figure demonstrates competent outdoor locomotion from video-only training, that bet starts paying, and the economics of every humanoid programme change, because the expensive input becomes cheap.
"How does this compare to what Figure showed before?" Earlier Helix work covered collaborative manipulation, including two Helix-driven robots tidying a bedroom together. That was upper-body, indoor and cooperative. Outdoor locomotion would be a different axis entirely, not an increment on the same one.
"Who else is close?" 1X's NEO with 25-DOF hands and a physical API is pushing dexterity and developer access. Tau Robotics ran a 30-hour cleaning session in San Francisco, which is an endurance claim rather than a capability one. The Shift free-cleaning programme in NYC is the purest example of the same underlying strategy: give the service away to harvest the data.
How to judge the demo when it drops
Robotics demos are the most misleading artifact in AI, because video hides exactly the variables that matter. Apply this checklist:
- Is it one continuous take? Cuts hide failures, resets and retries. A single unedited shot is worth more than five minutes of montage.
- Is the playback speed stated? Sped-up footage is standard practice and usually disclosed in small text. Check for it. Figure has previously noted its actuators can run over 5x faster than the software currently drives them, so raw speed is not the interesting variable anyway.
- Autonomous or teleoperated? Demand an explicit statement. Partial teleoperation is legitimate engineering and illegitimate marketing when unlabelled.
- Are the objects and the environment novel? Generalisation means handling things the model has not seen. A rehearsed route through a known courtyard proves much less than an unfamiliar street.
- Does it recover? The most informative footage in robotics is a stumble followed by a save. Polished success tells you about the best case; recovery tells you about the distribution.
- What is the success rate across attempts? One good run out of fifty is a highlight reel, not a capability.
- Does the claimed mechanism match the visible behaviour? If Figure says video-only training produced this, ask whether the behaviour looks like it generalises or looks like a memorised trajectory.
If the demo is genuinely video-trained outdoor locomotion with novel routes and visible recoveries, that is a real result worth updating on. If it is a polished cut of a known route at unstated speed, it is a marketing asset.
What this means if you build with AI
Most readers are not building humanoids. The transferable lesson is about data strategy, and it is the most copyable thing Figure is doing.
Figure's actual innovation may not be a model architecture at all. It may be recognising that the binding constraint was paired robot data, and then spending $1 billion to route around it by paying 44,000 people to film themselves doing chores. The model work follows from the data position, not the other way round.
That pattern generalises well beyond robotics: identify the input your competitors assume must be expensive, then find the substitute that already exists at scale. It is the same shape as world models learning environment dynamics from passive video rather than from interaction.
The demo lands today. Judge it on the checklist, not the adjective.
Related on explainx.ai
- Figure Index: the crowdsourced training dataset — the data engine behind the hypothesis
- Figure's robots now outnumber humans at the company — the deployment milestone that preceded this
- Figure Helix 02: two humanoids tidying a bedroom together — the previous Helix capability milestone
- Gemini Robotics 2 and whole-body intelligence — Google's parallel run at the same problem
- 1X NEO: 25-DOF hands and a physical API — the dexterity and developer-access angle
- OpenAI's humanoid hardware push — the newest well-funded entrant
- Tau Robotics' 30-hour cleaning run — endurance as a claim
- Shift's free cleaning as a data-collection play — give the service away, keep the data
- Yann LeCun on LLMs, physical agents and Moravec's paradox — why walking is harder than talking
- Mistral Robostral and embodied navigation — navigation as its own model problem
- World models explained — learning dynamics from passive video
Everything attributed to Figure in this post comes from statements the company or its CEO made publicly before September 17, 2026, plus contemporaneous reporting. The breakthrough itself is undisclosed at the time of writing and the locomotion hypothesis is this author's inference from surrounding evidence, not a confirmed fact. Index download, user and video counts, and the UBTECH plant specifications, come from secondary reporting and have not been independently audited. This post will be updated once the demo is public.
