GPT-6 Astra had a busy 24 hours. First, reports circulated a 7x lead over Claude Fable 5.1 on MazeBench, scoring 14% against Fable 5.1's apparent single-digit result. Then a video from Sharif Shameem went viral showing Astra clearing every one of the 48 levels in Neal Agarwal's browser game "I'm Not a Robot" — a puzzle that escalates from simple checkbox CAPTCHAs to visually identifying stop signs — using the model's computer-use tools. Both results are worth understanding on their own terms, and neither one, despite the online reaction, is evidence of general intelligence.
TL;DR
| Result | Reported figure | What it tests |
|---|---|---|
| MazeBench | 14% for Astra, ~7x Claude Fable 5.1's score | Spatial navigation and planning |
| "I'm Not a Robot" game | All 48 levels cleared | Real-time browser/UI interaction via computer-use tools |
| Is this AGI? | No, per Chollet and Schmidhuber's public pushback | Neither test measures invention or real-world mastery |
| Does it break real CAPTCHAs? | Not directly — production systems use additional signals | Device data, behavioral timing, account history |
| Who built the game? | Neal Agarwal, at neal.fun | A 2025 browser puzzle, not a security product |
MazeBench: a 7x lead, but a low absolute score
MazeBench tests spatial navigation and planning — tasks that require a model to reason about layout, position, and path-finding rather than pattern-match against text. Reports put GPT-6 Astra at 14%, described as roughly seven times Claude Fable 5.1's score on the same benchmark. Both numbers matter here, not just the multiplier: a 7x lead sounds dramatic, but 14% overall is still a low absolute score on a benchmark that's evidently hard for every current frontier model. That combination — large relative gains on a benchmark everyone still struggles with in absolute terms — is a pattern explainx.ai has flagged before, including in GPT-6 Astra's own ARC-AGI-3 launch numbers, where headline percentages needed harness-level context to interpret correctly.
The CAPTCHA gauntlet: what computer-use tools actually did
Neal Agarwal's "I'm Not a Robot" is a 2025 browser game, not a production security product — it's a deliberately escalating gauntlet of "prove you're human" puzzles, starting with a simple checkbox and ending in genuinely tricky visual-identification challenges. Astra reportedly cleared all 48 levels using its computer-use tools — the same category of capability explainx.ai covered in OpenAI Codex's computer-use expansion to Windows and mobile control — meaning it interacted with the live page directly: clicking, reading rendered visual content, and adapting to each new puzzle type as the game presented it, rather than following a scripted, pre-planned sequence.
That's a genuinely more impressive demonstration than a static benchmark score, because it requires the model to handle dynamic, unpredictable interface changes in real time rather than solving a fixed input. Reaction from prediction markets reflected that — Kalshi and Polymarket both flagged the clear as a notable capability milestone.
Why this doesn't mean CAPTCHAs are broken
It's worth being precise about what got tested. Neal Agarwal's game presents its puzzles entirely through the visible page — there's no hidden signal layer to defeat, because there isn't one to begin with; it's a demo, not a deployed defense. Production CAPTCHA systems like reCAPTCHA layer in signals a browser game can't replicate: device fingerprinting, mouse-movement and typing-cadence analysis, IP reputation, and account history accumulated over time. A model clearing a visible-puzzle-only demo is a real result, but it's not the same claim as "AI has defeated CAPTCHA," and several replies on the announcement thread made exactly this joke — that sites will soon need to verify whether users pay rather than whether they're human, since the visible-puzzle layer alone is no longer a reliable filter.
Why "cleared a human test" isn't an AGI claim
The most useful pushback in the reaction thread came from two credentialed voices, for different reasons:
- François Chollet, creator of ARC-AGI, argued the AGI conversation has always centered on invention — "it will cure cancer," "it will solve fusion energy" — and that a model shouldn't be called AGI until it demonstrates conceptual breakthroughs and genuinely novel real-world technology, not completion of an existing test someone else designed.
- Jürgen Schmidhuber argued separately that no AGI claim holds without mastery of the real physical world, distinguishing between software's existing capacity for self-improving meta-learning and the harder, unmet bar of self-improving hardware.
Both objections land on the same point from different angles: clearing a well-defined test — even an escalating, dynamic one like a 48-level CAPTCHA gauntlet — demonstrates strong, specific capability. It does not demonstrate the open-ended, generative capability the term AGI is actually meant to describe.
Related on explainx.ai
- GPT-6 Astra is live: every number that actually matters
- GPT-6 Astra: a SimpleBench win and a reasoning-monitor evasion problem
- OpenAI Codex computer-use: Windows and mobile control
- How to read AI benchmarks (without getting fooled)
- Jensen Huang: "AGI has arrived" with GPT-6 Astra
- GPT-6 Astra vs. Claude Fable 5.1 comparison
Sources
- Sharif Shameem on X — CAPTCHA gauntlet video, September 7-8, 2026
- Kalshi on X and Polymarket Money on X, September 8, 2026
- François Chollet on X, September 8, 2026
- Jürgen Schmidhuber on X, September 8, 2026
- Neal Agarwal's "I'm Not a Robot" — the original 2025 browser game
Benchmark figures and the CAPTCHA-clear claim reflect reporting and reactions circulating September 7-8, 2026. MazeBench methodology and the exact scoring for both models were not independently verified against a primary leaderboard at publication time.
