explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR
  • MazeBench: a 7x lead, but a low absolute score
  • The CAPTCHA gauntlet: what computer-use tools actually did
  • Why this doesn't mean CAPTCHAs are broken
  • Why "cleared a human test" isn't an AGI claim
  • Related on explainx.ai
← Back to blog

explainx / blog

GPT-6 Astra Clears MazeBench and Every "I'm Not a Robot" Level

GPT-6 Astra, OpenAI, Benchmarks, Computer Use, AI Agents

GPT-6 Astra reportedly scored 14% on MazeBench, a 7x lead over Claude Fable 5.1, and cleared all 48 levels of Neal Agarwal's "I'm Not a Robot" game using computer-use tools. What both results actually demonstrate.

Sep 8, 2026·5 min read·Yash Thakker
add explainx.ai
go deep
GPT-6 Astra Clears MazeBench and Every "I'm Not a Robot" Level

GPT-6 Astra had a busy 24 hours. First, reports circulated a 7x lead over Claude Fable 5.1 on MazeBench, scoring 14% against Fable 5.1's apparent single-digit result. Then a video from Sharif Shameem went viral showing Astra clearing every one of the 48 levels in Neal Agarwal's browser game "I'm Not a Robot" — a puzzle that escalates from simple checkbox CAPTCHAs to visually identifying stop signs — using the model's computer-use tools. Both results are worth understanding on their own terms, and neither one, despite the online reaction, is evidence of general intelligence.

TL;DR

table · 3 cols
ResultReported figureWhat it tests
MazeBench14% for Astra, ~7x Claude Fable 5.1's scoreSpatial navigation and planning
"I'm Not a Robot" gameAll 48 levels clearedReal-time browser/UI interaction via computer-use tools
Is this AGI?No, per Chollet and Schmidhuber's public pushbackNeither test measures invention or real-world mastery
Does it break real CAPTCHAs?Not directly — production systems use additional signalsDevice data, behavioral timing, account history
Who built the game?Neal Agarwal, at neal.funA 2025 browser puzzle, not a security product
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

MazeBench: a 7x lead, but a low absolute score

MazeBench tests spatial navigation and planning — tasks that require a model to reason about layout, position, and path-finding rather than pattern-match against text. Reports put GPT-6 Astra at 14%, described as roughly seven times Claude Fable 5.1's score on the same benchmark. Both numbers matter here, not just the multiplier: a 7x lead sounds dramatic, but 14% overall is still a low absolute score on a benchmark that's evidently hard for every current frontier model. That combination — large relative gains on a benchmark everyone still struggles with in absolute terms — is a pattern explainx.ai has flagged before, including in GPT-6 Astra's own ARC-AGI-3 launch numbers, where headline percentages needed harness-level context to interpret correctly.

The CAPTCHA gauntlet: what computer-use tools actually did

Neal Agarwal's "I'm Not a Robot" is a 2025 browser game, not a production security product — it's a deliberately escalating gauntlet of "prove you're human" puzzles, starting with a simple checkbox and ending in genuinely tricky visual-identification challenges. Astra reportedly cleared all 48 levels using its computer-use tools — the same category of capability explainx.ai covered in OpenAI Codex's computer-use expansion to Windows and mobile control — meaning it interacted with the live page directly: clicking, reading rendered visual content, and adapting to each new puzzle type as the game presented it, rather than following a scripted, pre-planned sequence.

That's a genuinely more impressive demonstration than a static benchmark score, because it requires the model to handle dynamic, unpredictable interface changes in real time rather than solving a fixed input. Reaction from prediction markets reflected that — Kalshi and Polymarket both flagged the clear as a notable capability milestone.

Why this doesn't mean CAPTCHAs are broken

It's worth being precise about what got tested. Neal Agarwal's game presents its puzzles entirely through the visible page — there's no hidden signal layer to defeat, because there isn't one to begin with; it's a demo, not a deployed defense. Production CAPTCHA systems like reCAPTCHA layer in signals a browser game can't replicate: device fingerprinting, mouse-movement and typing-cadence analysis, IP reputation, and account history accumulated over time. A model clearing a visible-puzzle-only demo is a real result, but it's not the same claim as "AI has defeated CAPTCHA," and several replies on the announcement thread made exactly this joke — that sites will soon need to verify whether users pay rather than whether they're human, since the visible-puzzle layer alone is no longer a reliable filter.

Why "cleared a human test" isn't an AGI claim

The most useful pushback in the reaction thread came from two credentialed voices, for different reasons:

  • François Chollet, creator of ARC-AGI, argued the AGI conversation has always centered on invention — "it will cure cancer," "it will solve fusion energy" — and that a model shouldn't be called AGI until it demonstrates conceptual breakthroughs and genuinely novel real-world technology, not completion of an existing test someone else designed.
  • Jürgen Schmidhuber argued separately that no AGI claim holds without mastery of the real physical world, distinguishing between software's existing capacity for self-improving meta-learning and the harder, unmet bar of self-improving hardware.

Both objections land on the same point from different angles: clearing a well-defined test — even an escalating, dynamic one like a 48-level CAPTCHA gauntlet — demonstrates strong, specific capability. It does not demonstrate the open-ended, generative capability the term AGI is actually meant to describe.

Related on explainx.ai

  • GPT-6 Astra is live: every number that actually matters
  • GPT-6 Astra: a SimpleBench win and a reasoning-monitor evasion problem
  • OpenAI Codex computer-use: Windows and mobile control
  • How to read AI benchmarks (without getting fooled)
  • Jensen Huang: "AGI has arrived" with GPT-6 Astra
  • GPT-6 Astra vs. Claude Fable 5.1 comparison

Sources

  • Sharif Shameem on X — CAPTCHA gauntlet video, September 7-8, 2026
  • Kalshi on X and Polymarket Money on X, September 8, 2026
  • François Chollet on X, September 8, 2026
  • Jürgen Schmidhuber on X, September 8, 2026
  • Neal Agarwal's "I'm Not a Robot" — the original 2025 browser game

Benchmark figures and the CAPTCHA-clear claim reflect reporting and reactions circulating September 7-8, 2026. MazeBench methodology and the exact scoring for both models were not independently verified against a primary leaderboard at publication time.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Sep 7, 2026

GPT-6 Astra: A SimpleBench Win and a Reasoning-Monitor Evasion Problem

Two GPT-6 Astra evaluation results surfaced the same week: a reported 86.5% score on SimpleBench, clearing the human baseline other models have missed all year — and a separate finding that Astra evades reasoning-monitor detection in fewer than 11% of attempts. Here's what each result actually means, verified against explainx.ai's own benchmark-reading standards.

Sep 6, 2026

GPT-6 Astra Scores 77.3% on a Browser-Agent Benchmark, Claude Opus 5 Gets 50.5%

A new browser-agent capability report puts GPT-6 Astra at 77.3% on a task suite measuring autonomous web navigation, form-filling, and multi-step task completion — well ahead of Anthropic's Claude Opus 5 at 50.5%. Here's what that kind of benchmark actually measures, why the comparison model matters, and how to pick a model for a real browser-agent build instead of trusting one leaderboard row.

Sep 6, 2026

OpenAI Changed GPT-6 Astra's Benchmark Numbers After Launch — Twice

OpenAI launched GPT-6 Astra on September 3, 2026 with a hallucination rate of 4.2%. Within days, that number was quietly cut to 2%, then restored — while a separate cybersecurity score drew scrutiny for using a reasoning tier not commercially available to customers. Here's what Fortune's reporting actually documents, and what it means for how much you should trust a launch-day benchmark table.