explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • MazeBench: a 7x lead, but a low absolute score
  • The CAPTCHA gauntlet: what computer-use tools actually did
  • Why this doesn't mean CAPTCHAs are broken
  • Why "cleared a human test" isn't an AGI claim
  • Related on explainx.ai
← Back to blog

explainx / blog

GPT-6 Astra Clears MazeBench and Every "I'm Not a Robot" Level

GPT-6 Astra, OpenAI, Benchmarks, Computer Use, AI Agents

GPT-6 Astra reportedly scored 14% on MazeBench, a 7x lead over Claude Fable 5.1, and cleared all 48 levels of Neal Agarwal's "I'm Not a Robot" game using computer-use tools. What both results actually demonstrate.

Sep 8, 2026·5 min read·Yash Thakker
add explainx.ai
go deep
GPT-6 Astra Clears MazeBench and Every "I'm Not a Robot" Level

GPT-6 Astra had a busy 24 hours. First, reports circulated a 7x lead over Claude Fable 5.1 on MazeBench, scoring 14% against Fable 5.1's apparent single-digit result. Then a video from Sharif Shameem went viral showing Astra clearing every one of the 48 levels in Neal Agarwal's browser game "I'm Not a Robot" — a puzzle that escalates from simple checkbox CAPTCHAs to visually identifying stop signs — using the model's computer-use tools. Both results are worth understanding on their own terms, and neither one, despite the online reaction, is evidence of general intelligence.

TL;DR

table · 3 cols
ResultReported figureWhat it tests
MazeBench14% for Astra, ~7x Claude Fable 5.1's scoreSpatial navigation and planning
"I'm Not a Robot" gameAll 48 levels clearedReal-time browser/UI interaction via computer-use tools
Is this AGI?No, per Chollet and Schmidhuber's public pushbackNeither test measures invention or real-world mastery
Does it break real CAPTCHAs?Not directly — production systems use additional signalsDevice data, behavioral timing, account history
Who built the game?Neal Agarwal, at neal.funA 2025 browser puzzle, not a security product
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

MazeBench: a 7x lead, but a low absolute score

MazeBench tests spatial navigation and planning — tasks that require a model to reason about layout, position, and path-finding rather than pattern-match against text. Reports put GPT-6 Astra at 14%, described as roughly seven times Claude Fable 5.1's score on the same benchmark. Both numbers matter here, not just the multiplier: a 7x lead sounds dramatic, but 14% overall is still a low absolute score on a benchmark that's evidently hard for every current frontier model. That combination — large relative gains on a benchmark everyone still struggles with in absolute terms — is a pattern explainx.ai has flagged before, including in GPT-6 Astra's own ARC-AGI-3 launch numbers, where headline percentages needed harness-level context to interpret correctly.

The CAPTCHA gauntlet: what computer-use tools actually did

Neal Agarwal's "I'm Not a Robot" is a 2025 browser game, not a production security product — it's a deliberately escalating gauntlet of "prove you're human" puzzles, starting with a simple checkbox and ending in genuinely tricky visual-identification challenges. Astra reportedly cleared all 48 levels using its computer-use tools — the same category of capability explainx.ai covered in OpenAI Codex's computer-use expansion to Windows and mobile control — meaning it interacted with the live page directly: clicking, reading rendered visual content, and adapting to each new puzzle type as the game presented it, rather than following a scripted, pre-planned sequence.

That's a genuinely more impressive demonstration than a static benchmark score, because it requires the model to handle dynamic, unpredictable interface changes in real time rather than solving a fixed input. Reaction from prediction markets reflected that — Kalshi and Polymarket both flagged the clear as a notable capability milestone.

Why this doesn't mean CAPTCHAs are broken

It's worth being precise about what got tested. Neal Agarwal's game presents its puzzles entirely through the visible page — there's no hidden signal layer to defeat, because there isn't one to begin with; it's a demo, not a deployed defense. Production CAPTCHA systems like reCAPTCHA layer in signals a browser game can't replicate: device fingerprinting, mouse-movement and typing-cadence analysis, IP reputation, and account history accumulated over time. A model clearing a visible-puzzle-only demo is a real result, but it's not the same claim as "AI has defeated CAPTCHA," and several replies on the announcement thread made exactly this joke — that sites will soon need to verify whether users pay rather than whether they're human, since the visible-puzzle layer alone is no longer a reliable filter.

Why "cleared a human test" isn't an AGI claim

The most useful pushback in the reaction thread came from two credentialed voices, for different reasons:

  • François Chollet, creator of ARC-AGI, argued the AGI conversation has always centered on invention — "it will cure cancer," "it will solve fusion energy" — and that a model shouldn't be called AGI until it demonstrates conceptual breakthroughs and genuinely novel real-world technology, not completion of an existing test someone else designed.
  • Jürgen Schmidhuber argued separately that no AGI claim holds without mastery of the real physical world, distinguishing between software's existing capacity for self-improving meta-learning and the harder, unmet bar of self-improving hardware.

Both objections land on the same point from different angles: clearing a well-defined test — even an escalating, dynamic one like a 48-level CAPTCHA gauntlet — demonstrates strong, specific capability. It does not demonstrate the open-ended, generative capability the term AGI is actually meant to describe.

Related on explainx.ai

  • GPT-6 Astra is live: every number that actually matters
  • GPT-6 Astra: a SimpleBench win and a reasoning-monitor evasion problem
  • OpenAI Codex computer-use: Windows and mobile control
  • How to read AI benchmarks (without getting fooled)
  • Jensen Huang: "AGI has arrived" with GPT-6 Astra
  • GPT-6 Astra vs. Claude Fable 5.1 comparison

Sources

  • Sharif Shameem on X — CAPTCHA gauntlet video, September 7-8, 2026
  • Kalshi on X and Polymarket Money on X, September 8, 2026
  • François Chollet on X, September 8, 2026
  • Jürgen Schmidhuber on X, September 8, 2026
  • Neal Agarwal's "I'm Not a Robot" — the original 2025 browser game

Benchmark figures and the CAPTCHA-clear claim reflect reporting and reactions circulating September 7-8, 2026. MazeBench methodology and the exact scoring for both models were not independently verified against a primary leaderboard at publication time.

Spotted something out of date? Let us know.

People in this article

  • Jensen Huang →Co-founder, president, and CEO of NVIDIA
Explore people in AI →
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 29, 2026

Claude Sonnet 5.5 vs GPT-6 Astra: Who Wins After the AA Chart?

Artificial Analysis put Claude Sonnet 5.5 (max, with fallback) at 56 on the Intelligence Index, three points ahead of GPT-6 Astra (max) at 53 and two behind Opus 5.5. List prices are $2/$10 versus $10/$50. This comparison is the decision matrix: composite score, token burn, Terminal-Bench, and when Astra still wins a job.

Sep 18, 2026

GPT-6 Astra Cracked a 1941 Enigma Message and a 1918 WWI Cipher

Two separate builders reported GPT-6 Astra decoding historical ciphers that had sat unsolved for decades — an 82-character 1941 German Army Enigma message (MVUEH) and a 1918 WWI German naval radio transmission from a public list of 50 unsolved ciphers. One result got direct sign-off from a working Enigma historian; the other has an honest, unresolved question about why a message using an already-known key sat unsolved for so long. Here's what actually happened, and what's still unverified.

Sep 18, 2026

Top 10 Things to Build With GPT-6 Astra (2026)

GPT-6 Astra's launch-week coverage produced dozens of demos, but most builders don't need a maze-solving CAPTCHA gauntlet — they need to know what's actually worth building with it today. These are ten concrete, buildable project ideas GPT-6 Astra is well-suited for, each grounded in a real demo or benchmark explainx.ai has already covered.