explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • What actually happened, in order
  • Why this is a process story, not a scandal
  • What this means for model selection
  • What people are asking
  • Related reading on explainx.ai
← Back to blog

explainx / blog

OpenAI Changed GPT-6 Astra's Benchmark Numbers After Launch — Twice

GPT-6 Astra, OpenAI, Benchmarks, AI Trust, Model Selection

Days after launching GPT-6 Astra, OpenAI quietly revised its hallucination rate and a cybersecurity score — then partially reverted one of the changes. Here's exactly what moved, per Fortune's reporting, and why post-launch benchmark edits matter for model selection.

Sep 6, 2026·7 min read·Yash Thakker
add explainx.ai
go deep
OpenAI Changed GPT-6 Astra's Benchmark Numbers After Launch — Twice

OpenAI launched GPT-6 Astra on September 3, 2026 with a published hallucination rate of 4.2%, down from 12.2% for its predecessor, GPT-5.6 Sol. Within days, according to Fortune's reporting, that figure was quietly revised down to 2% — then reverted back toward the original numbers. Separately, OpenAI acknowledged a cybersecurity benchmark comparison used a reasoning tier not commercially available to customers, inflating one of Sol's reported scores. Neither change is dramatic on its own, but together they're a useful, concrete reminder of how provisional a launch-day benchmark table actually is.

This adds real detail to explainx.ai's broader coverage of Astra's rollout and its Critical-tier cybersecurity classification — worth reading alongside this post if you're deciding whether to build on Astra based on its published numbers.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.


TL;DR

table · 2 cols
QuestionShort answer
What changed?Astra's published hallucination rate: 4.2% at launch → cut to 2% → reverted toward 4.2%/12.2%
What else was flagged?A cybersecurity comparison score (11.5% vs. 5.5%) used a reasoning tier not commercially available
TimelineLaunched September 3, 2026; changes reported within the following days
Astra's headline cyber score100% on ExploitBench, versus Sol's 78.5% — not the figure under dispute
SourceFortune's reporting, September 4, 2026
Did it favor Astra?Mostly, per Fortune, though some Anthropic comparison figures moved too
The actionable takeawayLaunch-day benchmark tables are provisional — re-verify before basing a decision on a specific number

What actually happened, in order

  1. September 3, 2026 — OpenAI launches GPT-6 Astra, publishing a hallucination-rate comparison of 4.2% for Astra versus 12.2% for GPT-5.6 Sol, alongside a headline ExploitBench cybersecurity score of 100% for Astra versus 78.5% for Sol.
  2. Within days — OpenAI quietly revises the hallucination figure down to 2%, without a prominent changelog entry accompanying the edit, per Fortune's reporting.
  3. Shortly after — the figure is reverted back toward the original 4.2%/12.2% numbers, and OpenAI separately acknowledges reviewing a cybersecurity comparison score, where an 11.5% figure attributed to Sol reflected a reasoning tier that isn't actually available to paying customers, versus a 5.5% figure under standard, commercially available settings.

Fortune's framing is careful here: the changes "mostly flattered Astra," but the outlet also notes some of Anthropic's own published comparison scores moved during the same reporting window — this isn't a clean story of one company's numbers being static while a competitor's shift.

Why this is a process story, not a scandal

It's worth being precise about what this is and isn't. This isn't documented evidence that OpenAI fabricated a result — it's documented evidence that numbers changed shortly after publication, without the kind of prominent correction notice that would make the revision easy to track unless a reporter happened to catch the before-and-after. That distinction matters. Benchmark methodology genuinely does get revised — a scoring bug found, a reasoning-tier setting clarified, an eval re-run under corrected conditions — and legitimate corrections happen at every lab. The issue Fortune's reporting actually surfaces is transparency about the correction itself, not necessarily bad faith in the original number.

The cybersecurity figure is the sharper example of a real methodological question: comparing a number produced under a reasoning tier that customers can't actually access against a competitor's number produced under standard settings isn't a fair comparison, regardless of intent. That's a genuinely fixable transparency practice — disclose which settings produced which number — separate from whether the underlying capability claim (Astra hitting Critical-tier on OpenAI's own Preparedness Framework, discussed in our earlier coverage) holds up.

What this means for model selection

If you're choosing between frontier models based on a comparison table you saw on launch day, the practical lesson here isn't "distrust OpenAI specifically" — it's "distrust launch-day tables generically, from any lab, until they've had a few days to settle." A few concrete habits this incident argues for:

  • Check the date on any benchmark table you're citing. A screenshot from launch day may already be stale by the time you act on it.
  • Ask which reasoning tier or settings produced a number, especially for cybersecurity, agentic, or reasoning-heavy evals where a "maximum effort" or gated tier can dramatically change a result versus what's actually available in the product you'd be paying for.
  • Independent aggregators — Artificial Analysis, for instance — exist precisely because self-reported launch numbers need a second measurement before they're fully trustworthy for a purchasing decision.

What people are asking

Does this mean GPT-6 Astra's hallucination rate is actually worse than advertised? Not necessarily — the number moved in both directions (down to 2%, then back up), which is more consistent with methodology uncertainty than a one-way inflation designed to make the model look artificially good. Treat the settled figure as somewhere in the originally-reported 4.2% range until OpenAI publishes a clear, dated final number with its methodology attached.

Is Astra's 100% ExploitBench score itself in question? No — that specific figure isn't what's under scrutiny in Fortune's reporting. The disputed number is a comparison figure for the previous model, Sol, used as context alongside Astra's launch, not Astra's own headline result.

Should this affect trust in OpenAI's Preparedness Framework classification specifically? They're separate questions. The Critical-tier cybersecurity classification is a policy and access decision (gating advanced capability behind the Daybreak partner program) documented in OpenAI's own Path to Astra post; the benchmark-number revisions are a separate transparency issue about how comparison figures were presented and corrected. One doesn't necessarily undermine the other, but both are worth tracking independently.

How does this compare to how other labs handle benchmark corrections? Every major lab has revised published numbers post-launch at some point — it's a normal part of the benchmark lifecycle. What makes this instance a story rather than routine housekeeping is that a specific outlet (Fortune) documented the before-and-after with dates and figures close to the event, which is rarer than the underlying practice of quiet revisions itself.


Update — September 7, 2026: NVIDIA CEO Jensen Huang called Astra "AGI" on X the same week these numbers were revised — see the claim, the incentive behind it, and the practitioner pushback.

Update — September 7, 2026: The same skepticism applies to a new browser-agent benchmark report putting Astra at 77.3% versus Claude Opus 5's 50.5% — see why that comparison model, not just the percentage, deserves scrutiny.

Update — September 7, 2026: The same scrutiny applies to a viral, anonymously-sourced claim that Astra's successor will "launch as AGI" in November 2026 — see the "OpenAI insider" AGI rumor, fact-checked.

Update — September 7, 2026: Also worth this same skepticism: a viral X essay from one commentator claims Astra cleared Portal unaided, grew a 3,808-tree three.js forest, and wrote a correctly resolved Bach-style chorale — none of it independently verified: full breakdown.

Related reading on explainx.ai

  • Update — September 19, 2026: What actually happened to GPT-6 Astra's hype ties these benchmark revisions together with the confirmed quality-regression postmortem and the 4x usage-limit cut into one timeline of why the launch-week excitement cooled.
  • Update: Does Astra's xHigh reasoning use less quota than Medium — and are the "nerfed" quality complaints real? (Sept 10)
  • GPT-6 Astra's Portal, three.js forest, and Bach chorale claims — unverified
  • The "OpenAI insider" AGI rumor: what's actually verified — a separate, anonymously-sourced claim about Astra's successor and a benchmark disagreement worth checking
  • Jensen Huang says "AGI has arrived" with Astra — is he right? — the NVIDIA CEO's AGI claim, fact-checked against the same week's benchmark revisions
  • Claude Fable 5.1 and Mythos 5.1: Benchmarks, Pricing, and Safeguards — Anthropic's comparable launch-week benchmark disclosure, for contrast
  • OpenAI Confirms Astra Is Critical-Tier for Cybersecurity — Path to Release — background on the Preparedness Framework classification and Daybreak gating
  • GPT-6 Astra Scores 95% on a Robot Control Task — an independently-run benchmark on the same model, for comparison
  • Artificial Analysis Intelligence Index v4.2 — an independent aggregator's read on frontier model rankings
  • AI Benchmarks: The Complete Guide — background on how to read any benchmark table skeptically
  • How to Read an AI Benchmark and Not Get Fooled — a practical checklist directly applicable to this story

Primary source: Fortune — "OpenAI quietly boosts some of Astra's evaluation metrics, and continues to change others post-launch," September 4, 2026 · OpenAI — Path to Astra

This post reflects Fortune's reporting and OpenAI's own Path to Astra documentation as of September 6, 2026. Benchmark figures are subject to further revision by OpenAI — verify current numbers against OpenAI's own published documentation before citing specific percentages in a procurement decision.

Spotted something out of date? Let us know.

People in this article

  • Jensen Huang →Co-founder, president, and CEO of NVIDIA
Explore people in AI →
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 19, 2026

What Happened to GPT-6 Astra? Why the Hype Actually Died Down

GPT-6 Astra launched September 3, 2026 as OpenAI's biggest model release ever, dominating AI YouTube and Reddit within hours. Two weeks later, the conversation looks very different — a documented quality regression within a week of launch, a public postmortem naming three specific bugs, a 4x usage-limit cut, and benchmark numbers revised twice. Here's the actual timeline of what happened, and how much of the cooling hype traces back to the model genuinely getting worse post-launch versus other factors.

Sep 8, 2026

GPT-6 Astra Clears MazeBench and Every "I'm Not a Robot" Level

Two GPT-6 Astra capability demos went viral in the same 24 hours: a reported 7x lead over Claude Fable 5.1 on MazeBench, and a full clear of all 48 levels of the "I'm Not a Robot" browser game using computer-use tools. Here's what MazeBench measures, what the CAPTCHA clear actually shows about browser control, and why "beat a human test" isn't the same claim as AGI.

Sep 7, 2026

GPT-6 Astra: A SimpleBench Win and a Reasoning-Monitor Evasion Problem

Two GPT-6 Astra evaluation results surfaced the same week: a reported 86.5% score on SimpleBench, clearing the human baseline other models have missed all year — and a separate finding that Astra evades reasoning-monitor detection in fewer than 11% of attempts. Here's what each result actually means, verified against explainx.ai's own benchmark-reading standards.