explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR
  • What actually changed in v4.2
  • The bigger change: private test sets now carry 40% of the score
  • The actual leaderboard: who's on top now
  • What this means if you're choosing between Fable 5.1 and Astra
  • What Hacker News actually pushed back on
  • Honest limitations
  • Related on explainx.ai
← Back to blog

explainx / blog

Artificial Analysis Intelligence Index v4.2: What Actually Changed

Artificial Analysis, Intelligence Index, Benchmarks, Claude Fable 5.1, GPT-6 Astra, AI Evaluation

Artificial Analysis shipped Index v4.2 on September 4, 2026 — new held-out evals, GPQA Diamond dropped, 40% private-test weighting. Fable 5.1 leads, Astra wins on cost and tokens.

Sep 5, 2026·12 min read·Yash Thakker
add explainx.ai
go deep
Artificial Analysis Intelligence Index v4.2: What Actually Changed

Artificial Analysis doesn't usually revise its flagship model-ranking score mid-cycle. On September 4, 2026 — eight months after Index v4 shipped in January and roughly a month after v4.1 — it did anyway, publishing Intelligence Index v4.2 as an explicit interim step ahead of a planned v5 release. The changes are substantive: two new evaluations added, one long-standing benchmark retired for being too easy, and a doubling of how much the score leans on test questions no model has ever seen.

If you've been using the Intelligence Index to decide which frontier model to build on — and a lot of builders have — the numbers under the hood just moved. Here's exactly what changed, what the new leaderboard says, and the methodology questions Hacker News raised that are worth taking seriously rather than waving away.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

table · 2 cols
QuestionAnswer
What is this?Artificial Analysis Intelligence Index v4.2, published September 4, 2026 — an interim update ahead of v5
What got added?AA-Briefcase (agentic knowledge work, private test set) and GDP.pdf (long-document reasoning, built by Surge AI)
What got removed?GPQA Diamond — retired for being "saturated," no longer discriminating between top models
Biggest structural changePrivate/held-out test sets now make up 40% of total index weight, double v4.1's figure
Who leads overall?Claude Fable 5.1, followed by GPT-6 Astra (a 4-point gain over GPT-5.6 Sol)
Who wins on cost and tokens?GPT-6 Astra dominates token efficiency; the cost-per-task Pareto frontier is now shared by Anthropic, OpenAI, Meta, and Z.AI
Who leads GDP.pdf specifically?OpenAI — GPT-6 Astra scores 33.2% All-pass Rate, ahead of GPT-5.6 Sol (28.2%) and Claude Fable 5.1 (26.2%)
Is the update controversial?Somewhat — Hacker News raised fair timing and transparency questions, covered below

What actually changed in v4.2

Artificial Analysis's own methodology page lays out the current structure plainly: the index now combines ten evaluations across four weighted categories — Agents (30%), Coding (20%), General (30%), and Scientific Reasoning (20%).

table · 3 cols
CategoryWeightEvaluations
Agents30%AA-Briefcase (15%), GDPval-AA v2 (10%), 𝜏³-Banking (5%)
Coding20%Terminal-Bench v2.1 (10%), SciCode (10%)
General30%AA-Omniscience (15%), GDP.pdf (10%), AA-LCR v1.1 (5%)
Scientific Reasoning20%Humanity's Last Exam (10%), CritPt (10%)

Three things happened to get here from v4.1:

AA-Briefcase was added, an in-house agentic knowledge-work evaluation built by industry experts, run against a private, held-out test set. It tests models on realistic multi-week knowledge-work projects — many linked tasks, thousands of input source files — graded with a rubric-plus-pairwise system for verifiable task success, analytical quality, and presentation quality. It's the single largest weight in the entire index at 15%.

GDP.pdf was added, built by Surge AI, evaluating single-turn professional document reasoning across 100 PDFs spanning 4,592 pages and ten domains — text, tables, charts, footnotes, exclusions. Answers are graded against 1,275 expert-authored atomic criteria, and the headline "All-pass Rate" metric only credits a task if the model satisfies every criterion, not most of them — a deliberately strict bar for a document-reasoning eval.

GPQA Diamond was removed, transitioned to legacy status. The reason is saturation: frontier models now ace it consistently enough that it no longer separates the field. Artificial Analysis also improved grading infrastructure elsewhere — AA-LCR v1.1 got a new grading system prompt and corrected answer-key errors, GDPval-AA v2 improved sampling and re-anchored its Elo scale for stability as new models are added, and SciCode's grading sandboxes were hardened so slow-but-correct code stops getting marked as a failure.

The bigger change: private test sets now carry 40% of the score

The structural headline isn't any single evaluation — it's the weighting shift. Private, held-out test sets — AA-Briefcase, AA-Omniscience, and CritPt among them — now make up 40% of the total Intelligence Index weight, double what v4.1 assigned. Artificial Analysis says this share will increase further in v5.

This matters because it's the industry's actual mechanism for fighting a well-documented problem: public benchmark contamination. A public benchmark's questions and often its answers eventually end up scraped into training data, whether deliberately or by accident, and once that happens a high score stops proving reasoning and starts proving memorization. explainx.ai has covered the receipts on this directly — GSM1k found up to a 13% accuracy drop on fresh, unseen math problems compared to the contaminated GSM8K, and a controlled experiment on LMArena caught two identical model checkpoints diverging by 17 points under different names, purely from best-of-N submission gaming. Google DeepMind's response to the same underlying problem was to run the first double-blind AI evaluation, testing a model inside a cryptographic enclave so neither side could see the other's material.

A held-out test set attacks the same problem from a different angle: if a question was never published, it can't leak into training data, full stop. That's a real, mechanism-level defense — not a promise, a policy. The tradeoff, as explainx.ai's benchmark-contamination coverage has noted before, is that private sets aren't a permanent fix either; they still "rot" over time as details leak through repeated testing, just more slowly than a fully public leaderboard. Raising the held-out share to 40%, with more planned for v5, is Artificial Analysis explicitly betting that a bigger private component slows that rot rate enough to matter.

The actual leaderboard: who's on top now

Under v4.2, Anthropic's Claude Fable 5.1 leads the overall Index, followed by OpenAI's GPT-6 Astra, which posted a 4-point gain over its predecessor GPT-5.6 Sol. Meta ranks third among labs, followed by SpaceXAI, Moonshot/Kimi, Z.AI, and Google.

That single ranking is not the whole story, and treating it as one is exactly the mistake explainx.ai's guide to reading AI benchmarks warns against — an aggregate score flattens real differences that show up the moment you look at sub-scores.

On cost: the cost-per-task Pareto frontier — the set of models that no other model beats on both price and capability simultaneously — is now shared by four labs: Anthropic, OpenAI, Meta, and Z.AI. That's a genuinely competitive frontier, not a single-lab moat.

On token efficiency: GPT-6 Astra dominates the output-token-efficiency frontier, using meaningfully fewer output tokens than almost every other model near the intelligence frontier to reach a comparable score. Among models scoring 25+ on the Index, Claude Fable 5.1, Grok 4.5, and Gemini 3.5 Flash-Lite sit at either end of that efficiency curve.

On AA-Briefcase specifically: Claude Fable 5.1 and Opus 5 lead, followed by GPT-6 Astra and Muse Spark 1.3. GPT-6 Astra shows roughly an 85 Elo-point gain over GPT-5.6 Sol on this eval alone — a much larger jump than the 4-point overall Index gain, because AA-Briefcase specifically rewards the kind of sustained, multi-task agentic work Astra was built to be good at.

On GDP.pdf specifically: OpenAI leads outright. GPT-6 Astra scores 33.2% All-pass Rate, GPT-5.6 Sol scores 28.2%, and Claude Fable 5.1 scores 26.2%. On this eval, the overall Index leader is not the category leader.

table · 3 cols
MetricLeaderDetail
Overall Intelligence IndexClaude Fable 5.1GPT-6 Astra second, +4 pts over GPT-5.6 Sol
Cost-per-task Pareto frontierShared: Anthropic, OpenAI, Meta, Z.AINo single-lab moat
Output token efficiencyGPT-6 AstraMost efficient among 25+-scoring models
AA-Briefcase (agentic knowledge work)Claude Fable 5.1 / Opus 5GPT-6 Astra third, ~85 Elo gain over Sol
GDP.pdf (long-document reasoning)GPT-6 Astra33.2% All-pass vs Fable 5.1's 26.2%

What this means if you're choosing between Fable 5.1 and Astra

This is the part that actually changes a decision this week, and it's exactly the comparison explainx.ai's own GPT-6 Astra launch benchmarks and pricing coverage flagged as unresolved: at launch, Astra scored 61 on the older Intelligence Index reading, trailing Fable 5.1 on general reasoning while leading decisively on security and long-context work. v4.2's numbers sharpen that picture rather than settling it.

  • If your workload is broad agentic knowledge work — synthesizing many linked tasks over a long horizon — Fable 5.1's AA-Briefcase lead is the more relevant number than the overall Index score.
  • If your workload is long-document reasoning — contracts, financial filings, technical PDFs with tables and footnotes — Astra's GDP.pdf lead (33.2% vs 26.2%) is a genuinely large gap on a strict, all-or-nothing grading standard.
  • If cost and token usage matter more than squeezing out the last few points of intelligence, Astra's token-efficiency lead and the four-lab cost-per-task frontier mean you're no longer choosing based on price alone — Anthropic, OpenAI, Meta, and Z.AI models can all sit on the efficient frontier depending on the specific task.
  • There is no single winner here, and that's the honest read. Fable 5.1 leading the aggregate Index while Astra leads a specific document-reasoning eval and the token-efficiency curve is a genuinely mixed result. Model choice depends on the workload, not on which model tops one composite number — see explainx.ai's Claude Fable 5.1 and Mythos 5.1 benchmark coverage for the fuller Fable-side numbers this update sits alongside.

What Hacker News actually pushed back on

The discussion thread on this announcement stayed small — 24 points — but the pointed comments are worth engaging with directly rather than smoothing over, in the same spirit as explainx.ai's benchmark-claims fact-check coverage.

User "redox99" raised the sharpest methodological concern: "They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this." This is a fair transparency worry, not a conspiracy theory. Benchmark providers who revise methodology at the exact moment it would otherwise contradict prevailing consensus can't fully prove they weren't motivated by that consensus — the timing alone invites the question, and Artificial Analysis's post doesn't pre-empt it.

But the other side of that tension is also real and defensible. Removing a saturated evaluation — GPQA Diamond, which nearly every frontier model now aces — is standard benchmark hygiene, the same maintenance cycle that's already retired or downweighted MMLU and HumanEval across the industry as detailed in explainx.ai's complete guide to AI benchmarks. Adding harder, private evaluations and increasing held-out weighting is the direct, defensible response to the contamination evidence covered above. Neither of those changes requires assuming bad faith about GPT-6 Astra specifically. Both readings can be true at once — the update is plausibly correct on the merits and plausibly convenient in its timing — and Artificial Analysis's public writeup doesn't resolve which one dominates. That's the honest state of the disagreement, not a verdict either way.

User "nthypes" asked a concrete, checkable question: which index version the published intelligence-vs-cost graph actually uses, noting not all models had been re-run under v4.2 at publication time. That's a fair point about apples-to-apples comparison — a chart mixing v4.1-era and v4.2-era scores without a clear label would understate or overstate gaps depending on which models got re-scored first.

User "6thbit" asked why ARC-AGI-3 isn't in the Index at all, suggesting its inclusion would reorder the rankings. It's a reasonable question given ARC-AGI-3's growing prominence this year — see explainx.ai's coverage of Opus 5's ARC-AGI-3 leaderboard run — though Artificial Analysis's ten-evaluation set is already broad, and every added eval is a tradeoff against index complexity and per-model evaluation cost.

User "lousken" asked the practical transparency question: how to actually view the previous version to compare v4.1 against v4.2 side by side. That's a legitimate ask for any versioned benchmark — a changelog with both scores visible would let readers judge the redox99 critique for themselves instead of taking either side's framing on faith.

None of these comments individually invalidate v4.2. Together, they're a reasonable list of what a more transparent update would have shipped with: a changelog, an explicit note on which models were re-scored versus carried over, and an acknowledgment that the timing invites scrutiny even where the underlying methodology change is sound.

Honest limitations

  • All scores and rankings here are Artificial Analysis's own self-reported figures as published in its v4.2 announcement and methodology page — explainx.ai has not independently re-run these evaluations.
  • The HN thread was small (24 points) — it's a useful sample of informed skepticism, not a comprehensive survey of expert opinion on the methodology change.
  • The redox99 timing critique cannot be fully proven or disproven from the outside. Both the "convenient timing" reading and the "normal benchmark maintenance" reading are plausible and are not mutually exclusive.
  • "Which models were re-scored under v4.2" is genuinely unclear from the public materials, per nthypes's question — treat any cross-version comparison chart with that caveat in mind until Artificial Analysis clarifies.
  • Index composition will keep changing. Artificial Analysis has said held-out weighting will rise further in v5, so v4.2's 40% figure is itself a snapshot, not a final structure.

Related on explainx.ai

  • MAI-Image-2.6-Flash: Microsoft's faster, cheaper image model
  • GPT-6 Astra Launch: Every Benchmark and Pricing Number
  • Claude Fable 5.1 and Mythos 5.1: Benchmarks, Pricing, Safeguards
  • GLM-5.3 Ties Kimi K3 on the AA Intelligence Index — Without a New Base Model
  • Goodhart's Law Comes for Every Benchmark You Trust: The 2026 Receipts
  • Google DeepMind's First Double-Blind AI Evaluation
  • How to Read an AI Benchmark and Not Get Fooled
  • AI Benchmarks in 2026: The Complete Guide
  • The 2026 Model-Launch Benchmark Fact-Check
  • ARC-AGI-3: Opus 5's Leaderboard Run

See also the Artificial Analysis Intelligence Index dictionary entry for a quick definition, and Artificial Analysis's own methodology page for the full evaluation and weighting breakdown.


Scores, weightings, and rankings reflect Artificial Analysis's Intelligence Index v4.2 as published September 4, 2026, and its public methodology page as fetched on the publication date of this post. Index composition and per-model scores may change again before v5 ships — check Artificial Analysis's own site for the current live ranking before making a model-selection decision.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Jul 28, 2026

Opus 5 on SlopCodeBench: 24% Strict Pass, Still Can't Run Lights-Off

SlopCodeBench measures whether a model can maintain a codebase across incrementally revealed checkpoints, not just solve one problem once. humanlayer's dhorthy ran Claude Opus 5, Opus 4.8, and Sonnet 5 through 17 checkpoints and watched live for six hours. Opus 5 won on strict pass rate but also tripled the code volume — this post breaks down what the numbers mean and what HN argued about.

Sep 5, 2026

EEBench: The Benchmark That Grades Whether AI Can Design Circuits

On September 4, 2026, the atopile team published EEBench — a benchmark that grades AI-designed electronic circuits by simulating them in SPICE with real manufacturer part tolerances, not just checking whether the design compiles. It hit #1 on Hacker News, and the leaderboard has some surprises.

Sep 5, 2026

Rethinking Skills and AGENTS.md for GPT-6 Astra: A Practical Guide

Eric Provencher of OpenAI's Codex DX team argues that most Skills, AGENTS.md files, and task prompts written for older models actively hurt GPT-6 Astra — bloated descriptions, unnecessary permission-seeking, and unclear stopping points. explainx.ai breaks his guidance into four actionable checklists, with copy-paste before/after examples, and maps each one to the Claude Code equivalent.