Artificial Analysis doesn't usually revise its flagship model-ranking score mid-cycle. On September 4, 2026 — eight months after Index v4 shipped in January and roughly a month after v4.1 — it did anyway, publishing Intelligence Index v4.2 as an explicit interim step ahead of a planned v5 release. The changes are substantive: two new evaluations added, one long-standing benchmark retired for being too easy, and a doubling of how much the score leans on test questions no model has ever seen.
If you've been using the Intelligence Index to decide which frontier model to build on — and a lot of builders have — the numbers under the hood just moved. Here's exactly what changed, what the new leaderboard says, and the methodology questions Hacker News raised that are worth taking seriously rather than waving away.
TL;DR
| Question | Answer |
|---|---|
| What is this? | Artificial Analysis Intelligence Index v4.2, published September 4, 2026 — an interim update ahead of v5 |
| What got added? | AA-Briefcase (agentic knowledge work, private test set) and GDP.pdf (long-document reasoning, built by Surge AI) |
| What got removed? | GPQA Diamond — retired for being "saturated," no longer discriminating between top models |
| Biggest structural change | Private/held-out test sets now make up 40% of total index weight, double v4.1's figure |
| Who leads overall? | Claude Fable 5.1, followed by GPT-6 Astra (a 4-point gain over GPT-5.6 Sol) |
| Who wins on cost and tokens? | GPT-6 Astra dominates token efficiency; the cost-per-task Pareto frontier is now shared by Anthropic, OpenAI, Meta, and Z.AI |
| Who leads GDP.pdf specifically? | OpenAI — GPT-6 Astra scores 33.2% All-pass Rate, ahead of GPT-5.6 Sol (28.2%) and Claude Fable 5.1 (26.2%) |
| Is the update controversial? | Somewhat — Hacker News raised fair timing and transparency questions, covered below |
What actually changed in v4.2
Artificial Analysis's own methodology page lays out the current structure plainly: the index now combines ten evaluations across four weighted categories — Agents (30%), Coding (20%), General (30%), and Scientific Reasoning (20%).
| Category | Weight | Evaluations |
|---|---|---|
| Agents | 30% | AA-Briefcase (15%), GDPval-AA v2 (10%), 𝜏³-Banking (5%) |
| Coding | 20% | Terminal-Bench v2.1 (10%), SciCode (10%) |
| General | 30% | AA-Omniscience (15%), GDP.pdf (10%), AA-LCR v1.1 (5%) |
| Scientific Reasoning | 20% | Humanity's Last Exam (10%), CritPt (10%) |
Three things happened to get here from v4.1:
AA-Briefcase was added, an in-house agentic knowledge-work evaluation built by industry experts, run against a private, held-out test set. It tests models on realistic multi-week knowledge-work projects — many linked tasks, thousands of input source files — graded with a rubric-plus-pairwise system for verifiable task success, analytical quality, and presentation quality. It's the single largest weight in the entire index at 15%.
GDP.pdf was added, built by Surge AI, evaluating single-turn professional document reasoning across 100 PDFs spanning 4,592 pages and ten domains — text, tables, charts, footnotes, exclusions. Answers are graded against 1,275 expert-authored atomic criteria, and the headline "All-pass Rate" metric only credits a task if the model satisfies every criterion, not most of them — a deliberately strict bar for a document-reasoning eval.
GPQA Diamond was removed, transitioned to legacy status. The reason is saturation: frontier models now ace it consistently enough that it no longer separates the field. Artificial Analysis also improved grading infrastructure elsewhere — AA-LCR v1.1 got a new grading system prompt and corrected answer-key errors, GDPval-AA v2 improved sampling and re-anchored its Elo scale for stability as new models are added, and SciCode's grading sandboxes were hardened so slow-but-correct code stops getting marked as a failure.
The bigger change: private test sets now carry 40% of the score
The structural headline isn't any single evaluation — it's the weighting shift. Private, held-out test sets — AA-Briefcase, AA-Omniscience, and CritPt among them — now make up 40% of the total Intelligence Index weight, double what v4.1 assigned. Artificial Analysis says this share will increase further in v5.
This matters because it's the industry's actual mechanism for fighting a well-documented problem: public benchmark contamination. A public benchmark's questions and often its answers eventually end up scraped into training data, whether deliberately or by accident, and once that happens a high score stops proving reasoning and starts proving memorization. explainx.ai has covered the receipts on this directly — GSM1k found up to a 13% accuracy drop on fresh, unseen math problems compared to the contaminated GSM8K, and a controlled experiment on LMArena caught two identical model checkpoints diverging by 17 points under different names, purely from best-of-N submission gaming. Google DeepMind's response to the same underlying problem was to run the first double-blind AI evaluation, testing a model inside a cryptographic enclave so neither side could see the other's material.
A held-out test set attacks the same problem from a different angle: if a question was never published, it can't leak into training data, full stop. That's a real, mechanism-level defense — not a promise, a policy. The tradeoff, as explainx.ai's benchmark-contamination coverage has noted before, is that private sets aren't a permanent fix either; they still "rot" over time as details leak through repeated testing, just more slowly than a fully public leaderboard. Raising the held-out share to 40%, with more planned for v5, is Artificial Analysis explicitly betting that a bigger private component slows that rot rate enough to matter.
The actual leaderboard: who's on top now
Under v4.2, Anthropic's Claude Fable 5.1 leads the overall Index, followed by OpenAI's GPT-6 Astra, which posted a 4-point gain over its predecessor GPT-5.6 Sol. Meta ranks third among labs, followed by SpaceXAI, Moonshot/Kimi, Z.AI, and Google.
That single ranking is not the whole story, and treating it as one is exactly the mistake explainx.ai's guide to reading AI benchmarks warns against — an aggregate score flattens real differences that show up the moment you look at sub-scores.
On cost: the cost-per-task Pareto frontier — the set of models that no other model beats on both price and capability simultaneously — is now shared by four labs: Anthropic, OpenAI, Meta, and Z.AI. That's a genuinely competitive frontier, not a single-lab moat.
On token efficiency: GPT-6 Astra dominates the output-token-efficiency frontier, using meaningfully fewer output tokens than almost every other model near the intelligence frontier to reach a comparable score. Among models scoring 25+ on the Index, Claude Fable 5.1, Grok 4.5, and Gemini 3.5 Flash-Lite sit at either end of that efficiency curve.
On AA-Briefcase specifically: Claude Fable 5.1 and Opus 5 lead, followed by GPT-6 Astra and Muse Spark 1.3. GPT-6 Astra shows roughly an 85 Elo-point gain over GPT-5.6 Sol on this eval alone — a much larger jump than the 4-point overall Index gain, because AA-Briefcase specifically rewards the kind of sustained, multi-task agentic work Astra was built to be good at.
On GDP.pdf specifically: OpenAI leads outright. GPT-6 Astra scores 33.2% All-pass Rate, GPT-5.6 Sol scores 28.2%, and Claude Fable 5.1 scores 26.2%. On this eval, the overall Index leader is not the category leader.
| Metric | Leader | Detail |
|---|---|---|
| Overall Intelligence Index | Claude Fable 5.1 | GPT-6 Astra second, +4 pts over GPT-5.6 Sol |
| Cost-per-task Pareto frontier | Shared: Anthropic, OpenAI, Meta, Z.AI | No single-lab moat |
| Output token efficiency | GPT-6 Astra | Most efficient among 25+-scoring models |
| AA-Briefcase (agentic knowledge work) | Claude Fable 5.1 / Opus 5 | GPT-6 Astra third, ~85 Elo gain over Sol |
| GDP.pdf (long-document reasoning) | GPT-6 Astra | 33.2% All-pass vs Fable 5.1's 26.2% |
What this means if you're choosing between Fable 5.1 and Astra
This is the part that actually changes a decision this week, and it's exactly the comparison explainx.ai's own GPT-6 Astra launch benchmarks and pricing coverage flagged as unresolved: at launch, Astra scored 61 on the older Intelligence Index reading, trailing Fable 5.1 on general reasoning while leading decisively on security and long-context work. v4.2's numbers sharpen that picture rather than settling it.
- If your workload is broad agentic knowledge work — synthesizing many linked tasks over a long horizon — Fable 5.1's AA-Briefcase lead is the more relevant number than the overall Index score.
- If your workload is long-document reasoning — contracts, financial filings, technical PDFs with tables and footnotes — Astra's GDP.pdf lead (33.2% vs 26.2%) is a genuinely large gap on a strict, all-or-nothing grading standard.
- If cost and token usage matter more than squeezing out the last few points of intelligence, Astra's token-efficiency lead and the four-lab cost-per-task frontier mean you're no longer choosing based on price alone — Anthropic, OpenAI, Meta, and Z.AI models can all sit on the efficient frontier depending on the specific task.
- There is no single winner here, and that's the honest read. Fable 5.1 leading the aggregate Index while Astra leads a specific document-reasoning eval and the token-efficiency curve is a genuinely mixed result. Model choice depends on the workload, not on which model tops one composite number — see explainx.ai's Claude Fable 5.1 and Mythos 5.1 benchmark coverage for the fuller Fable-side numbers this update sits alongside.
What Hacker News actually pushed back on
The discussion thread on this announcement stayed small — 24 points — but the pointed comments are worth engaging with directly rather than smoothing over, in the same spirit as explainx.ai's benchmark-claims fact-check coverage.
User "redox99" raised the sharpest methodological concern: "They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this." This is a fair transparency worry, not a conspiracy theory. Benchmark providers who revise methodology at the exact moment it would otherwise contradict prevailing consensus can't fully prove they weren't motivated by that consensus — the timing alone invites the question, and Artificial Analysis's post doesn't pre-empt it.
But the other side of that tension is also real and defensible. Removing a saturated evaluation — GPQA Diamond, which nearly every frontier model now aces — is standard benchmark hygiene, the same maintenance cycle that's already retired or downweighted MMLU and HumanEval across the industry as detailed in explainx.ai's complete guide to AI benchmarks. Adding harder, private evaluations and increasing held-out weighting is the direct, defensible response to the contamination evidence covered above. Neither of those changes requires assuming bad faith about GPT-6 Astra specifically. Both readings can be true at once — the update is plausibly correct on the merits and plausibly convenient in its timing — and Artificial Analysis's public writeup doesn't resolve which one dominates. That's the honest state of the disagreement, not a verdict either way.
User "nthypes" asked a concrete, checkable question: which index version the published intelligence-vs-cost graph actually uses, noting not all models had been re-run under v4.2 at publication time. That's a fair point about apples-to-apples comparison — a chart mixing v4.1-era and v4.2-era scores without a clear label would understate or overstate gaps depending on which models got re-scored first.
User "6thbit" asked why ARC-AGI-3 isn't in the Index at all, suggesting its inclusion would reorder the rankings. It's a reasonable question given ARC-AGI-3's growing prominence this year — see explainx.ai's coverage of Opus 5's ARC-AGI-3 leaderboard run — though Artificial Analysis's ten-evaluation set is already broad, and every added eval is a tradeoff against index complexity and per-model evaluation cost.
User "lousken" asked the practical transparency question: how to actually view the previous version to compare v4.1 against v4.2 side by side. That's a legitimate ask for any versioned benchmark — a changelog with both scores visible would let readers judge the redox99 critique for themselves instead of taking either side's framing on faith.
None of these comments individually invalidate v4.2. Together, they're a reasonable list of what a more transparent update would have shipped with: a changelog, an explicit note on which models were re-scored versus carried over, and an acknowledgment that the timing invites scrutiny even where the underlying methodology change is sound.
Honest limitations
- All scores and rankings here are Artificial Analysis's own self-reported figures as published in its v4.2 announcement and methodology page — explainx.ai has not independently re-run these evaluations.
- The HN thread was small (24 points) — it's a useful sample of informed skepticism, not a comprehensive survey of expert opinion on the methodology change.
- The redox99 timing critique cannot be fully proven or disproven from the outside. Both the "convenient timing" reading and the "normal benchmark maintenance" reading are plausible and are not mutually exclusive.
- "Which models were re-scored under v4.2" is genuinely unclear from the public materials, per nthypes's question — treat any cross-version comparison chart with that caveat in mind until Artificial Analysis clarifies.
- Index composition will keep changing. Artificial Analysis has said held-out weighting will rise further in v5, so v4.2's 40% figure is itself a snapshot, not a final structure.
Related on explainx.ai
- MAI-Image-2.6-Flash: Microsoft's faster, cheaper image model
- GPT-6 Astra Launch: Every Benchmark and Pricing Number
- Claude Fable 5.1 and Mythos 5.1: Benchmarks, Pricing, Safeguards
- GLM-5.3 Ties Kimi K3 on the AA Intelligence Index — Without a New Base Model
- Goodhart's Law Comes for Every Benchmark You Trust: The 2026 Receipts
- Google DeepMind's First Double-Blind AI Evaluation
- How to Read an AI Benchmark and Not Get Fooled
- AI Benchmarks in 2026: The Complete Guide
- The 2026 Model-Launch Benchmark Fact-Check
- ARC-AGI-3: Opus 5's Leaderboard Run
See also the Artificial Analysis Intelligence Index dictionary entry for a quick definition, and Artificial Analysis's own methodology page for the full evaluation and weighting breakdown.
Scores, weightings, and rankings reflect Artificial Analysis's Intelligence Index v4.2 as published September 4, 2026, and its public methodology page as fetched on the publication date of this post. Index composition and per-model scores may change again before v5 ships — check Artificial Analysis's own site for the current live ranking before making a model-selection decision.
