explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR — What People Are Asking
  • The Verified Row That Matters
  • What ARC-AGI-3 Is (and Isn’t)
  • Why the Jump Feels Unnatural (and Why That Doesn’t Settle the Argument)
  • The Harness Debate (Why Exclusion Exists)
  • Why Fable 5 Is Missing
  • Cost: $20K Feels Like a Third-World SWE — Until You Compare
  • How This Fits the Opus 5 Story
  • What Builders Should Do
  • Honest Limitations
  • Related on explainx.ai
← Back to blog

explainx / blog

Opus 5 Hits 30.2% on ARC-AGI-3 — What the Jump Means

Claude Opus 5 scores 30.2% on ARC-AGI-3 — ~4× GPT-5.6 Sol Max and ~20× Opus 4.8. ARC Prize verified results, cost/task, Fable’s absence, and the HN contamination debate.

Jul 25, 2026·8 min read·Yash Thakker
ARC-AGI-3Claude Opus 5BenchmarksAnthropicARC Prize
go deep
Opus 5 Hits 30.2% on ARC-AGI-3 — What the Jump Means

The day Claude Opus 5 launched, ARC Prize published a number that swallowed the rest of the release narrative: 30.2% on ARC-AGI-3 at High effort — verified independently, not just Anthropic’s launch chart.

On the public ARC-AGI leaderboard, that is not a small bump. The previous frontier cluster sat near GPT-5.6 Sol (Max) ~7.8%, with Opus 4.8 (High) at 1.5% and most other systems under 1%. Hacker News noticed — the ARC-AGI Leaderboard thread hit ~120 points and immediately split into breakthrough vs benchmaxxing.

This explainx.ai post is the builder read: what the table says, what ARC-AGI-3 actually tests, cost, why Fable is missing, and how to treat the HN contamination fight without cope or worship.

Update — July 30, 2026: OpenAI says the official harness discarded Sol’s reasoning each move — retained reasoning + compaction tripled public-set scores. Read that before treating Sol’s ~7.8% board row as a hard ceiling.

TL;DR — What People Are Asking

table · 2 cols
QuestionAnswer
Top ARC-AGI-3 score?Opus 5 (High) · 30.2% (Jul 24, 2026)
Prior best (approx.)GPT-5.6 Sol Max · ~7.8%
Opus 4.8?~1.5% (High)
Max effort on AGI-3?Not run (short testing window)
Also cleared?5 public demo envs no prior model beat
Cost/task (High)?~$1.45 · ~$20.7K Cost (V3)
Fable on board?No (retention / ZDR for semi-private)
HN vibe?Jump real · contamination vs RL debate
Official?arcprize.org · Opus 5 results
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

The Verified Row That Matters

From ARC Prize’s Claude Opus 5 results page and the public leaderboard:

table · 6 cols
SystemARC-AGI-1ARC-AGI-2ARC-AGI-3Cost/TaskCost (V3)
Claude Opus 5 (High)97.5%88.3%30.2%~$1.45~$20.7K
Claude Opus 5 (Max)97.5%90.4%N/A~$2.06N/A
GPT-5.6 Sol (Max)96.5%92.5%7.8%~$1.44~$25.1K
GPT-5.6 Sol (xHigh)97.5%90.0%7.0%~$1.04~$19.2K
Claude Opus 4.8 (High)92.0%72.1%1.5%~$2.74~$10.0K
Grok 4.5 (High)85.7%52.6%0.3%~$0.78~$6.9K

ARC’s writeup: Opus 5 (High) is the highest-performing model on ARC-AGI-3 as of July 24, 2026; it completed five additional Public Demo environments no model had previously beaten. Max was competitive on ARC-AGI-1/2 but ARC-AGI-3 was only evaluated at High because of the short pre-launch testing window.

That last caveat matters for fairness: Sol’s 7.8% is a Max ladder point; Opus 5’s 30.2% is High. The gap is still enormous — but it is not an effort-matched duel.

Anthropic’s own launch materials already highlighted the same ARC-AGI-3 step-change on a cost/score plot — see our Opus 5 launch charts.

What ARC-AGI-3 Is (and Isn’t)

ARC Prize’s framing: ARC-AGI evolved from passive fluid intelligence (1 and 2) to interactive adaptation (3). Agents face novel environments; they must explore, form hypotheses, and act efficiently — often scored relative to human median steps.

Important protocol notes that show up in both official docs and the HN thread:

  • Primary leaderboard rows are model + CoT / reasoning settings, not full Claude Code / Codex-style harnesses.
  • Community / Kaggle tracks exist under different constraints (e.g. compute budgets).
  • Semi-private evaluation depends on retention assurances so tasks aren’t quietly absorbed into future training.

So when commenters say “test Claude Code, not Opus,” they are asking for a different benchmark class — agentic coding suites like Frontier-Bench, not a replacement for ARC’s chosen protocol.

Why the Jump Feels Unnatural (and Why That Doesn’t Settle the Argument)

On ARC-AGI-1/2, Opus 5 is strong but in the same league as Sol Max / prior frontiers (high 90s / high 80s–90s). On ARC-AGI-3, it leaves everyone in the dust. That outlier shape is exactly what triggers contamination theories.

Hypothesis A — Real capability on interactive adaptation

ARC Prize and Anthropic-adjacent commentary emphasize stronger logical reasoning → autonomous exploration and planning in unfamiliar envs. Clearing five never-before-beaten public demos is hard to dismiss as a single lucky seed.

HN practitioners who played the public games argue 30% is not “solving the first two easy levels only” in a trivial sense — scoring includes efficiency vs human medians, gotchas on harder levels, and environments many humans fail.

Hypothesis B — Training distribution overlap / “right RL envs”

A common HN guess: the leap is less “general IQ” and more RL / curricula that look like ARC-AGI-3-style interactive puzzles. That can be legitimate capability and still be specialized. Those are not mutually exclusive.

Hypothesis C — Contamination / leakage / naughty system prompts

Threads floated: memorized solutions, stating “hidden rules” then playing byte-identical optimal traces, markdown cheat sheets in system context, session leakage across resets. Some of that discussion mixes public-demo harness claims (e.g. schema-harness 99% talk) with frontier API leaderboard rows — different animals.

ARC’s design (private sets, retention rules, no external harness on the main CoT board) exists specifically to make C harder. It cannot make C impossible for closed models.

explainx.ai posture: celebrate the verified score; refuse both “AGI arrived” and “all benches are fake.” Use private evals for product decisions. Pair ARC with agent harness reality checks.

The Harness Debate (Why Exclusion Exists)

HN split hard:

table · 2 cols
CampClaim
Include harnessesReal products are systems; pen-and-paper analogy; excluding tools makes the bench less relevant
Exclude harnessesOtherwise you measure the DSL / search / A* wrapper; “creating the harness is the work”; inductive bias saturates

ARC’s public protocol currently sides with exclusion for the main model board, while allowing models to write tools inside an episode in some setups. That choice keeps the metric closer to “first contact with a novel env” — and farther from “shipped agent product.”

If you care about shipping, track both: ARC-style fluid adaptation and harnessed coding/computer-use benches (model vs effort).

Why Fable 5 Is Missing

Commenters: ARC only runs the semi-private set when providers guarantee retention policies that won’t feed those tasks back into training. Fable’s data-retention story reportedly didn’t clear that bar — so no Fable datapoint, even though Anthropic materials sometimes cite Fable-class public-demo approx. ~20% vs Opus 5’s 30.2% in secondary coverage.

Absence ≠ “Fable can’t do it.” Absence = protocol couldn’t run it under ARC’s trust constraints.

Cost: $20K Feels Like a Third-World SWE — Until You Compare

Leaderboard Cost (V3) for Opus 5 High is ~$20.7K; Sol Max ~$25.1K. Commenters joked that’s a software engineer’s salary in some countries. Fair sticker shock — and also the point of ARC’s cost axes: intelligence without efficiency is incomplete.

Per-task ~$1.45 for Opus 5 High sits near Sol Max’s ~$1.44 while scoring ~4× higher on ARC-AGI-3. That is the efficiency story Anthropic pushed on launch day.

Subscription “included usage” vs API list price muddies casual cost comparisons — another HN rabbit hole. For lab-to-lab fairness, trust ARC’s published cost methodology more than Reddit math.

How This Fits the Opus 5 Story

table · 2 cols
SignalRead
ARC-AGI-3 30.2%Outlier interactive adaptation
Frontier-Bench 43.3%Agentic coding SOTA on Anthropic’s chart
Same $5/$25 as Opus 4.8Capability density, not a new SKU tax
Rocket League demoAnecdotal long-horizon game/UI agency
HN “back to 4.5 after 3 weeks”Hedonic adaptation + workload saturation

If your day job is saturated around “Opus 4.5-class CRUD,” you may feel little uplift even when benches move. That is a workload ceiling problem, not necessarily a fake bench.

What Builders Should Do

  1. Don’t route prod solely on ARC-AGI-3 — great signal, narrow skill.
  2. Keep a private interactive suite you never publish (games puzzles, internal games, mystery APIs).
  3. Track effort ladders — High vs Max vs Fast mode economics (developer guide).
  4. Separate model vs harness evals — score Claude Code / Codex sessions distinctly.
  5. Re-read ARC notes when filtering ($10K run caps, preview flags, partial tests).
  6. Assume labs optimize for public boards — that’s rational; your moat is private.
text
Private ARC-style smoke test:
- Novel interactive toy with undocumented rules
- Score: success + steps vs a human baseline
- Never post the env online
- Re-run after each model upgrade

Honest Limitations

  • Leaderboard snapshots change; always re-check arcprize.org.
  • High vs Max effort mismatch vs Sol Max.
  • Closed-model contamination can never be fully ruled out from outside.
  • Harness exclusion means this is not “agent product capability.”
  • Cost (V3) and subscription economics are easy to misread.
  • Solving ARC ≠ being useful at your job (HN said it plainly).

Related on explainx.ai

  • DeepSeek V4 Flash 0731 verified on ARC-AGI: 89% at $0.02/task
  • OpenAI — retained reasoning + compaction on ARC-AGI-3
  • Opus 5 on SlopCodeBench — 24% strict pass, 5x more code
  • Claude Opus 5 launch — benches, charts, ARC cost plot
  • Opus 5 for developers — migrate, effort, Fast mode
  • Opus 5 Rocket League clone demo
  • Claude Code model vs effort
  • What is an agent harness?
  • AI benchmarks complete guide 2026
  • Recursive reasoning — HRM / TRM
  • Inkling / Thinking Machines open weights
  • Claude 5 context engineering

Sources: ARC-AGI leaderboard · Claude Opus 5 ARC results · HN — ARC-AGI Leaderboard · Anthropic — Introducing Claude Opus 5


Scores and costs as published by ARC Prize around July 24–25, 2026. Re-verify live leaderboard rows, effort labels, and testing policy before citing in investor or product docs.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Jul 28, 2026

Opus 5 on SlopCodeBench: 24% Strict Pass, Still Can't Run Lights-Off

SlopCodeBench measures whether a model can maintain a codebase across incrementally revealed checkpoints, not just solve one problem once. humanlayer's dhorthy ran Claude Opus 5, Opus 4.8, and Sonnet 5 through 17 checkpoints and watched live for six hours. Opus 5 won on strict pass rate but also tripled the code volume — this post breaks down what the numbers mean and what HN argued about.

Jul 25, 2026

Claude Opus 5 Launch: Near Fable 5 at Half the Price

Claude Opus 5 is live — default on Max, strongest on Pro, same price as Opus 4.8. explainx.ai unpacks Anthropic’s benches, effort/cost charts, alignment story, and when to pick Opus 5 vs Fable 5.

Aug 15, 2026

Why Does Claude Opus 5 Feel Worse to Work With? The HN Debate

"Why does Opus 5 feel worse to work with?" hit 778 points and 717 comments on Hacker News this week. The original post's theory: reinforcement learning from verifiable rewards trains models to commit to an answer instead of pausing to ask, and that trade-off shows up as a model that makes bold assumptions instead of checking them. explainx.ai breaks down the thesis, the recurring complaints from the thread, and how to prompt around it in Claude Code.