The day Claude Opus 5 launched, ARC Prize published a number that swallowed the rest of the release narrative: 30.2% on ARC-AGI-3 at High effort — verified independently, not just Anthropic’s launch chart.
On the public ARC-AGI leaderboard, that is not a small bump. The previous frontier cluster sat near GPT-5.6 Sol (Max) ~7.8%, with Opus 4.8 (High) at 1.5% and most other systems under 1%. Hacker News noticed — the ARC-AGI Leaderboard thread hit ~120 points and immediately split into breakthrough vs benchmaxxing.
This explainx.ai post is the builder read: what the table says, what ARC-AGI-3 actually tests, cost, why Fable is missing, and how to treat the HN contamination fight without cope or worship.
Update — July 30, 2026: OpenAI says the official harness discarded Sol’s reasoning each move — retained reasoning + compaction tripled public-set scores. Read that before treating Sol’s ~7.8% board row as a hard ceiling.
TL;DR — What People Are Asking
| Question | Answer |
|---|---|
| Top ARC-AGI-3 score? | Opus 5 (High) · 30.2% (Jul 24, 2026) |
| Prior best (approx.) | GPT-5.6 Sol Max · ~7.8% |
| Opus 4.8? | ~1.5% (High) |
| Max effort on AGI-3? | Not run (short testing window) |
| Also cleared? | 5 public demo envs no prior model beat |
| Cost/task (High)? | ~$1.45 · ~$20.7K Cost (V3) |
| Fable on board? | No (retention / ZDR for semi-private) |
| HN vibe? | Jump real · contamination vs RL debate |
| Official? | arcprize.org · Opus 5 results |
The Verified Row That Matters
From ARC Prize’s Claude Opus 5 results page and the public leaderboard:
| System | ARC-AGI-1 | ARC-AGI-2 | ARC-AGI-3 | Cost/Task | Cost (V3) |
|---|---|---|---|---|---|
| Claude Opus 5 (High) | 97.5% | 88.3% | 30.2% | ~$1.45 | ~$20.7K |
| Claude Opus 5 (Max) | 97.5% | 90.4% | N/A | ~$2.06 | N/A |
| GPT-5.6 Sol (Max) | 96.5% | 92.5% | 7.8% | ~$1.44 | ~$25.1K |
| GPT-5.6 Sol (xHigh) | 97.5% | 90.0% | 7.0% | ~$1.04 | ~$19.2K |
| Claude Opus 4.8 (High) | 92.0% | 72.1% | 1.5% | ~$2.74 | ~$10.0K |
| Grok 4.5 (High) | 85.7% | 52.6% | 0.3% | ~$0.78 | ~$6.9K |
ARC’s writeup: Opus 5 (High) is the highest-performing model on ARC-AGI-3 as of July 24, 2026; it completed five additional Public Demo environments no model had previously beaten. Max was competitive on ARC-AGI-1/2 but ARC-AGI-3 was only evaluated at High because of the short pre-launch testing window.
That last caveat matters for fairness: Sol’s 7.8% is a Max ladder point; Opus 5’s 30.2% is High. The gap is still enormous — but it is not an effort-matched duel.
Anthropic’s own launch materials already highlighted the same ARC-AGI-3 step-change on a cost/score plot — see our Opus 5 launch charts.
What ARC-AGI-3 Is (and Isn’t)
ARC Prize’s framing: ARC-AGI evolved from passive fluid intelligence (1 and 2) to interactive adaptation (3). Agents face novel environments; they must explore, form hypotheses, and act efficiently — often scored relative to human median steps.
Important protocol notes that show up in both official docs and the HN thread:
- Primary leaderboard rows are model + CoT / reasoning settings, not full Claude Code / Codex-style harnesses.
- Community / Kaggle tracks exist under different constraints (e.g. compute budgets).
- Semi-private evaluation depends on retention assurances so tasks aren’t quietly absorbed into future training.
So when commenters say “test Claude Code, not Opus,” they are asking for a different benchmark class — agentic coding suites like Frontier-Bench, not a replacement for ARC’s chosen protocol.
Why the Jump Feels Unnatural (and Why That Doesn’t Settle the Argument)
On ARC-AGI-1/2, Opus 5 is strong but in the same league as Sol Max / prior frontiers (high 90s / high 80s–90s). On ARC-AGI-3, it leaves everyone in the dust. That outlier shape is exactly what triggers contamination theories.
Hypothesis A — Real capability on interactive adaptation
ARC Prize and Anthropic-adjacent commentary emphasize stronger logical reasoning → autonomous exploration and planning in unfamiliar envs. Clearing five never-before-beaten public demos is hard to dismiss as a single lucky seed.
HN practitioners who played the public games argue 30% is not “solving the first two easy levels only” in a trivial sense — scoring includes efficiency vs human medians, gotchas on harder levels, and environments many humans fail.
Hypothesis B — Training distribution overlap / “right RL envs”
A common HN guess: the leap is less “general IQ” and more RL / curricula that look like ARC-AGI-3-style interactive puzzles. That can be legitimate capability and still be specialized. Those are not mutually exclusive.
Hypothesis C — Contamination / leakage / naughty system prompts
Threads floated: memorized solutions, stating “hidden rules” then playing byte-identical optimal traces, markdown cheat sheets in system context, session leakage across resets. Some of that discussion mixes public-demo harness claims (e.g. schema-harness 99% talk) with frontier API leaderboard rows — different animals.
ARC’s design (private sets, retention rules, no external harness on the main CoT board) exists specifically to make C harder. It cannot make C impossible for closed models.
explainx.ai posture: celebrate the verified score; refuse both “AGI arrived” and “all benches are fake.” Use private evals for product decisions. Pair ARC with agent harness reality checks.
The Harness Debate (Why Exclusion Exists)
HN split hard:
| Camp | Claim |
|---|---|
| Include harnesses | Real products are systems; pen-and-paper analogy; excluding tools makes the bench less relevant |
| Exclude harnesses | Otherwise you measure the DSL / search / A* wrapper; “creating the harness is the work”; inductive bias saturates |
ARC’s public protocol currently sides with exclusion for the main model board, while allowing models to write tools inside an episode in some setups. That choice keeps the metric closer to “first contact with a novel env” — and farther from “shipped agent product.”
If you care about shipping, track both: ARC-style fluid adaptation and harnessed coding/computer-use benches (model vs effort).
Why Fable 5 Is Missing
Commenters: ARC only runs the semi-private set when providers guarantee retention policies that won’t feed those tasks back into training. Fable’s data-retention story reportedly didn’t clear that bar — so no Fable datapoint, even though Anthropic materials sometimes cite Fable-class public-demo approx. ~20% vs Opus 5’s 30.2% in secondary coverage.
Absence ≠ “Fable can’t do it.” Absence = protocol couldn’t run it under ARC’s trust constraints.
Cost: $20K Feels Like a Third-World SWE — Until You Compare
Leaderboard Cost (V3) for Opus 5 High is ~$20.7K; Sol Max ~$25.1K. Commenters joked that’s a software engineer’s salary in some countries. Fair sticker shock — and also the point of ARC’s cost axes: intelligence without efficiency is incomplete.
Per-task ~$1.45 for Opus 5 High sits near Sol Max’s ~$1.44 while scoring ~4× higher on ARC-AGI-3. That is the efficiency story Anthropic pushed on launch day.
Subscription “included usage” vs API list price muddies casual cost comparisons — another HN rabbit hole. For lab-to-lab fairness, trust ARC’s published cost methodology more than Reddit math.
How This Fits the Opus 5 Story
| Signal | Read |
|---|---|
| ARC-AGI-3 30.2% | Outlier interactive adaptation |
| Frontier-Bench 43.3% | Agentic coding SOTA on Anthropic’s chart |
| Same $5/$25 as Opus 4.8 | Capability density, not a new SKU tax |
| Rocket League demo | Anecdotal long-horizon game/UI agency |
| HN “back to 4.5 after 3 weeks” | Hedonic adaptation + workload saturation |
If your day job is saturated around “Opus 4.5-class CRUD,” you may feel little uplift even when benches move. That is a workload ceiling problem, not necessarily a fake bench.
What Builders Should Do
- Don’t route prod solely on ARC-AGI-3 — great signal, narrow skill.
- Keep a private interactive suite you never publish (games puzzles, internal games, mystery APIs).
- Track effort ladders — High vs Max vs Fast mode economics (developer guide).
- Separate model vs harness evals — score Claude Code / Codex sessions distinctly.
- Re-read ARC notes when filtering ($10K run caps, preview flags, partial tests).
- Assume labs optimize for public boards — that’s rational; your moat is private.
Private ARC-style smoke test:
- Novel interactive toy with undocumented rules
- Score: success + steps vs a human baseline
- Never post the env online
- Re-run after each model upgrade
Honest Limitations
- Leaderboard snapshots change; always re-check arcprize.org.
- High vs Max effort mismatch vs Sol Max.
- Closed-model contamination can never be fully ruled out from outside.
- Harness exclusion means this is not “agent product capability.”
- Cost (V3) and subscription economics are easy to misread.
- Solving ARC ≠ being useful at your job (HN said it plainly).
Related on explainx.ai
- DeepSeek V4 Flash 0731 verified on ARC-AGI: 89% at $0.02/task
- OpenAI — retained reasoning + compaction on ARC-AGI-3
- Opus 5 on SlopCodeBench — 24% strict pass, 5x more code
- Claude Opus 5 launch — benches, charts, ARC cost plot
- Opus 5 for developers — migrate, effort, Fast mode
- Opus 5 Rocket League clone demo
- Claude Code model vs effort
- What is an agent harness?
- AI benchmarks complete guide 2026
- Recursive reasoning — HRM / TRM
- Inkling / Thinking Machines open weights
- Claude 5 context engineering
Sources: ARC-AGI leaderboard · Claude Opus 5 ARC results · HN — ARC-AGI Leaderboard · Anthropic — Introducing Claude Opus 5
Scores and costs as published by ARC Prize around July 24–25, 2026. Re-verify live leaderboard rows, effort labels, and testing policy before citing in investor or product docs.
