September 29, 2026 — Artificial Analysis put Claude Sonnet 5.5 at 56 on the Intelligence Index (max effort, default fallback). GPT-6 Astra (max) sits at 53, tied with Claude Fable 5.1 (max) in the same bar chart. Opus 5.5 (max) still leads at 58. The chart people are screenshotting is real. The routing decision is not “Sonnet beats the OpenAI flagship, switch everything.”
This is Sonnet 5.5 vs live GPT-6 Astra, not vs cancelled GPT-6.1 Astra. If your Slack thread still says “wait for 6.1,” that SKU is off the October calendar.
TL;DR — questions people actually ask
| Question | Direct answer |
|---|---|
| Who is ahead on AA Intelligence Index? | Sonnet 5.5 (max + fallback) 56 vs Astra (max) 53. Opus 5.5 is 58. |
| Who is cheaper on the API sticker? | Sonnet 5.5 $2/$10 vs Astra $10/$50. Same sticker as GPT-6 Sol, not the same model. |
| Who is cheaper per hard index task? | Not automatic. AA measured ~$7.60/task for Sonnet 5.5 at max because of output-token volume. Astra uses far fewer tokens per index task. |
| Terminal-Bench 4.0? | AA: Sonnet ~64%, Astra/Opus ~60%. Anthropic table: Sonnet 70.6%. Pick a source. |
| Default in Claude Code? | /model sonnet is Sonnet 5.5 medium from v2.1.284. Default model is still Opus 5.5. See the building guide. |
| Is 6.1 in this comparison? | No. |
The Artificial Analysis chart (what the bars actually are)

Source: Artificial Analysis public leaderboard screenshot, Intelligence Index (labels truncated on the image: “max with fallback” on the Claude 5.5 rows). Index version in AA’s Sonnet 5.5 model card is v4.3.2 — ten evals including AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. That is a newer mix than explainx.ai’s v4.2 write-up from Fable-week.
Read the bars as one composite, not as “IQ.” Fable 5.1 and Astra tie at 53 on this snapshot. Sonnet 5.5 is not a new flagship; it is Anthropic’s $2/$10 tier sitting two points under Opus and three above Astra.
AA’s own article is the part the screenshot leaves out: Sonnet 5.5 (max) used about 193k output tokens per Intelligence Index task, the highest they have measured, roughly 7× GPT-6 Astra (max). Quality at that setting is purchased with decode.
List price vs tokens vs dollars per task
| Claude Sonnet 5.5 | GPT-6 Astra | |
|---|---|---|
| Role | Everyday 5.5 tier (launch guide) | OpenAI flagship (launch numbers) |
| Input / output | $2 / $10 per MTok | $10 / $50 per MTok |
| vs Sonnet 5 | Same sticker; Anthropic claims up to ~30% less per task on typical work because of fewer tokens and speed | N/A |
| Context | 1M native | 1M class (OpenAI launch; some aggregators show ~1.05M) |
| AA Intelligence (max) | 56 (with fallback) | 53 |
| AA effort ladder (Sonnet) | Low 36 → Med 41 → High 47 → Xhigh 52 → Max 56 | Low 46 → Med 50 → High 51 → Xhigh 52 → Max 53 |
| AA output speed (card) | Up to 171 t/s at Xhigh (AA’s fastest Sonnet 5.5 row) | About 61 t/s at max (AA card) |
Five times cheaper on the sticker is the headline finance teams will quote. It is true for input and output rates. It is false as a promise that a max-effort Sonnet job costs one-fifth of a max-effort Astra job.
AA says Sonnet 5.5 at max costs about $7.60 per Intelligence Index task, ~50% more than Sonnet 5 on that same cost-per-task metric, even though the per-token price did not change. The extra quality is extra tokens. Astra’s AA model table lists a $7.7 “price” column next to 53 intelligence — treat AA’s own article ($7.60/task, 193k output tokens) as the Sonnet cost story, and do not smash unlike columns into one spreadsheet.
Practical billing rule:
- High-volume, bounded tasks (bugs, docs, slides, scoped agents) → start Sonnet 5.5 medium/high. That is the job Anthropic staffed the Osmani guide for.
- Max effort as a habit → you are paying Opus-adjacent quality with Sonnet’s rate card but Opus-like decode. Measure $/accepted PR, not $/MTok.
- Astra still makes sense if your harness is Codex / ChatGPT, you need Astra-specific computer-use history from Astra vs Fable, or AA’s science-terminal split still favors OpenAI on your workload.
For spend math on the Anthropic side, keep Opus 5.5 task cost next to this post. Sonnet cache read is $0.20/M (same as Sonnet 5); Astra cache read was $1.00/M in the Fable comparison — cache-heavy loops tilt Anthropic even before the 5× sticker.
Terminal-Bench and the two leaderboards problem
Anthropic’s Sonnet 5.5 launch table (in our building guide) has Terminal-Bench 4.0 at 70.6% for Sonnet vs 66.4% Opus at Xhigh and no Astra cell.
Artificial Analysis’s Sonnet 5.5 article has Terminal-Bench 4.0 at 64% for Sonnet 5.5 vs 60% for Opus 5.5 and GPT-6 Astra. Same name, different numbers.
Do not average them. Vendor evals and AA evals disagree by six points on Sonnet. Both still put Sonnet ≥ Astra on Terminal-Bench 4.0. That is the only claim that survives both tables.
AA also reports Terminal-Bench-Science (not in the Intelligence Index): Sonnet 53%, behind GPT-6 Astra and Opus 5.5. If your agents look like lab workflows in a shell, the composite index overstates Sonnet vs Astra.
FrontierCode / CursorBench on Anthropic’s chart still favor Opus, and Sonnet drops at Max on FrontierCode because of Claude Code review-skill timeouts. That is a Sonnet vs Opus issue, not Astra — but it is why “max everything” is a bad default even when the AA index loves max.
Effort ladders: Sonnet climbs, Astra is already high
AA’s Sonnet 5.5 family on the Intelligence Index: 36 / 41 / 47 / 52 / 56 from Low → Max. Astra’s family: 46 / 50 / 51 / 52 / 53.
Sonnet Low (36) is not in the same league as Astra Low (46). If you “save money” by slamming Sonnet to Low, you can lose the 56-vs-53 story entirely. Astra’s score is flat across effort compared with Sonnet’s steep climb.
That matches how you should configure APIs:
- Sonnet: effort is a real knob. Claude Code maps
sonnetto medium. API default effort on Sonnet 5.5 is high. Max is an eval setting, not a production default. - Astra: you are already near the index ceiling at medium/high. Paying for max buys less index movement than on Sonnet.
What GPT-6 Astra still owns
The Astra vs Fable 5.1 matrix still applies at the flagship price: computer use, some math/security lanes, token efficiency, launch-week GUI demos. Sonnet 5.5 did not delete that product.
Astra vs Sol is the OpenAI-internal trade: Sol is $2/$10, same sticker as Sonnet 5.5, weaker than Astra. Do not swap “Sonnet 5.5 vs Astra” with “Sonnet 5.5 vs Sol.” Sol’s AA bar in the screenshot is 48 (max). Sonnet 5.5 at 56 is a different comparison.
If the question is OpenAI mid-tier vs Anthropic mid-tier, pair this post with Sol vs Opus 5.5 — that is Sol vs flagship Claude, not Sonnet vs Astra.
GPT-6.1 is not a third column
OpenAI shelved October GPT-6.1 Astra after tests on scope authorization, deception, and how the model reports work. Full story: cancellation post. This comparison uses shipping GPT-6 Astra. A 6.1 bar is not on the AA chart you pasted.
How to pick this week
Use Sonnet 5.5 when:
- You are in Claude Code or the Claude API and the task is scoped (bugs, features, docs, slides).
- You can live with adaptive thinking and the migration breaks (
between_tools, no forcedtool_choice, toolset IDs). - You will hillclimb on your eval, not AA’s index — build-eval / hillclimb.
- Sticker $2/$10 matters and you will cap effort so you do not accidentally buy 193k-token max runs.
Use GPT-6 Astra when:
- The job is already on Codex/ChatGPT and switching vendors is the expensive part.
- AA Terminal-Bench-Science (or your own science-shell eval) still prefers Astra.
- You need Astra’s computer-use / efficiency profile from the Fable-week head-to-head.
- You want a model whose AA score does not depend on lighting money on fire at max.
Use Opus 5.5 when:
- You were about to set Sonnet to max anyway. At 58 vs 56, Opus is the honest flagship at $4/$20, not a mystery third screenshot bar.
Run one held-out set. If train improves and test does not, you overfit the index. That is the same warning as the hillclimb playbook.
Related reading
- Opus 5.5 vs Sonnet 5.5
- Claude Sonnet 5.5 building guide
- GPT-6 Astra launch benchmarks
- GPT-6 Astra vs Fable 5.1
- GPT-6 Astra vs Sol
- GPT-6.1 Astra October cancellation
- AA Intelligence Index v4.2 (older mix)
- How to read AI benchmarks
- Official: AA on Sonnet 5.5 · Building with Sonnet 5.5
Intelligence Index scores and token counts follow Artificial Analysis’s Sonnet 5.5 article and model cards as of September 29, 2026. Anthropic Terminal-Bench figures follow Anthropic’s launch table. They are different harnesses.
