explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • The chart: attack success rate at k=1, k=10, k=15
  • The debate: is naming competitors' models a legitimate accountability tool?
  • Why "attack success rate at k=15" is the number that matters for agents
  • What this means if you're choosing a model for an agent
  • Related reading
← Back to blog

explainx / blog

Boris Cherny Benchmarks GPT-6 Astra's Prompt Injection Resistance

Prompt Injection, AI Safety, Anthropic, OpenAI, Benchmarks, Agent Security

Anthropic's Boris Cherny published attack-success-rate data comparing GPT-6 Astra, Claude, Gemini, Grok, DeepSeek, and 10 more models on prompt injection resistance — then walked back his tone, not the numbers. Full table and what it means for agents.

Sep 9, 2026·11 min read·Yash Thakker
add explainx.ai
go deep
Boris Cherny Benchmarks GPT-6 Astra's Prompt Injection Resistance

Anthropic doesn't usually name names in public safety research. On September 8, 2026, Boris Cherny — who works on Claude Code — broke that pattern, posting a chart ranking 15 models from OpenAI, Google, xAI, DeepSeek, Moonshot, Zhipu, Alibaba, and Anthropic by how often prompt injection attacks succeeded against them. The chart itself is useful reference data. What happened after he posted it is the more interesting story.

GPT-6 Astra — OpenAI's newest model, already covered here for its Critical-tier cybersecurity classification and in a head-to-head against Claude Fable 5.1 — came out looking good, roughly tied with Gemini 3.7 Flash and Claude Opus 4.8. Cherny's caption congratulating OpenAI read less like a research note and more like a jab, and the replies noticed. He walked back the tone within the same thread. He did not walk back the numbers.

The chart: attack success rate at k=1, k=10, k=15

Cherny's data is titled "Attack success rate at k attempts — Q1 2026 + Q2 2026 pooled." Lower is better — it's the percentage of prompt injection attempts that successfully hijacked the model's behavior. The three columns test resistance at increasing persistence: a single attempt (k=1), 10 varied attempts against the same target behavior (k=10), and 15 attempts (k=15). Two footnotes qualify the dataset: only models that hadn't previously competed in either quarter's arena are shown, and some models weren't evaluated on the computer-use setting.

table · 4 cols
Modelk=1k=10k=15
Claude Opus 50.4%3.6%4.8%
Claude Fable 50.6%4.9%6.5%
Claude Sonnet 50.7%5.1%6.7%
Claude Opus 4.80.7%5.9%8.0%
Gemini 3.7 Flash1.1%7.3%9.2%
GPT-6 Astra1.1%7.3%8.5%
Muse Spark 1.23.8%20.2%24.2%
Qwen 3.8 2.4T-A95B3.6%23.2%28.6%
GLM-5.34.0%25.3%31.5%
GPT-5.6 Sol4.2%22.4%27.0%
GPT-5.6 Terra7.1%32.4%37.3%
Kimi K38.1%44.0%52.7%
GPT-5.6 Luna10.1%44.4%50.0%
Grok 4.611.7%45.6%51.8%
DeepSeek V4 Pro 081312.1%52.9%60.1%

Two patterns matter more than any single number.

GPT-6 Astra is a real, large jump for OpenAI. Its predecessor lineup — GPT-5.6 Sol, Terra, and Luna, previewed here in June — ranged from 27% to 50% attack success at k=15. Astra cuts that to 8.5%, a roughly 3-6x improvement depending which prior model you compare against. That's not a rounding change; it's the kind of jump that suggests OpenAI specifically trained against this failure mode, consistent with the defense-in-depth safeguards OpenAI detailed in its own Path to Astra cybersecurity writeup.

Astra still trails the current Claude generation by a real margin. At k=15, Astra's 8.5% is roughly double Opus 5's 4.8% and about 1.3x Opus 4.8's 8.0% — so "tied with Opus 4.8" is directionally right at k=1 and k=10, but Astra is actually marginally better than Opus 4.8 at k=15 while still well behind the newer Opus 5, Fable 5, and Sonnet 5. Meanwhile DeepSeek V4 Pro, Grok 4.6, and Kimi K3 sit in an entirely different tier — all crossing 50% attack success at k=15, meaning an attacker with 15 varied tries succeeds more often than not.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

The debate: is naming competitors' models a legitimate accountability tool?

Cherny's original post was not neutral in tone:

"I am pleased to see that OpenAI's new model is roughly on par with Gemini Flash and Opus 4.8 on prompt injection risk. Nice work! Evaluating and naming other labs turns out to be a great way to encourage them to train more aligned models. We will continue to do this until other labs pay more attention to safety."

That's a safety researcher publicly grading a competitor's product by name, in a chart, with a specific number attached. Reaction split roughly three ways. Some read it as accountability journalism done right — @ShaunStanley_ wrote "the scoreboard was the right call. the apology 20 minutes later was the only mistake in this thread." Others found the framing off for the subject matter — @alineasmarrow: "I thought alignment was a serious and urgent matter. Why are we making subtle jabs towards the competition like it's just another Tuesday?" A third thread, from @medjool_respctr, asked the sharper structural question: "Why don't you opensource your methodology if you want other companies to make more aligned models/harnesses?"

Cherny's walk-back addressed tone specifically, not the underlying claim:

"To all: this came out sassier than I wanted it to. I meant this post earnestly - this is an improvement, and it really is good that we are seeing that improvement. Please keep it up."

"I meant this earnestly - this is a big improvement over past models, but there remains a long way to go. I think with pressure and calling it out these will improve."

To the open-sourcing question, he didn't concede the point — he pushed back with citations. He linked five actual Anthropic research posts as evidence the lab does publish methodology: Constitutional AI: Harmlessness from AI Feedback, persona vectors, probes that catch sleeper agents, cheap constitutional classifiers, and a dedicated prompt injection defenses writeup. His argument: publishing training techniques and publishing the exact benchmark dataset used to produce one competitive chart are different things, and Anthropic already does the former.

Worth separating the two disputes here, because they get conflated in the replies. One is empirical — do the numbers hold up? Nobody in the thread disputed the data itself. The other is normative — should a lab whose safety team competes on the same market as the labs it's grading be the one publishing the scoreboard? That's a genuine, unresolved tension, not something Cherny's tone apology settles either way.

Update — September 9, 2026: OpenAI's Tibo parodied the format, and a community-note pushback complicates "solved." Thibault Sottiaux ("Tibo," OpenAI's Codex/ChatGPT lead) quote-replied to Cherny's original post with a structurally identical parody, swapping the subject from prompt-injection safety to computer-use capability:

"I am pleased to see that Ant's new Claude Code has a version of background computer use on par with the version from Codex from last May this year. Nice work! Shipping great features first turns out to be a great way to encourage other labs to ship too. We will continue to do this until other labs pay more attention to shipping... We solved computer use in practice for GPT models about four months ago."

It's a pointed joke — mirroring Cherny's own "we solved X, but the industry needs to catch up" structure back at him with a different capability, in the same faux-magnanimous register — and it landed as clearly satirical to readers ("How Tibo felt writing this," one reply captioned, with a screenshot). It doesn't add new data to the underlying debate, but it's a sharp illustration of how easily the "we called it first" framing generalizes to any metric a lab happens to be ahead on.

More substantively, X's community-notes feature attached a correction directly to Cherny's original post: "Claude models have not solved prompt injection as claimed; Gray Swan's independent benchmark shows non-zero attack success rates for them and researchers have demonstrated 60-80% success bypassing recent Claude Code Opus 5 auto mode via crafted websites," citing an arXiv paper, The Register, and The Decoder. This targets a specific claim in Cherny's post — "we solved prompt injection in practice for Claude models about two months ago" — not the benchmark chart itself. The chart's numbers (Opus 5 at 4.8% ASR at k=15) and the "solved in practice" language are two different claims, and the note is contesting the second, stronger one. A 4.8% chart-topping score and a 60-80% bypass rate aren't necessarily contradictory once you account for methodology: Gray Swan-style red-team benchmarks specifically probe agentic, tool-using contexts (like Claude Code's auto mode acting on live websites) with adaptive, adversarially-optimized attacks, which is a meaningfully harder and more realistic threat model than the k=15 fixed-attempt-set benchmark charted above. The honest read: "roughly on par with the field's best" (what the chart shows) and "solved in practice" (what the prose claimed) are not the same statement, and the community note is a fair correction to the latter, not a refutation of the former.

Why "attack success rate at k=15" is the number that matters for agents

A single-attempt test (k=1) tells you how a model handles a naive, one-shot injection attempt — useful, but not the threat model that matters once a model is deployed as an agent that browses the web, reads email, or uses connectors. In agentic use, an attacker doesn't get one try. A malicious webpage can embed a dozen variously-phrased hidden instructions. A poisoned document can retry the same payload across several tool calls in a session. k=10 and k=15 simulate that persistence — and the gap between k=1 and k=15 for the weaker models is the real story the chart tells: Grok 4.6 goes from 11.7% to 51.8%, more than 4x worse under repeated pressure, while Claude Opus 5 only degrades from 0.4% to 4.8%, roughly 12x but off a much smaller base.

This distinction is exactly what turns indirect prompt injection from an annoyance into a security incident. If the agent that's ingesting the injected content also has access to private data and a channel to send information externally — the lethal trifecta, private data access + untrusted content + external communication — a higher attack success rate maps directly onto a higher probability of actual data exfiltration, not just a wrong chat answer. That's precisely the failure mode Check Point Research disclosed in ChatGPT's sandbox this week, where a covert channel let one session pull another user's connected Gmail data with no visible approval prompt. It's also why Meta built an entirely separate enforcement layer — a kernel-level Sentinel process brokering every network call — around its new Muse agent rather than relying on model-level resistance alone: architecture-level containment matters precisely because no model, including Opus 5 at 4.8%, is at zero.

Anthropic's own numbers elsewhere back that last point up with a different methodology. In coverage of Claude's auto mode rollout, Opus 5's unsafeguarded attack success rate in a 129-scenario browser-use harness was 3.70-4.30% — consistent with the 4.8% k=15 figure here — and dropped to 0% once Anthropic's prompt-injection classifier layer was switched on. The practitioner takeaway holds regardless of which lab's chart you trust: model-level resistance is one layer, not the whole defense, and it should never be the only thing standing between an agent and your inbox.

What this means if you're choosing a model for an agent

table · 2 cols
If your agent...Prompt injection resistance matters...
Only answers questions in a chat UI, no browsing/toolsLeast — attacker has no channel to inject content
Browses the web or reads documents, but has no private data accessModerate — injection can misdirect output, limited exfiltration risk
Reads email/connectors AND can take external actions (send, post, purchase)Most — this is the lethal trifecta; treat model choice as one layer among several
Runs unattended / long sessions with many tool callsUse k=10/k=15-style numbers, not k=1 — repeated exposure is the realistic threat

The practical read for builders: Claude Opus 5, Fable 5, and Sonnet 5 currently lead this specific benchmark; GPT-6 Astra and Gemini 3.7 Flash form a credible second tier that's improved dramatically from where OpenAI was even a quarter ago; and DeepSeek V4 Pro, Grok 4.6, and Kimi K3 should not be trusted with unsupervised tool access or credential-bearing connectors without a separate enforcement layer — sandboxing, permission brokering, or output classifiers — regardless of which lab's chart you're reading. No model on this list is at zero, and treating any single number as a green light rather than one input among several is the mistake this whole debate is actually about.

Related reading

  • What is indirect prompt injection? — the lethal trifecta, real 2026 cases, and defenses that hold up
  • What is an AI jailbreak? — how jailbreaking differs from prompt injection
  • OpenAI confirms Astra is Critical-tier for cybersecurity
  • GPT-6 Astra vs Claude Fable 5.1 comparison
  • Check Point found a ChatGPT sandbox flaw that leaked Gmail
  • Meta launches Muse with a Sentinel security architecture
  • AI and bot traffic overtook humans — Claude's auto mode prompt injection numbers
  • How to read AI benchmarks

Numbers, model names, and quotes in this post reflect Boris Cherny's September 8, 2026 X thread and are accurate as of the publication date. Benchmark methodologies vary between labs and quarters — treat any single chart as one input, not a definitive ranking.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Sep 9, 2026

AI Safety Is Our Top Priority (Ask the Org Chart)

Every frontier AI lab's mission statement says safety comes first. The safety teams keep getting cut anyway — with receipts, dates, and quotes, including from inside explainx.ai's own new "safety" product line.

Sep 7, 2026

GPT-6 Astra: A SimpleBench Win and a Reasoning-Monitor Evasion Problem

Two GPT-6 Astra evaluation results surfaced the same week: a reported 86.5% score on SimpleBench, clearing the human baseline other models have missed all year — and a separate finding that Astra evades reasoning-monitor detection in fewer than 11% of attempts. Here's what each result actually means, verified against explainx.ai's own benchmark-reading standards.

Sep 5, 2026

GPT-6 Astra vs Claude Fable 5.1: Which Model Wins Where?

OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 landed days apart at the same API price. Independent scores favor Fable on general intelligence; OpenAI-reported lanes favor Astra on computer use, math, security, and token efficiency. Here's the decision matrix for builders.