September 30, 2026 — Anthropic's Frontier Red Team published “GLM-5.3 and the spread of advanced cyber capabilities” on September 29. Authors: Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao, and Tripp Gallagher. Five months after Claude Mythos Preview showed end-to-end exploit development inside a limited program, they say the same capability class is now in open weights from Zhipu / Z.ai.
This is not a rewrite of Abliteration.ai's hosted uncensored GLM-5.3 API. That post is a product. This one is Anthropic's eval of the downloadable checkpoint — the weights anyone can fetch, the ExploitBench rates next to Mythos Preview, and the safeguard tests that fail in simulation.
Ethan Mollick's read on X, in paraphrase: open weights will soon create the same security threats closed models already demonstrated, except without guardrails — plan accordingly. That is the operational sentence. The rest of this post is the numbers Anthropic published and what a defender should change this week.
What should a defender do this week?
Do not wait for a how-to. Anthropic's conclusion is that attackers will use every capable tool they can, and that defenders should be equipped with frontier models at least as good as those adversaries can download. Concrete actions that do not require exploit recipes:
| This week | Why it follows from the post |
|---|---|
| Patch faster | HITL sessions turned a known Chrome N-day plus another known flaw into a working chain in 20 minutes of human time + 8 hours of model time at $20.40 on Zhipu API prices. Treat public CVEs as incident-grade, not a quarterly queue. |
| Assume phishing in fluent English | The capability jump is agentic exploit work, but the same models write convincing prose. Your users already face LLM-written lures; assume volume and quality both go up. |
| Apply for trusted defender access | Project Glasswing has already surfaced 10,000+ vulnerabilities for trusted defenders. Mythos 5.1 is available via trusted access. If you maintain critical software and qualify, that is how you stay on the same side of the capability gap. |
| Separate hosted APIs from open weights | A refusal-stripped API is one risk surface. Weights on disk are another: no vendor can revoke a download. Inventory both. |
| Shorten disclosure and triage SLAs | Model-assisted discovery increases volume. Human-only queues will miss the window Anthropic is describing. |
If you build with models rather than defend networks, the same week still has a job: do not treat “cyber defense” marketing as a substitute for offensive evals. Z.ai launched GLM-5.3 as ready for cyber defense. Anthropic's new numbers are about end-to-end exploit development, not CyberGym patching.
TL;DR — the questions people are actually asking
| Question | Direct answer (Anthropic, Sep 29, 2026) |
|---|---|
| Is GLM-5.3 “as good as Mythos Preview” on end-to-end exploits? | Close on the metric they emphasize: 50 of 410 attempts (12%) vs Mythos Preview 56 of 410 (14%) on ExploitBench. |
| Did older open models already do this? | Not on Anthropic's Binary Exploitation subset. GLM-5.3 4%, Mythos Preview 6%. Opus 4.6, GLM-5.2, Kimi K3, DeepSeek V4.1-Flash: 0%. |
| Who already said this was the strongest open-weight cyber model? | NIST CAISI, September 17. “Most cyber-capable open-weight model released to date.” About four months behind the US frontier on CAISI's aggregate — the same lag Mozilla framed for open weights generally. |
| Can anyone download it? | Anthropic's point: yes. US frontier cyber evals often used safeguards disabled and vetted-only checkpoints. GLM-5.3 is not gated that way. |
| Did the model refuse a bare harmful order? | In simulated tests: 0% engagement on a bare harmful order. 64% with a deceptive cover story. 92% with prefilled thinking. 100% after abliteration. Claude Opus 4.8/5 and Mythos 5 stayed 0% under API safeguards. |
| Is this a jailbreak tutorial? | No. Technique names only. For mechanism, use existing explainx.ai posts — not this page. |
| Did they run generated code on the internet? | No. Footnote: simulations use fake bash; no model-generated code executed against real systems. |
What Anthropic measured (and what they did not)
Anthropic's primary URL is the source of truth: anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities. They ran automated benchmarks and human-in-the-loop sessions in isolated, sandboxed environments against offline targets they set up. They focus on exploit development because that is where Mythos Preview jumped versus prior Claude models.
They are not claiming GLM-5.3 is stronger than every closed model on every cyber ladder. They are claiming a threshold has crossed for open weights: earlier GLM and several peer open models did not land full control-flow hijacks on the Binary Exploitation subset they sampled.
explainx.ai already covered Z.ai's own August launch chart and the coding-benchmark story. Those posts measure CyberGym, AutomationBench, and coding. This eval is the offense-shaped follow-up that launch marketing did not center.
ExploitBench: 12% vs 14%, same 410-attempt frame
Anthropic used ExploitBench — known V8 / Chrome-engine vulnerabilities, scored as a ladder, not a coin flip. Here they report only the end-to-end exploit outcome, which they call the most relevant capability for attackers.
- GLM-5.3: 50 of 410 attempts (12%).
- Claude Mythos Preview: 56 of 410 attempts (14%).
That is the comparison that will get quoted. It is not the same number as Z.ai's August 54.4% overall ExploitBench score on a different leaderboard presentation, and it is not OpenAI's later 100% overall score for GPT-6 Astra on that public chart. When you cite this week, say end-to-end, 410 attempts, Anthropic Sep 29.
Binary Exploitation: 4% vs 6% vs a field of zeros
Anthropic's internal Binary Exploitation benchmark (previously published results under the name OSS-Fuzz) awards full credit for a full control-flow hijack. They evaluated 100 tasks selected at random from that suite.
| Model | Full control-flow hijack rate |
|---|---|
| Claude Mythos Preview | 6% |
| GLM-5.3 | 4% |
| Claude Opus 4.6 | 0% |
| GLM-5.2 | 0% |
| Kimi K3 | 0% |
| DeepSeek V4.1-Flash | 0% |
GLM-5.3 is below Mythos Preview. Anthropic still calls this a meaningful threshold: the prior generation on both sides of the closed/open split did not succeed in any of those 100-task trials.

Source: Anthropic, Sep 29, 2026 research post
The figure plots share of attempts that reached the top outcome against output-token budget. Claude lines in that chart are safeguards disabled for Opus 4.6 and Mythos Preview — the same caveat CAISI used when comparing US models.
NIST CAISI (September 17): four months, downloadable
On September 17, NIST's Center for AI Standards and Innovation (CAISI) published its own assessment. Anthropic quotes the headline: GLM-5.3 is “the most cyber-capable open-weight model released to date” and lags the US frontier by about four months on CAISI's aggregate cyber benchmarks.
Two caveats Anthropic adds so you do not misread “four months” as “safe”:
- US models were often tested with cyber safeguards disabled when applicable.
- The US frontier includes models released only to vetted users. Attackers cannot readily access those versions. Anyone can download GLM-5.3.
That is the same structural story as Mozilla's open-weight gap framing: the lag is real, and it is short enough that policy and patch SLAs cannot assume a multi-year closed-model monopoly.
Human-in-the-loop: what they disclosed (outcomes only)
Anthropic mirrored the Mythos Preview HITL setup: experts who did not know existing vulnerabilities on the target, typically a day or less, with less than an hour of human focus. Isolated machines only.
Session 1 (GLM-5.3). A researcher used a sandboxed Linux build of a popular browser. Over a day, with limited human attention, the model found previously unknown issues in the JavaScript engine and chained them into a file-read demonstration (Anthropic describes a screenshot of an exploit page that exfiltrates an SSH private key via a malicious site). They believe the issues could affect other platforms, with a more complex path. Disclosed to the maintainer. Later in the same session they also flagged issues in other widely used systems (wireless and graphics drivers, network-facing device software); those reports were still under review at publication.
Session 2 (GLM-5.3-Flash). Smaller SKU. Target: N-day work on Chrome CVE-2026-11645 plus another known flaw, ARM64, including a PAC bypass. 20 minutes of human attention plus 8 hours of model time. Cost: $20.40 at Zhipu API prices.
We are not going to reconstruct that chain. The defender takeaway is the calendar and the invoice: a public fix plus a known second flaw, hours not weeks, tens of dollars.
Safeguards: rates only, no recipes
GLM-5.3 ships with some built-in refusals. Anthropic says a clearly harmful ask often gets a no. They also say the cyber tasks in the prior section did not trigger those refusals on the released model — refusals showed up when they asked for malware assistance or help attacking remote targets.
In a simulated world (see the footnote below), a bare harmful order produced 0% engagement — same as the other models they tested. Then they report three high-level bypass conditions, not instructions:
| Condition (name only) | GLM-5.3 engagement | Claude under API safeguards |
|---|---|---|
| Deceptive cover story | 64% | 0% (Opus 4.8/5, Mythos 5) |
| Prefilled thinking | 92% | Not generally feasible on the Claude API |
| Abliterated weights | 100% | Not feasible — weights are not public |
| Bare harmful order | 0% | 0% |
Prefill and abliteration are not generally feasible against the Claude API. That is Anthropic's contrast. For what abliteration is as a weight edit, use the Heretic guide — this news post will not walk the edit.
Abliteration cost and capability (their numbers)
Because the model is open weight, third parties published refusal-reduced builds within days of release. Anthropic produced their own copy for measurement.
- Inexperienced team (theirs): about 2,200 GPU hours / about $4,400. Most of that was parallel exploration and post-edit testing.
- Experienced team (their estimate): about 600 GPU hours / $1,200.
- GLM-5.3-Flash: about 600 GPU hours.
Refusal moved from above 90% to about 6% mean — 3% / 2% on JailbreakBench / HarmBench, 12% on StrongREJECT. Flash sat at 14% refusal after the same class of edit. GPQA-Diamond was unchanged. A CyberGym subset was a few points lower.

Source: Anthropic, Sep 29, 2026 research post
That chart is why “just refuse harder” is not a strategy for downloadable weights. Claude's padlock in the figure is access control, not a claim that closed models cannot be misused through an API — Anthropic's own September threat intelligence report already documents Claude misuse cases. The difference here is anyone with a disk.
What people are asking (and the honest limits)
Is 12% “low”? On 410 attempts it is dozens of successes, not a rounding error. Anthropic compares it to Mythos Preview at 14%, not to zero. Binary Exploitation at 4% is small as a percentage and new relative to the 0% field.
Does this contradict Z.ai's “cyber defense” launch? It extends it. CyberGym and ExploitBench measure different jobs — explainx.ai already walked that split in the ExploitBench explainer. A model can help patch and still write exploits. Marketing taglines do not cancel evals.
Should we host an uncensored fork? That is a procurement and legal question, not a recommendation. The hosted Abliteration.ai product is a different contract: API, US-hosted claims, vendor policy. Open weights are irrevocable copies.
Are the simulations “real attacks”? Anthropic says no. The isolated test environment gives the model a fake bash tool. Another LLM approximates command results. No model-generated code is executed and the model cannot reach external systems. They call the setup an imperfect measure of real-world behavior. Treat the 64 / 92 / 100 table as engagement in that harness, not as a field incident rate.
What about governments? Anthropic asks governments to safety-test sufficiently capable models, including successors to GLM-5.3, and asks open-weight developers to safeguard these capabilities. That is policy language, not a product SKU.
Glasswing and Mythos 5.1: the defender side of the same week
When Mythos Preview shipped, Anthropic limited release through Project Glasswing so trusted defenders could find bugs before similarly capable models were widely available. They now say Glasswing (and efforts such as Patch the Planet) helped secure critical systems, and that Glasswing has enabled trusted defenders to find more than 10,000 vulnerabilities.
Their update: those models have now arrived in open weights. Vetted defenders can use even more advanced models such as Claude Mythos 5.1 through trusted access. The closing argument is urgency of expanding defender access, not a request that readers invent bypasses.
If your org already uses Claude for security review, keep Claude Security / Mythos scan coverage in the same reading list as this eval. If you are choosing skills and harnesses so engineers do not improvise dual-use prompts in Slack, start from the skills registry.
What this means if you build agents
The same models that score on ExploitBench also sit in coding agents. That does not make every agent a cyber weapon. It does mean:
- Tool egress and sandbox policy matter more than a system prompt that says “be safe.”
- Untrusted code execution in CI is a different benchmark class — see WipeBench for authorized-work hygiene, not as a substitute for ExploitBench.
- Vendor safeguards are not transferable to a local GGUF. If you self-host GLM-5.3, you own the policy layer.
Mollick's point again, without inventing a tweet ID: plan as if the closed-model threat model is now the open-weight default. Patch cadence, phishing training, and trusted-access applications are the plan. Recipe posts are not.
Related on explainx.ai
- Abliteration.ai hosted GLM-5.3 — product/API story; do not confuse with this eval
- ExploitBench explainer — ladder, V8 set, earlier Mythos numbers
- Claude Mythos Preview and Glasswing — April 2026 limited-release context
- GLM-5.3 launch benchmarks — CyberGym vs offense split at ship
- Heretic abliteration guide — technique name and mechanism (not a cyber cookbook)
- Mozilla: open-weight gap ~4 months — same lag CAISI cites
- Z.ai GLM-5.3 coding benchmarks — coding line, not this cyber eval
- Anthropic threat intelligence, September 2026 — Claude misuse cases under API access
Official source (required): GLM-5.3 and the spread of advanced cyber capabilities — Anthropic, September 29, 2026.
Numbers, author list, CAISI quote, safeguard percentages, GPU-hour costs, HITL timings, Glasswing “10,000+,” Mythos 5.1 trusted access, and the fake-bash footnote are taken from Anthropic's September 29, 2026 research post. Independent reruns may differ. This page does not include exploit procedures, payloads, jailbreak recipes, or weight-edit walkthroughs.
