Cloudflare has open-sourced the agent skill that its security team used as the seed for a fleet-wide vulnerability discovery system. The repo, cloudflare/security-audit-skill, is MIT licensed and showed about 25.6k stars at the time of writing, with a Hacker News thread at roughly 213 points. It turns a coding agent into an auditor that runs six phases and files only findings that an independent agent has tried and failed to disprove.
This post covers what the skill does, how to install it, the sandbox requirement you should not skip, and what developers say about its cost. For background on how skills work, see our agent skills guide.
TL;DR: questions people are asking
| Question | Answer |
|---|---|
| Who made it? | Cloudflare's security AI research team |
| License? | MIT |
| Install? | npx skills add https://github.com/cloudflare/security-audit-skill --skill security-audit |
| Do I need API keys of my own? | You need a coding agent with a model that supports tool use and parallel sub-agents |
| Is it safe to run on untrusted code? | Only inside an OS-enforced sandbox with no external network |
| Does one run find everything? | No. Cloudflare says one run found roughly half of what repeated runs found |
| Is it cheap? | One HN commenter reported heavy token use |
Where the skill came from
In June, Cloudflare published Build your own vulnerability harness, describing how it moved from a single skill to a multi-stage system. The post says the team started with a roughly 450-line security-audit skill run against one repository, tuning prompts until it surfaced real bugs, then added orchestration around it. Cloudflare says its prompts "continue to carry the initial skill's attacker scenarios, bug classes, and anti-pattern detections nearly unchanged."
The same post says going from the first slash-command run to a scanner covering 128 repositories took about six weeks, and that the team uses one model for discovery and a different model for validation so the two double-check each other. It also describes three walls the single skill hit: context exhaustion, no persistence after a crash, and no cross-repo reasoning. That is the context the new repo sits in. It is explicitly "the single-repo starting point" the harness evolved from.
This fits a wider pattern we have followed since the Glasswing preview, covered in our posts on Claude Mythos Preview and Glasswing and on VulnCheck's analysis of exploit rates for AI-found bugs: labs and vendors are moving from "can a model find a bug" to "how do you run this as a repeatable pipeline."
The six phases
The README describes the workflow this way:
- Reconnaissance. Map architecture, trust boundaries, input surfaces, prior evidence, and deterministic coverage into
architecture.mdandcoverage-ledger.json. - Coverage-led hunting. Assign isolated hunter agents from ledger units, record their checks, and use coverage critics to find gaps.
- Candidate validation. Give every unique candidate to a fresh verifier that tries to disprove it.
- Structured output. Write
confirmed,needs_validation, andrejectedrecords tofindings.json, validated againstreport-schema.json. - Independent record verification. Fresh agents verify final source claims, and material replacements get another independent verifier.
- Target-neutral reporting. Derive
REPORT.md,FINDINGS-DETAIL.md, andNEEDS-VALIDATION.mdfrom verified records and the coverage ledger.
Two small scripts enforce structure. validate-coverage-ledger.cjs runs after the ledger is created and after each update. validate-findings.cjs runs in Phase 4 and after every Phase 5 replacement. Both are zero-dependency Node scripts, so the discipline does not depend on the model remembering the schema.
The three verdicts
| Verdict | Meaning |
|---|---|
confirmed | Complete source trace and a bounded observed result |
needs_validation | An exact unresolved fact, and no severity assigned |
rejected | A candidate that was disproved |
That middle category is the useful one. Many AI audit tools present every suspicion as a finding. Here, a lead that cannot be proven stays labeled as unproven, with the specific question that blocks it.
The design principles
The README lists five principles that read like a checklist for any agentic review:
- Only confirm established boundary failures. A blocked but source-grounded lead stays
needs_validation. - Adversarial validation. The agent that checks a finding is never the agent that found it.
- Severity requires impact. Likelihood times impact, not deviation from a checklist.
- Defense-in-depth gaps are not vulnerabilities. If layer A prevents the attack, the missing layer B is a hardening note.
- Multiple runs improve coverage. Cloudflare says a single run found roughly half of what repeated runs found.
Separate files cover attack classes by target type, including memory safety and binaries, AI and LLM targets, web protocols and auth, client-side code, supply chain and release, cloud and deployment, RPC and messaging, resource exhaustion, data isolation, and desktop, mobile, and local IPC. The skill picks the classes that match the target, so a TypeScript web app does not get kernel-fuzzing prompts.
How to install and run it
Install with the Skills CLI:
npx skills add https://github.com/cloudflare/security-audit-skill \
--skill security-audit
Add --global for a user-level install. Then start your agent in, or pointed at, the codebase and ask in plain language:
security audit this codebase
find security vulnerabilities in ./src
do a security review, output to ~/audits/my-project
A direct audit or pen-test request uses full audit mode. Security questions and focused vulnerability work use a lighter guidance mode unless you ask for report artifacts. In full mode, an unspecified output directory defaults to ~/security-audit-skill/<repo-name>/run-<N>, and the workflow writes inside the target repository only if you explicitly choose a directory there.
The requirement people skip
The README is explicit about what you need: a coding agent whose model supports tool use and parallel sub-agents, Node.js for the validators, and an OS-enforced sandbox for anything target-controlled, such as builds, tests, processes, browsers, emulators, fuzzers, and fixtures. That sandbox must disable external networking, use a sanitized allowlisted environment, enforce resource limits, and allow writes only to assigned scratch paths. Without those controls, the workflow keeps the lead as needs_validation instead of executing target code.
Why it matters: to validate a bug the skill may build and run the code under audit, and an untrusted repo can execute anything in its build scripts. If you point this at code you did not write, treat the sandbox as mandatory. This is the same class of risk we described in our post on agent skills as a threat surface, only here the skill is the one running code.
What developers are saying
The HN discussion has substance on cost, noise, and framing. These are commenters' reports, not verified measurements.
- Cost. One commenter, drchaim, reported spending about 1M tokens on a medium codebase without a result to show for it. Cloudflare does not publish a per-run cost in the README.
- Skill sprawl. prodigycorp asked Cloudflare to consolidate its many platform skills into one routed skill, arguing that too many skill descriptions pollute the context window.
- Refusals. wslh suggested that security-framed skills can trigger refusals from top models, and that splitting work into bug-class skills without the security framing, plus a combining skill, works better for them. This is one user's workaround.
- Skepticism. acedTrex was sarcastic about a repo and post being dedicated to a markdown file. Our read, not a Cloudflare claim: the value lies in the orchestration and validators rather than any one prompt.
If you are tempted to dismiss it as only prompts, note the two validators and the schema: they make the output machine-checkable, which a plain code-review prompt does not.
How it compares to what you may already use
| Approach | Strength | Weakness |
|---|---|---|
| A one-shot review prompt | Cheap and quick | One hypothesis at a time; weak on coverage |
| This skill | Coverage ledger, adversarial validators, schema-checked output | Heavy token use, needs a sandbox |
| Cloudflare's full harness | Persistence, cross-repo tracing, dedup | Not released; you build it |
| Vendor tools such as OpenAI Codex Security | Integrated product | Tied to one stack |
Cloudflare's own advice in the harness post is to start with a skill, get the prompts working, and build the next architectural stage only when not having it is the thing slowing you down.
What to do this week
- Run it on a repo you own, inside a container or VM with networking disabled.
- Run it at least twice. Cloudflare says repeat runs find more, and the skill uses prior ledgers to target gaps.
- Read
NEEDS-VALIDATION.mdfirst. The unresolved leads show where the agent could not prove something and what fact is missing. - Do not file bug reports from
confirmedrecords without reading the source trace yourself. - Track token cost per run before scheduling it in CI.
What remains unclear
- Cloudflare has not published precision or false-positive rates for the open-source skill.
- Token cost per run is unreported by Cloudflare.
- Which models perform best in the skill is not stated in the README; the harness post says Cloudflare swaps models and uses different ones for discovery and validation.
- How many real vulnerabilities it has found outside Cloudflare's own fleet is unknown.
We will update this post if Cloudflare publishes benchmarks.
Related reading
- What are agent skills? Complete guide
- Claude Mythos Preview and Project Glasswing
- VulnCheck: AI-found bugs and exploit rates
- OpenAI Codex Security CLI and SDK go open source
- OpenAI Daybreak and Codex Security for cyber defense
- Agent skills as a security threat
- Anthropic on GLM-5.3 open-weight cyber capabilities
- Top prompts for security review
Repository details, star counts, and requirements reflect the cloudflare/security-audit-skill README as of October 7, 2026 and may change.
