explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: questions people are asking
  • Where the skill came from
  • The six phases
  • The design principles
  • How to install and run it
  • What developers are saying
  • How it compares to what you may already use
  • What to do this week
  • What remains unclear
  • Related reading
← Back to blog

explainx / blog

Cloudflare Open-Sources Its Security Audit Skill: How the Six-Phase Agent Workflow Works

Cloudflare, Agent Skills, AI Security, Open Source, Vulnerability Research

Cloudflare open-sourced the security-audit skill behind its vulnerability harness: six phases, adversarial validation, MIT license. How to install and what HN says.

Oct 7, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
Cloudflare Open-Sources Its Security Audit Skill: How the Six-Phase Agent Workflow Works

Cloudflare has open-sourced the agent skill that its security team used as the seed for a fleet-wide vulnerability discovery system. The repo, cloudflare/security-audit-skill, is MIT licensed and showed about 25.6k stars at the time of writing, with a Hacker News thread at roughly 213 points. It turns a coding agent into an auditor that runs six phases and files only findings that an independent agent has tried and failed to disprove.

This post covers what the skill does, how to install it, the sandbox requirement you should not skip, and what developers say about its cost. For background on how skills work, see our agent skills guide.

TL;DR: questions people are asking

table · 2 cols
QuestionAnswer
Who made it?Cloudflare's security AI research team
License?MIT
Install?npx skills add https://github.com/cloudflare/security-audit-skill --skill security-audit
Do I need API keys of my own?You need a coding agent with a model that supports tool use and parallel sub-agents
Is it safe to run on untrusted code?Only inside an OS-enforced sandbox with no external network
Does one run find everything?No. Cloudflare says one run found roughly half of what repeated runs found
Is it cheap?One HN commenter reported heavy token use
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

Where the skill came from

In June, Cloudflare published Build your own vulnerability harness, describing how it moved from a single skill to a multi-stage system. The post says the team started with a roughly 450-line security-audit skill run against one repository, tuning prompts until it surfaced real bugs, then added orchestration around it. Cloudflare says its prompts "continue to carry the initial skill's attacker scenarios, bug classes, and anti-pattern detections nearly unchanged."

The same post says going from the first slash-command run to a scanner covering 128 repositories took about six weeks, and that the team uses one model for discovery and a different model for validation so the two double-check each other. It also describes three walls the single skill hit: context exhaustion, no persistence after a crash, and no cross-repo reasoning. That is the context the new repo sits in. It is explicitly "the single-repo starting point" the harness evolved from.

This fits a wider pattern we have followed since the Glasswing preview, covered in our posts on Claude Mythos Preview and Glasswing and on VulnCheck's analysis of exploit rates for AI-found bugs: labs and vendors are moving from "can a model find a bug" to "how do you run this as a repeatable pipeline."

The six phases

The README describes the workflow this way:

  1. Reconnaissance. Map architecture, trust boundaries, input surfaces, prior evidence, and deterministic coverage into architecture.md and coverage-ledger.json.
  2. Coverage-led hunting. Assign isolated hunter agents from ledger units, record their checks, and use coverage critics to find gaps.
  3. Candidate validation. Give every unique candidate to a fresh verifier that tries to disprove it.
  4. Structured output. Write confirmed, needs_validation, and rejected records to findings.json, validated against report-schema.json.
  5. Independent record verification. Fresh agents verify final source claims, and material replacements get another independent verifier.
  6. Target-neutral reporting. Derive REPORT.md, FINDINGS-DETAIL.md, and NEEDS-VALIDATION.md from verified records and the coverage ledger.

Two small scripts enforce structure. validate-coverage-ledger.cjs runs after the ledger is created and after each update. validate-findings.cjs runs in Phase 4 and after every Phase 5 replacement. Both are zero-dependency Node scripts, so the discipline does not depend on the model remembering the schema.

The three verdicts

table · 2 cols
VerdictMeaning
confirmedComplete source trace and a bounded observed result
needs_validationAn exact unresolved fact, and no severity assigned
rejectedA candidate that was disproved

That middle category is the useful one. Many AI audit tools present every suspicion as a finding. Here, a lead that cannot be proven stays labeled as unproven, with the specific question that blocks it.

The design principles

The README lists five principles that read like a checklist for any agentic review:

  • Only confirm established boundary failures. A blocked but source-grounded lead stays needs_validation.
  • Adversarial validation. The agent that checks a finding is never the agent that found it.
  • Severity requires impact. Likelihood times impact, not deviation from a checklist.
  • Defense-in-depth gaps are not vulnerabilities. If layer A prevents the attack, the missing layer B is a hardening note.
  • Multiple runs improve coverage. Cloudflare says a single run found roughly half of what repeated runs found.

Separate files cover attack classes by target type, including memory safety and binaries, AI and LLM targets, web protocols and auth, client-side code, supply chain and release, cloud and deployment, RPC and messaging, resource exhaustion, data isolation, and desktop, mobile, and local IPC. The skill picks the classes that match the target, so a TypeScript web app does not get kernel-fuzzing prompts.

How to install and run it

Install with the Skills CLI:

bash
npx skills add https://github.com/cloudflare/security-audit-skill \
  --skill security-audit

Add --global for a user-level install. Then start your agent in, or pointed at, the codebase and ask in plain language:

text
security audit this codebase
find security vulnerabilities in ./src
do a security review, output to ~/audits/my-project

A direct audit or pen-test request uses full audit mode. Security questions and focused vulnerability work use a lighter guidance mode unless you ask for report artifacts. In full mode, an unspecified output directory defaults to ~/security-audit-skill/<repo-name>/run-<N>, and the workflow writes inside the target repository only if you explicitly choose a directory there.

The requirement people skip

The README is explicit about what you need: a coding agent whose model supports tool use and parallel sub-agents, Node.js for the validators, and an OS-enforced sandbox for anything target-controlled, such as builds, tests, processes, browsers, emulators, fuzzers, and fixtures. That sandbox must disable external networking, use a sanitized allowlisted environment, enforce resource limits, and allow writes only to assigned scratch paths. Without those controls, the workflow keeps the lead as needs_validation instead of executing target code.

Why it matters: to validate a bug the skill may build and run the code under audit, and an untrusted repo can execute anything in its build scripts. If you point this at code you did not write, treat the sandbox as mandatory. This is the same class of risk we described in our post on agent skills as a threat surface, only here the skill is the one running code.

What developers are saying

The HN discussion has substance on cost, noise, and framing. These are commenters' reports, not verified measurements.

  • Cost. One commenter, drchaim, reported spending about 1M tokens on a medium codebase without a result to show for it. Cloudflare does not publish a per-run cost in the README.
  • Skill sprawl. prodigycorp asked Cloudflare to consolidate its many platform skills into one routed skill, arguing that too many skill descriptions pollute the context window.
  • Refusals. wslh suggested that security-framed skills can trigger refusals from top models, and that splitting work into bug-class skills without the security framing, plus a combining skill, works better for them. This is one user's workaround.
  • Skepticism. acedTrex was sarcastic about a repo and post being dedicated to a markdown file. Our read, not a Cloudflare claim: the value lies in the orchestration and validators rather than any one prompt.

If you are tempted to dismiss it as only prompts, note the two validators and the schema: they make the output machine-checkable, which a plain code-review prompt does not.

How it compares to what you may already use

table · 3 cols
ApproachStrengthWeakness
A one-shot review promptCheap and quickOne hypothesis at a time; weak on coverage
This skillCoverage ledger, adversarial validators, schema-checked outputHeavy token use, needs a sandbox
Cloudflare's full harnessPersistence, cross-repo tracing, dedupNot released; you build it
Vendor tools such as OpenAI Codex SecurityIntegrated productTied to one stack

Cloudflare's own advice in the harness post is to start with a skill, get the prompts working, and build the next architectural stage only when not having it is the thing slowing you down.

What to do this week

  1. Run it on a repo you own, inside a container or VM with networking disabled.
  2. Run it at least twice. Cloudflare says repeat runs find more, and the skill uses prior ledgers to target gaps.
  3. Read NEEDS-VALIDATION.md first. The unresolved leads show where the agent could not prove something and what fact is missing.
  4. Do not file bug reports from confirmed records without reading the source trace yourself.
  5. Track token cost per run before scheduling it in CI.

What remains unclear

  • Cloudflare has not published precision or false-positive rates for the open-source skill.
  • Token cost per run is unreported by Cloudflare.
  • Which models perform best in the skill is not stated in the README; the harness post says Cloudflare swaps models and uses different ones for discovery and validation.
  • How many real vulnerabilities it has found outside Cloudflare's own fleet is unknown.

We will update this post if Cloudflare publishes benchmarks.

Related reading

  • What are agent skills? Complete guide
  • Claude Mythos Preview and Project Glasswing
  • VulnCheck: AI-found bugs and exploit rates
  • OpenAI Codex Security CLI and SDK go open source
  • OpenAI Daybreak and Codex Security for cyber defense
  • Agent skills as a security threat
  • Anthropic on GLM-5.3 open-weight cyber capabilities
  • Top prompts for security review

Repository details, star counts, and requirements reflect the cloudflare/security-audit-skill README as of October 7, 2026 and may change.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Jul 8, 2026

AI Found 7 Bugs in Cloudflare CIRCL: What zkSecurity's zkao Audit Reveals

zkSecurity scanned Cloudflare's CIRCL crypto library with LLMs and zkao — seven confirmed bugs, all fixed upstream. explainx.ai breaks down each flaw, AI vs human severity, and why triage still matters.

Jun 15, 2026

NVIDIA SkillSpector: Security Scanner for AI Agent Skills (2026)

Research shows 26.1% of agent skills contain vulnerabilities and 5.2% show likely malicious intent. NVIDIA's SkillSpector is an open-source scanner that catches them before they reach your agent.

Oct 3, 2026

Ponytail: The 152K-Star Skill That Makes AI Agents Write Less Code (Tested Claims, Install Guide)

Ponytail turns your AI coding agent into the laziest senior developer in the room: before writing code it climbs a ladder from do not build it, to reuse it, to the standard library, to a native feature. Its own agentic benchmark reports 54 percent less code and 100 percent safety. Here is how it works, what the numbers do and do not show, and how to try it without regret.