If you want to know how well an AI can attack a network, you can wait for criminals to show you, or hire the AI to do it first. Armadin, led by Kevin Mandia, the founder of Mandiant, sells the second option. This week it announced a $255.5 million Series B at a valuation above $2.5 billion, and with it a set of numbers that stood out: more than 90 zero-day vulnerabilities at Fortune 500 companies since January, found by swarms of AI agents attacking from the outside.
This post explains what Armadin does, what the numbers do and do not show, how autonomous attack swarms differ from traditional penetration testing, and the questions a defender should ask. Everything about Armadin's results is company-reported. We found no independent verification.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is it? | A swarm of AI agents that attack your network from outside to find real, exploitable flaws. |
| Who runs it? | Kevin Mandia, founder of Mandiant. |
| Headline claim? | 90+ zero-days at Fortune 500 firms since January. |
| August exercise? | 1,300 attacks, 26,000 agents, ~17M actions, 25,000+ services, 238 findings, 38 attack paths. |
| Funding? | $255.5M Series B, valuation above $2.5B, per reports. |
| Verified independently? | Not that we found. |
| Is it a good idea? | Plausibly useful, if tightly scoped and authorized. |
What Armadin says it does
The pitch, as described in investor and press materials, is continuous, autonomous offensive testing. Armadin's agents map every service, route and system from the outside, then re-attack whenever something changes. The platform deploys an "autonomous swarm" of specialized agents that "reason like a skilled adversary across the attack surface, chaining individually low-severity weaknesses into validated kill chains." The stated long-term goal is not only to find flaws but to patch them autonomously.
That differs from a traditional penetration test in three ways. Human tests happen periodically, perhaps once or twice a year, while a swarm can run continuously. Human testers are scarce and expensive, while agents can be scaled by parallelism. And human testers bring judgment about what matters, which swarms must approximate by chaining findings into demonstrated paths.
The numbers
| Claim | Figure | Context |
|---|---|---|
| Zero-days found | 90+ | Since January 2026, at Fortune 500 companies |
| Method | Black-box, from the internet | No source code access |
| August exercise attacks | 1,300 | Three-day exercise |
| Agents involved | 26,000 | Across the exercise |
| Offensive actions | About 17 million | Total |
| Services targeted | 25,000+ | Services in scope |
| Findings | 238 | 98 reported as significant |
| Attack paths | 38 | Validated chains |
Two cautions on reading these. First, "zero-day" is used loosely. A true zero-day is a vulnerability unknown to the vendor. Many findings in a network test are misconfigurations, exposed services or weak credentials, which are serious but are not zero-days in the strict sense. The coverage does not break down how many of the 90 are vendor-product bugs versus customer-specific flaws. Second, counts are not severity. Ninety low-impact findings and ninety critical ones are very different outcomes. The August exercise, with 98 of 238 findings called significant, is more informative because it reports a ratio, but it is still self-reported.
Why chaining matters
The most interesting claim is chaining. Real breaches rarely depend on one dramatic flaw. They combine a minor information leak, a weak internal permission and a forgotten service into a path to something valuable. Human testers do this, but it is time-consuming, and automated scanners usually report each issue separately, leaving the chain to the reader.
If agents can reliably reason across findings and validate a chain by actually walking it, that is useful, because it ranks fixes by real risk rather than by a checklist. It also reflects a broader trend: frontier models are getting better at multi-step offensive reasoning, which cuts both ways. Mistral's recent launch pitched its model as strong at defensive cyber work without refusals, as covered in our Mistral Large 4 post, and we have tracked attackers using local models in offline operations.
The dual-use problem
Autonomous offensive agents are dual-use by nature. The same swarm that finds a customer's flaws could, in the wrong hands, attack. So responsible products rely on tight controls: verified customers, signed scope documents, enforced boundaries on what can be touched, rate limits to avoid disrupting production, and kill switches. We have already seen how an autonomous agent can wander beyond its intended scope, as in the Hugging Face autonomous agent breach. That is why a buyer should ask about containment as closely as about capability.
There is also the downstream question of disclosure. If an agent finds a zero-day in a third-party product while testing a customer, someone must report it to the vendor, and the timeline and handling matter. A swarm that finds dozens of such issues creates a large coordination burden. The coverage did not describe Armadin's disclosure process in detail.
Questions to ask before buying autonomous red-teaming
- Scope control. How is the target scope defined and enforced technically? Can the agents ever touch systems outside it?
- Safety for production. How does it avoid outages, data corruption or lockouts? Can I run it in a staging mirror first?
- Kill switch. Can I stop all agents instantly, and how fast does that take effect?
- Verification. How are findings validated, and what is the false-positive rate? Can I see the evidence for each?
- Disclosure. Who handles reporting to third-party vendors, and on what timeline?
- Data handling. What does the platform store about my network, credentials found and traffic?
- Human review. Where do experts review before anything is exploited or patched?
- Autonomous patching. If it can change my systems, what approvals and rollbacks exist?
- Benchmarks. What independent evidence supports the capability claims, such as third-party tests or customer references?
- Legal. Do contracts and authorizations cover testing of third-party infrastructure that my systems depend on?
Context: benchmarks versus real networks
Public benchmarks show offensive capability climbing fast. Mistral's launch cited an 82 percent score on a test that asks a model to reproduce and patch a real vulnerability, and cybersecurity benchmarks overall have risen sharply this year. But a benchmark is a controlled task, while a Fortune 500 network is messy, large and full of surprises. A swarm's real-world hit rate depends on reconnaissance quality, handling of unusual systems and avoiding noise. That gap is why customer-verifiable results matter more than leaderboard numbers. For a related caution about taking security claims at face value, see our discussion of verifying AI math and science claims, where the same logic applies: ask what artifact exists and who checked it.
What defenders should do regardless of vendor
Whether or not you buy an autonomous testing product, the same defensive basics determine how much a swarm, or an attacker, can find. Keep an accurate inventory of externally exposed services, since agents that map everything from the outside will find forgotten systems first. Patch internet-facing software quickly, and remove or restrict anything that does not need to be public. Enforce least privilege and strong authentication so that a chain of small weaknesses ends early. Monitor for unusual enumeration and lateral movement, because a swarm that re-attacks whenever something changes looks different from ordinary traffic. And rehearse incident response, since the faster attackers get at finding paths, the less time defenders have between discovery and exploitation.
A useful exercise is to ask what your own team would find if it pointed a capable agent at your external surface with a narrow, authorized scope for one afternoon. Even a small, tightly controlled experiment tends to surface exposed admin panels, stale subdomains and credentials in old repositories. Document what you find, fix the cheapest high-impact items first and repeat the exercise after changes to confirm the paths are closed.
What this means for what you build or pay
If you run a security program, continuous automated testing is attractive because the attack surface changes daily, and annual tests go stale. Treat vendors in this space as tools that raise your testing frequency, not as replacements for your own hygiene: patching, least privilege, monitoring and incident response still decide outcomes. If you build software, expect your external attack surface to be probed more often and more cleverly, by defenders and attackers alike, and consider running your own agent-based tests with strict scope. And if you read the headline number, remember it is a company claim awaiting outside confirmation.
Related reading
- Mistral Large 4: cyber capability and open weights
- Security-One: an open decision model for security triage
- Hugging Face autonomous agent breach
- Kimsuky and offline local AI operations
- OpenAI Daybreak: Codex security and cyber defense
- How to check an unverified AI claim
Primary: Armadin Series B announcement (PR Newswire) and SecurityWeek coverage · a16z post on Armadin · interviews with Kevin Mandia
Details are accurate as of October 7, 2026 and come from company announcements and press coverage. The zero-day count, exercise figures and valuation are company-reported or reported by press, and we found no independent verification.
