Prompt injection is the problem that does not go away. As agents read web pages, emails and repository files, any of that text can try to hijack them. The standard defenses are least privilege, sandboxing and human approval, and increasingly a screening model that looks at every input and tool call and flags the suspicious ones. On October 5, 2026, Superagent's Alan Zabihi released one as open weights: Security-One, a 27B decision model under Apache 2.0.
The headline number is striking, 599 of 600 prompt-injection attacks caught on one benchmark, and the less quoted numbers are more instructive. This post covers what the model is, how to use it, what the benchmarks say and do not say, and how to fit it into an agent defense without over-trusting it.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is it? | A 27B decision model returning unsafe probabilities for security-relevant events. |
| What can it screen? | Prompts, agent tool calls, code changes and alerts. |
| License? | Apache 2.0, weights on Hugging Face, plus an API. |
| Best number? | 599 of 600 attacks on BIPIA, 1 of 200 benign flagged, at threshold 0.70. |
| Weaker numbers? | 47 of 60 on Deepset; higher false positives than Jev on NotInject. |
| Can it enforce? | No. The vendor says use it for screening. |
| Price? | Not specified for the API in the coverage we reviewed. |
What a decision model for security does
We explained the category in our decision models guide and covered a commercial example in Liquid AI's d1. A decision model answers a closed question with probabilities in one forward pass, with no generated text. For security triage that means you hand it an event and a question, such as "is this input an attempt to override the agent's instructions?", and get back a probability for each answer.
Security-One is built for always-on use. Because it generates no tokens, it is cheap and quick enough to run on every prompt, tool call or alert, not only on a sample. Superagent describes it as a routing layer that identifies suspicious activity and sends those events to human analysts or more powerful models for deeper investigation. It follows the "System One" decision-model idea, popularized by the Jev family, of fast intuitive classification with deliberate review reserved for exceptions.
The benchmark numbers, in context
Superagent reports results against public prompt-injection datasets. Here they are in one place.
| Dataset | Result | Notes |
|---|---|---|
| BIPIA | 599 of 600 attacks detected (99.83%); 1 of 200 benign flagged (0.50%) | Threshold 0.70 |
| Deepset | 47 of 60 attacks detected (78.33%); 0 of 56 benign flagged | Same model, different dataset |
| NotInject | 0 of 339 false positives, but a 12.39% false-positive rate versus 2.36% for Jev in the vendor's comparison | Over-flagging benign text that looks suspicious |
The spread across datasets is the real finding. A 99.83 percent catch rate on one benchmark and 78 percent on another says that performance depends on how attacks are written, and that no single number describes how it will behave on your traffic. Public datasets are also known to be imperfect proxies: attackers adapt, benchmarks age, and models can be tuned to them. The over-flagging result matters too, because a screening model that blocks benign requests teaches users to bypass it.
Why you must tune the threshold on your own traffic
The benchmark threshold is 0.70 unsafe probability. That is a trade-off, not a truth. Lower the threshold and you catch more attacks and flag more benign inputs. Raise it and the reverse happens. The right setting depends on:
- the cost of a missed attack in your system (a read-only chatbot versus an agent that can send money),
- the cost of a false alarm (a blocked customer versus a delayed review),
- the base rate of attacks in your traffic, which is usually far lower than in a benchmark, so even a small false-positive rate can swamp your analysts.
Superagent provides a calibration recipe and says explicitly that thresholds must be tuned on your own data. Do that with a labeled sample that includes your real benign traffic, especially the odd-looking but legitimate inputs, like pasted logs or security discussions.
Where it fits in an agent defense
A screening model is one layer. A sensible stack looks like this:
- Least privilege. Give agents the narrowest credentials that work. A hijacked agent that cannot do anything dangerous is a nuisance, not an incident.
- Sandboxing. Run tools in isolated environments, as in our guides to agent sandbox isolation and the Codex Auto-review reviewer.
- Screening. Run Security-One, or a similar model, on inbound content and on each proposed tool call.
- Escalation. Send medium-probability events to a stronger model or a human, and block the highest ones.
- Logging. Record the probability and the decision, so you can audit and retune.
The reason to screen tool calls and not only inputs is that injections usually succeed through an action: an agent tries to send data somewhere or run a command. Catching the action is more robust than guessing every phrasing of the attack. For real-world examples of what goes wrong, see our coverage of prompt injection in GitHub agentic workflows and the rise of AI-driven web traffic.
How to evaluate it this week
- Collect two sets. Attack examples that resemble your threat model, and a large sample of benign inputs from your logs.
- Run Security-One and plot the distribution of unsafe probabilities for each set.
- Choose a threshold that meets your false-positive budget, then record the catch rate at that threshold.
- Red-team it. Try rephrased attacks, other languages, encoded text and multi-step instructions. Note which pass.
- Test tool-call screening by showing it proposed actions, not just prompts.
- Measure latency and cost in your serving setup. A 27B model needs a capable GPU, so compare with the API.
- Plan for drift. Re-run the evaluation periodically with fresh attacks.
Limits to keep in mind
- It is a prediction model. Superagent itself says it can be confidently wrong.
- Dataset-dependent performance. The gap between BIPIA and Deepset shows how uneven it can be.
- Over-flagging. A higher false-positive rate on benign-but-odd text can hurt usability.
- Unclear base model. The vendor post does not name the base model in the content we reviewed, so check the model card for lineage and licensing of the training data.
- Adaptive attackers. Anyone who knows you use a given screener can test against it. Treat it as a speed bump, not a wall.
- No pricing details for the hosted API in the coverage we reviewed.
A worked example of threshold choice
Suppose your agent processes 100,000 inbound items a day and real attacks are rare, say 20 a day. A screener with a 0.5 percent false-positive rate flags about 500 benign items daily, which means 25 false alarms for each real attack even if it catches every attack. If each review costs a minute, that is more than eight analyst-hours a day. Raise the threshold and the false alarms drop, but so does the catch rate. The point of the exercise is that base rates, not benchmark accuracy, drive operational cost. Measure your own base rate, set a review budget, and choose the threshold that fits it, then monitor drift as attackers adapt.
It also helps to separate tiers. Events above a high threshold can be blocked automatically, events in a middle band can be sent to a stronger model for a second opinion, and events below it can pass with logging. A three-tier policy gets more value from a cheap screener than a single cutoff does.
Logging for audit and retuning
Whatever threshold you choose, log the probability, the input hash, the decision and the eventual human verdict for escalated items. Those records let you measure real precision, find recurring false-alarm patterns and build a labeled set for retuning. They also give you evidence if an incident occurs and you need to show how the screening layer behaved. Mind privacy: store hashes or redacted excerpts when the inputs may contain personal data.
What this means for what you build or pay
If you ship agents that read untrusted content, a cheap open screening model is a reasonable addition, and an Apache 2.0 license means you can run it yourself and inspect the behavior. Budget the work to calibrate it, because an uncalibrated threshold is the most common way these tools disappoint. And keep the main defenses where they belong: scoped credentials, isolation and human review for irreversible actions. A good screener reduces the number of events that reach those layers. It does not replace them.
Related reading
- What are decision models? AI classifiers guide
- Liquid AI d1: a decision model with vision
- Codex Auto-review: a reviewer agent for approvals
- Prompt injection in GitHub agentic workflows
- AI traffic overtakes humans and prompt injection
- Agent sandbox isolation: five things to know
- Perplexity's decision API
Primary: Superagent, "Introducing Security-One: a decision model for security" (superagent.sh) · HuggingNews and Datastudios coverage (October 5 to 6, 2026)
Details are accurate as of October 6, 2026 and come from vendor materials and press reports. Benchmark results are self-reported, thresholds must be tuned on your own traffic, and the model is a screening aid, not a security guarantee.
