Cases may perturb normal data, exploit instruction conflicts, or target tool and application boundaries. Repeating the tests after fixes checks whether mitigations generalize beyond a single prompt.
Adversarial testing evaluates a model or system with inputs intentionally designed to trigger errors or bypass controls.
Cases may perturb normal data, exploit instruction conflicts, or target tool and application boundaries. Repeating the tests after fixes checks whether mitigations generalize beyond a single prompt.