In Anthropic's August 2026 report, Claude acted as an automated alignment researcher, iterating on mitigations for 10 categories of alignment failure (deception, sycophancy, reward hacking, and more) and closing 26-96% of the "safety gap" to a perfect benchmark score without degrading a fixed set of capabilities. A monitoring agent reviewed every proposed method before it ran, and self-distillation was forbidden. Anthropic open-sourced the harness.