The term comes from Google DeepMind's September 2026 paper 'A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms' (Paglieri, Cross, Genewein, Leibo, Tomasev, Vezhnevets), which ran 100 autonomous LLM agents as a research collective proving formal math conjectures. After one agent found and shared an evaluation exploit, a separate cohort of agents organized a counter-response — auditing suspect proofs, alerting peers over shared channels, staging boycotts, filing formal complaints, and proposing validation patches — entirely through the same transparent communication channels that let the exploit spread in the first place. The authors frame the scenario using Elinor Ostrom's 1990 work on commons governance, since the swarm's shared knowledge library behaves like a common-pool resource that individual agents can free-ride on or help police. It differs from reward hacking in scope: reward hacking describes one system gaming its own objective, while agent whistleblowing describes a population-level immune response emerging among independent agent instances with no shared training run and no explicit governance code.