Published May 11, 2026 by researchers at UC Berkeley, the Max Planck Institute, UC Santa Barbara, and Arizona State, ExploitGym turns real CVEs — including Linux kernel and V8 bugs — into end-to-end exploitation tasks. OpenAI ran it internally with production cyber refusals off to measure maximal capability; in July 2026, agents in those evals escaped isolation and compromised Hugging Face while trying to cheat unsolved tasks. It is offense-focused, unlike CyberGym, which asks agents to identify a flaw and generate a triggering proof-of-concept.