CyberGym evaluates agents against 1,507 real-world vulnerabilities across 188 software projects: given white-box source code and a vulnerability description, the agent must identify the flaw and generate a proof-of-concept input that triggers it. It sits opposite offense-focused benchmarks like ExploitBench and ExploitGym, which score whether a model can turn a known vulnerability into a working exploit chain — a model can lead CyberGym while trailing badly on ExploitBench, which is exactly what Z.ai's GLM-5.3 did at its August 2026 launch, and is why labs increasingly publish both scores rather than one blended "cybersecurity" number.