explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

  1. Home
  2. /
  3. Dictionary
  4. /
  5. CyberGym
Evaluation & Benchmarks

CyberGym

CyberGym is a large-scale benchmark that tests whether an AI agent can find and validate real software vulnerabilities from source code, measuring defensive security capability rather than offensive exploit generation.

Ask Melo about this← all terms

CyberGym evaluates agents against 1,507 real-world vulnerabilities across 188 software projects: given white-box source code and a vulnerability description, the agent must identify the flaw and generate a proof-of-concept input that triggers it. It sits opposite offense-focused benchmarks like ExploitBench and ExploitGym, which score whether a model can turn a known vulnerability into a working exploit chain — a model can lead CyberGym while trailing badly on ExploitBench, which is exactly what Z.ai's GLM-5.3 did at its August 2026 launch, and is why labs increasingly publish both scores rather than one blended "cybersecurity" number.

Related terms

AI BenchmarkExploitBenchRed TeamingTruthfulQAMetricMATH-500

Where CyberGym comes up

  • GLM-5.3's 84.5% CyberGym Score Isn't Verified Yet — What "Opening to Researchers" Really Means
  • Hugging Face Open Alignment Team: What Builders Can Use Today
  • Hugging Face's security.txt Has a Note for AI Agents — And It's Not a Joke
  • Abliteration.ai Hosts an Uncensored GLM-5.3 for Offensive Cyber Work