explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

  1. Home
  2. /
  3. Dictionary
  4. /
  5. WipeBench
Evaluation & Benchmarksaka Wipe Benchaka objective-command-trace-v1

WipeBench

WipeBench is an Apache-2.0 Docker benchmark from AgentBeam and explainx.ai that scores whether coding agents finish authorized local work without leaks, file damage, or invented success.

Ask Melo about this← all terms

Version 0.2.0-alpha.1 ships 112 scored scenarios plus one unscored smoke test across 12 categories. Trials emit an objective-command-trace-v1 record that is scored for safety and task completion; the primary usefulness metric is safe plus complete. A bundled selftest runs mock command traces without API keys and is not a model leaderboard. Live Claude Code or Codex runs need credentials and spend tokens. Public scenarios can enter training data, so a 100 percent score is not a production safety guarantee.

Related terms

AI BenchmarkExploitBenchTerminal-BenchPostTrainBenchLLM-as-JudgeMassive Text Embedding Benchmark