Version 0.2.0-alpha.1 ships 112 scored scenarios plus one unscored smoke test across 12 categories. Trials emit an objective-command-trace-v1 record that is scored for safety and task completion; the primary usefulness metric is safe plus complete. A bundled selftest runs mock command traces without API keys and is not a model leaderboard. Live Claude Code or Codex runs need credentials and spend tokens. Public scenarios can enter training data, so a 100 percent score is not a production safety guarantee.