ARC-AGI-3 evaluates agents in interactive, game-like environments with no stated rules or goals, so raw pass/fail alone does not capture how efficiently an agent explored and solved a level. RHAE aggregates task completion with per-level action efficiency relative to first-time human performance across levels and environments, producing a single score comparable across agents that differ in how many environment actions they needed to reach the same outcome.