Built to correlate with human preference rankings from Chatbot Arena without requiring live human voting for every new model, Arena-Hard selects prompts that best separate strong models and scores head-to-head win rates via an LLM judge. Like other LLM-as-judge benchmarks, it inherits the judge model's own biases, and its win-rate scale is nearing saturation at the frontier.