Drawn from the same competition-math source as the full MATH benchmark, the 500-problem subset is cheap enough to run in every evaluation cycle. It is largely saturated at the 2026 frontier, with the best recorded scores leaving only a few points of headroom on its 100-point scale — a state where the metric struggles to separate strong models from each other.