A benchmark is a standardized test set and evaluation protocol used to compare model performance — results are only meaningful when methodology, splits, and metrics are held constant. Popular benchmarks include MMLU for broad knowledge, HumanEval for code generation, and GPQA for graduate-level reasoning. Benchmark saturation (models nearing perfect scores) drives the community to create harder evaluations. Criticism of benchmarks focuses on data contamination, narrow task coverage, and the gap between benchmark scores and real-world usefulness.