A model selects answers under a specified zero-shot or few-shot prompt, and accuracy is aggregated across subjects. The score measures performance on this question format, not general intelligence or deployment reliability.
MMLU is a benchmark that tests language models with multiple-choice questions drawn from many academic and professional subjects.
A model selects answers under a specified zero-shot or few-shot prompt, and accuracy is aggregated across subjects. The score measures performance on this question format, not general intelligence or deployment reliability.