The evaluator receives a rubric and may compare outputs with references or with each other. Calibration against human judgments, order controls, and bias checks are needed because the judge can share model errors or favor superficial traits.
Reported agreement with human raters spans a wide band depending on task structure and rubric quality. Mello et al. (LAK'25) found GPT-4 with prompt engineering reached QWK ~0.94, rivaling fine-tuned BERT and outperforming SVM (0.80). AutoSCORE (AAAI'26) showed that an explicit rubric-recognition step before scoring fixes single-prompt inconsistency. Technology, Knowledge and Learning (2025) found the human–LLM alignment gap shrinks substantially when high-quality rubrics are supplied. Emirtekin et al. (JCAL 2026) reported moderate agreement (QWK 0.585–0.640) from GPT-4o, Gemini 2.0 Flash, DeepSeek V3, and LLaMA 3.3 on Bloom's-taxonomy rubrics — below human inter-rater reliability (ICC 0.667–0.800). Flodén (2025) found ~70% of ChatGPT's Master's-exam scores fell within 10% of human scores. Treat sub-0.6 QWK grading as formative-only unless you validate on your own rubric.