explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

  1. Home
  2. /
  3. Dictionary
  4. /
  5. LLM as a Judge
Evaluation & Benchmarksaka Model-Based Evaluationaka LLM-as-Judge

LLM as a Judge

LLM as a judge uses a language model to score, rank, or critique other model outputs.

Ask Melo about this← all terms

The evaluator receives a rubric and may compare outputs with references or with each other. Calibration against human judgments, order controls, and bias checks are needed because the judge can share model errors or favor superficial traits.

Reported agreement with human raters spans a wide band depending on task structure and rubric quality. Mello et al. (LAK'25) found GPT-4 with prompt engineering reached QWK ~0.94, rivaling fine-tuned BERT and outperforming SVM (0.80). AutoSCORE (AAAI'26) showed that an explicit rubric-recognition step before scoring fixes single-prompt inconsistency. Technology, Knowledge and Learning (2025) found the human–LLM alignment gap shrinks substantially when high-quality rubrics are supplied. Emirtekin et al. (JCAL 2026) reported moderate agreement (QWK 0.585–0.640) from GPT-4o, Gemini 2.0 Flash, DeepSeek V3, and LLaMA 3.3 on Bloom's-taxonomy rubrics — below human inter-rater reliability (ICC 0.667–0.800). Flodén (2025) found ~70% of ChatGPT's Master's-exam scores fell within 10% of human scores. Treat sub-0.6 QWK grading as formative-only unless you validate on your own rubric.

Related terms

BLEU ScoreHuman EvaluationRed TeamingAdversarial TestingMeloLLM