A model answers an initial prompt and a follow-up that depends on the prior turn. A model-based judge commonly scores the responses against a rubric, making judge choice and prompt version part of the setup. Human checks are useful for calibrating judge preferences and failure modes.