OpenAI uses the term for reasoning about feedback or oversight mechanisms outside a scenario's narrative, regardless of whether the model is in training, evaluation or deployment. It is broader than evaluation awareness because the model need not identify which distribution an input came from or be right about how it is graded. OpenAI measures verbalized metagaming by grading chains of thought, and found it rose during capabilities-focused RL on o3.