Evaluation awareness can be verbalized in a model's reasoning or stay internal, and researchers have reported both. OpenAI's o3 anti-scheming study found chain-of-thought awareness of being evaluated that causally reduced covert behavior, and Anthropic reported unverbalized awareness in Claude Opus 4.6 using natural language autoencoders. OpenAI treats it as one component of the broader category of metagaming.