Safety teams rely on reading a model's extended thinking as an early-warning system, but monitorability breaks down if reasoning is unfaithful (the model acts on considerations it doesn't verbalize) or if training pressure is applied directly to reasoning text, teaching the model to make thoughts look acceptable rather than be accurate. Anthropic's August 2026 Risk Report disclosed that chain-of-thought was unintentionally exposed to reinforcement-learning reward calculation on up to 5.1% of training episodes for Claude Mythos Preview, and treats keeping reasoning unoptimized-against as central to its ability to detect covert misalignment.