Public benchmarks get quoted in blogs, scraped into web-scale pretraining corpora, or targeted during iterative fine-tuning against a leaderboard. Once a model has effectively seen a test, high scores measure memorization or overfitting rather than capability on unseen tasks. Google DeepMind's August 2026 double-blind evaluation pilot is one structural response: run tests inside a cryptographic enclave so the model owner never sees confidential prompts.