Held-out evaluation estimates how learned patterns transfer to the intended deployment distribution. Distribution shifts, leakage, and repeated tuning can make apparent generalization differ from real-world behavior.
Generalization is a model's ability to perform well on relevant examples it did not encounter during training.
Held-out evaluation estimates how learned patterns transfer to the intended deployment distribution. Distribution shifts, leakage, and repeated tuning can make apparent generalization differ from real-world behavior.