A checkpoint is a saved copy of model weights, optimizer state, and training metadata at a point during training, enabling resumption after failure or selection of the best-performing snapshot. Regular checkpointing is critical for large training runs that can take weeks and cost millions of dollars, as hardware failures are inevitable at scale.