A checkpoint is a persisted snapshot of an agent's complete state at a specific point in its execution — including conversation history, tool results, scratchpad contents, step count, and any accumulated artifacts. Checkpoints enable fault tolerance: if an agent crashes, times out, or is preempted, it can resume from the last checkpoint rather than restarting the entire task. They are essential infrastructure for long-running agents and are analogous to checkpoints in distributed computing and ML training.