The model optimizes a scalable objective such as next-token prediction, masked-token recovery, or contrastive matching. The resulting parameters provide reusable representations and capabilities that later stages can refine.
Pretraining teaches a model general patterns from a broad dataset before it is adapted to a specific task or behavior.
The model optimizes a scalable objective such as next-token prediction, masked-token recovery, or contrastive matching. The resulting parameters provide reusable representations and capabilities that later stages can refine.