Synthetic data generation is using a model to produce training examples — prompts, completions, or preference pairs — that augment or replace human-written data, especially for scarce domains. This technique has become essential for scaling training data in specialized areas where human annotation is expensive or insufficient.