Preference learning is any training method that uses human or model comparisons between outputs (A is better than B) rather than absolute scores, including RLHF, DPO, and KTO. This approach captures nuanced quality judgments that are easier for humans to express as comparisons than as numerical ratings.