Its objective increases the relative likelihood of preferred responses while comparing the policy with a reference model. The reference term limits drift and turns preference learning into a supervised-style optimization problem.
Direct preference optimization trains a model from preferred and rejected response pairs without first fitting a separate reward model.
Its objective increases the relative likelihood of preferred responses while comparing the policy with a reference model. The reference term limits drift and turns preference learning into a supervised-style optimization problem.