Direct Preference Optimization
AI and Machine Learning · Language Models · 2023 · also: DPO · direct-preference-optimization.yaml
Shows the preference objective can be optimised as a classification loss on the model itself, with no reward model and no rollouts.
- supersedescorrects · extends
- classifiesspecializes · part-of
- substitutes forapproximates · alternative-to
- depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
This node
- correctsfixes a defect in RLHFa reward model and an RL loop to fit a preference dataset
References
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model — Rafailov et al. — 2023 · link