Direct Preference Optimization

AI and Machine Learning · Language Models · 2023 · also: DPO · direct-preference-optimization.yaml

Shows the preference objective can be optimised as a classification loss on the model itself, with no reward model and no rollouts.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G direct-preference-optimization Direct Preference Optimization rlhf RLHF direct-preference-optimization->rlhf a reward model and an RL loop to fit a preference dataset

This node

References