Fits a reward model to human preference comparisons, then optimises the language model against it. The step that turned a next-token predictor into something that follows instructions.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it