Reward Model

AI and Machine Learning · Language Models · 2022 · reward-model.yaml

A model trained on pairwise human preferences to score responses, standing in for the objective nobody can write down. Also the component that gets gamed once the policy optimises against it hard enough.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G reward-model Reward Model rlhf RLHF reward-model->rlhf

This node

References