Supervised Fine-Tuning

AI and Machine Learning · Language Models · 2021 · also: SFT, Instruction tuning · supervised-fine-tuning.yaml

Continues next-token training on curated demonstrations of the behaviour wanted. Cheap, stable, and bounded by how well the desired behaviour can be written down.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G autoregressive-language-model Autoregressive Language Model rlhf RLHF supervised-fine-tuning Supervised Fine- Tuning rlhf->supervised-fine-tuning the policy is initialised from a supervised pass supervised-fine-tuning->autoregressive-language-model trains on demonstrations of the behaviour wanted

This node

Referenced by

References