Trains a small model against the large model's full output distribution rather than the hard labels. The relative probabilities of the wrong answers carry most of the transferred information.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
requiresdoes not work without Cross-Entropy Lossthe target is the teacher's whole distribution
References
Distilling the Knowledge in a Neural Network — Hinton, Vinyals, Dean — 2015 · link