Knowledge Distillation

AI and Machine Learning · Efficiency and Deployment · 2015 · knowledge-distillation.yaml

Trains a small model against the large model's full output distribution rather than the hard labels. The relative probabilities of the wrong answers carry most of the transferred information.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G cross-entropy-loss Cross-Entropy Loss knowledge-distillation Knowledge Distillation knowledge-distillation->cross-entropy-loss the target is the teacher's whole distribution

This node

References