Normalise each activation across the batch, then rescale by learned parameters. Allowed much higher learning rates, which is what made very deep networks trainable in practice.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in Stochastic Gradient Descenteach layer's input distribution shifts as earlier layers update
Referenced by
correctsLayer Normalization fixes a defect in thisbatch statistics make training and inference disagree
References
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift — Ioffe, Szegedy — 2015 · link