Routes each token to a few of many expert subnetworks, so parameter count grows without the arithmetic growing with it. The routing is the hard part, and load balance is a training objective of its own.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in Transformercompute grows with capacity if every weight is used
References
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — Shazeer et al. — 2017 · link