Refits the compute-optimal frontier and finds parameters and tokens should grow together in roughly equal proportion. Implied the models of the day were several times too large for the data they saw.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in Scaling Lawsthe earlier fits left models too large for their data
References
Training Compute-Optimal Large Language Models — Hoffmann et al. — 2022 · link