A small draft model proposes several tokens and the large model verifies them in one batched pass, keeping the longest correct prefix. Exact: the output distribution is unchanged.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in KV Cachedecoding stays serial, one token per pass
References
Fast Inference from Transformers via Speculative Decoding — Leviathan, Kalman, Matias — 2022 · link