Speculative Decoding

AI and Machine Learning · Efficiency and Deployment · 2022 · speculative-decoding.yaml

A small draft model proposes several tokens and the large model verifies them in one batched pass, keeping the longest correct prefix. Exact: the output distribution is unchanged.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G kv-cache KV Cache speculative-decoding Speculative Decoding speculative-decoding->kv-cache decoding stays serial, one token per pass

This node

References