Attention

AI and Machine Learning · Sequence Models · 2014 · attention.yaml

Compute a weighted sum over all positions, with weights from a learned compatibility score. Every position can reach every other in one hop, regardless of distance.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G attention Attention recurrent-network Recurrent Network attention->recurrent-network a fixed-size hidden state cannot carry a whole sequence flash-attention FlashAttention flash-attention->attention it is memory-bandwidth bound, not compute bound transformer Transformer transformer->attention it is the only operation left that mixes positions

This node

Referenced by

References