FlashAttention

AI and Machine Learning · Sequence Models · 2022 · flash-attention.yaml

Tiles attention so the score matrix is never written to memory, recomputing it in fast on-chip storage instead. Exact, not approximate: the same result, bounded by arithmetic rather than bandwidth.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G attention Attention cache-hierarchy Cache Hierarchy flash-attention FlashAttention flash-attention->attention it is memory-bandwidth bound, not compute bound flash-attention->cache-hierarchy it tiles attention to fit in on-chip memory

This node

References