Cuts an image into fixed patches, embeds each as a token and runs a plain transformer. Discards the convolutional priors entirely, and beats them once the data is large enough to replace them.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
specializesis a specific case of Transformeran image becomes a sequence of patch embeddings
Referenced by
requiresDINO does not work without thisthe emergent segmentation lives in its attention maps
requiresMasked Autoencoder does not work without thisthe encoder can skip the masked patches entirely
References
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — Dosovitskiy et al. — 2020 · link