Vision Transformer

AI and Machine Learning · Representation Learning · 2020 · also: ViT · vision-transformer.yaml

Cuts an image into fixed patches, embeds each as a token and runs a plain transformer. Discards the convolutional priors entirely, and beats them once the data is large enough to replace them.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G dino DINO vision-transformer Vision Transformer dino->vision-transformer the emergent segmentation lives in its attention maps masked-autoencoder Masked Autoencoder masked-autoencoder->vision-transformer the encoder can skip the masked patches entirely transformer Transformer vision-transformer->transformer an image becomes a sequence of patch embeddings

This node

Referenced by

References