Learning what an image or a signal means without anyone labelling it.
Colour is the family; a dashed line is the second member of it.
Keeps a queue of embeddings from previous batches, encoded by a slowly moving copy of the network, so the pool of negatives is decoupled from the bat… · 2019
An online network predicts the output of a slowly updated target copy of itself, with no negatives at all. Should collapse to a constant and does not… · 2020
Pulls two augmented views of one image together in embedding space and pushes other images apart. Learns without labels, and the augmentations quietl… · 2020
Cuts an image into fixed patches, embeds each as a token and runs a plain transformer. Discards the convolutional priors entirely, and beats them onc… · 2020
Trains an image encoder and a text encoder so that a picture and its caption land together. Gives a shared space where a class can be named rather th… · 2021
Self-distillation without labels: a student matches a momentum teacher across views, held away from collapse by centring and sharpening. Its attentio… · 2021
Hides most of the patches, encodes only the visible ones and reconstructs the rest with a light decoder. The high mask ratio is what makes it both a… · 2021
Scales that recipe on a curated corpus into features good enough to use frozen, with no fine-tuning, across depth, segmentation and retrieval. · 2023
Predicts the representation of a masked region rather than its pixels. The argument is that pixel reconstruction forces the model to model detail tha… · 2023
Extends the same idea to video, predicting the features of masked spacetime regions. Learns motion and object permanence from unlabelled footage alon… · 2024
Compresses input through a bottleneck and reconstructs it. The bottleneck is the whole mechanism: whatever survives it is what the data could not do…