AI and Machine Learning

Representation Learning

Learning what an image or a signal means without anyone labelling it.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it · hover one for the reasons on its edges, and to bring it forward
G autoencoder Autoencoder byol BYOL contrastive-learning Contrastive Learning byol->contrastive-learning clip CLIP clip->contrastive-learning dino DINO dino->byol vision-transformer Vision Transformer dino->vision-transformer dinov2 DINOv2 dinov2->dino jepa JEPA masked-autoencoder Masked Autoencoder jepa->masked-autoencoder masked-autoencoder->autoencoder masked-autoencoder->vision-transformer bert BERT masked-autoencoder->bert momentum-contrast Momentum Contrast momentum-contrast->contrastive-learning v-jepa V-JEPA v-jepa->jepa transformer Transformer vision-transformer->transformer
The 12 reasons on these edges, as text

11 nodes

Momentum Contrast

Keeps a queue of embeddings from previous batches, encoded by a slowly moving copy of the network, so the pool of negatives is decoupled from the bat… · 2019

BYOL

An online network predicts the output of a slowly updated target copy of itself, with no negatives at all. Should collapse to a constant and does not… · 2020

Contrastive Learning

Pulls two augmented views of one image together in embedding space and pushes other images apart. Learns without labels, and the augmentations quietl… · 2020

Vision Transformer

Cuts an image into fixed patches, embeds each as a token and runs a plain transformer. Discards the convolutional priors entirely, and beats them onc… · 2020

CLIP

Trains an image encoder and a text encoder so that a picture and its caption land together. Gives a shared space where a class can be named rather th… · 2021

DINO

Self-distillation without labels: a student matches a momentum teacher across views, held away from collapse by centring and sharpening. Its attentio… · 2021

Masked Autoencoder

Hides most of the patches, encodes only the visible ones and reconstructs the rest with a light decoder. The high mask ratio is what makes it both a… · 2021

DINOv2

Scales that recipe on a curated corpus into features good enough to use frozen, with no fine-tuning, across depth, segmentation and retrieval. · 2023

JEPA

Predicts the representation of a masked region rather than its pixels. The argument is that pixel reconstruction forces the model to model detail tha… · 2023

V-JEPA

Extends the same idea to video, predicting the features of masked spacetime regions. Learns motion and object permanence from unlabelled footage alon… · 2024

Autoencoder

Compresses input through a bottleneck and reconstructs it. The bottleneck is the whole mechanism: whatever survives it is what the data could not do…