Trains an image encoder and a text encoder so that a picture and its caption land together. Gives a shared space where a class can be named rather than trained.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
specializesis a specific case of Contrastive Learningthe positive pair is an image and its own caption
Referenced by
requiresLatent Diffusion does not work without thisthe text conditioning comes from an image-text encoder
References
Learning Transferable Visual Models From Natural Language Supervision — Radford et al. — 2021 · link