CLIP

AI and Machine Learning · Representation Learning · 2021 · clip.yaml

Trains an image encoder and a text encoder so that a picture and its caption land together. Gives a shared space where a class can be named rather than trained.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G clip CLIP contrastive-learning Contrastive Learning clip->contrastive-learning the positive pair is an image and its own caption latent-diffusion Latent Diffusion latent-diffusion->clip the text conditioning comes from an image-text encoder

This node

Referenced by

References