Runs the diffusion process in the compressed latent space of an autoencoder rather than on pixels. The reason text-to-image generation fits on consumer hardware.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in Diffusion Modeldenoising at pixel resolution is prohibitively expensive
requiresdoes not work without Variational Autoencoderthe latent space it works in comes from one
requiresdoes not work without CLIPthe text conditioning comes from an image-text encoder
References
High-Resolution Image Synthesis with Latent Diffusion Models — Rombach, Blattmann, Lorenz, Esser, Ommer — 2021 · link