Runs the diffusion process in the compressed latent space of an autoencoder rather than on pixels. The reason text-to-image generation fits on consumer hardware.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in Diffusion Modeldenoising at pixel resolution is prohibitively expensive
requiresdoes not work without Variational Autoencoderthe latent space it works in comes from one
References
High-Resolution Image Synthesis with Latent Diffusion Models — Rombach, Blattmann, Lorenz, Esser, Ommer — 2021 · link