Predicts the representation of a masked region rather than its pixels. The argument is that pixel reconstruction forces the model to model detail that is genuinely unpredictable and carries no meaning.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in Masked Autoencoderpixel targets spend capacity on unpredictable detail
Referenced by
extendsV-JEPA adds capability to thispredicts masked spacetime regions across video frames
References
Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture — Assran et al. — 2023 · link