Hides most of the patches, encodes only the visible ones and reconstructs the rest with a light decoder. The high mask ratio is what makes it both a hard task and a cheap one.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
specializesis a specific case of Autoencoderreconstructs patches hidden from the encoder
extendsadds capability to BERTmasked prediction moved from tokens to image patches
requiresdoes not work without Vision Transformerthe encoder can skip the masked patches entirely
Referenced by
correctsJEPA fixes a defect in thispixel targets spend capacity on unpredictable detail
References
Masked Autoencoders Are Scalable Vision Learners — He, Chen, Xie, Li, Dollar, Girshick — 2021 · link