Extends the same idea to video, predicting the features of masked spacetime regions. Learns motion and object permanence from unlabelled footage alone.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
extendsadds capability to JEPApredicts masked spacetime regions across video frames
Referenced by
requiresGenie does not work without thisthe dynamics are learned over video representations
References
Revisiting Feature Prediction for Learning Visual Representations from Video — Bardes et al. — 2024 · link