Trains on unlabelled internet video and infers a latent action space, producing a model that can be driven frame by frame. A playable environment learned from footage of environments.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
extendsadds capability to World Modelinfers latent actions from video, so none need labelling
requiresdoes not work without V-JEPAthe dynamics are learned over video representations
References
Genie: Generative Interactive Environments — Bruce et al. — 2024 · link