AI and Machine Learning

World Models

Learning the dynamics of an environment so a policy can be trained inside the model.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G dreamer Dreamer world-model World Model dreamer->world-model trains the policy inside the model, not the environment genie Genie genie->world-model infers latent actions from video, so none need labelling v-jepa V-JEPA genie->v-jepa the dynamics are learned over video representations model-based-rl Model-Based RL q-learning Q-Learning model-based-rl->q-learning plans with learned dynamics instead of caching values monte-carlo-tree-search Monte Carlo Tree Search markov-decision-process Markov Decision Process monte-carlo-tree-search->markov-decision-process it searches the tree those transitions define upper-confidence-bound Upper Confidence Bound monte-carlo-tree-search->upper-confidence-bound every node of the tree is its own bandit problem muzero MuZero muzero->monte-carlo-tree-search it plans by searching inside the learned model muzero->world-model predicting observations is capacity planning cannot use world-model->model-based-rl the learned model is a latent-space simulator variational-autoencoder Variational Autoencoder world-model->variational-autoencoder the latent its dynamics run in comes from one

6 nodes

Monte Carlo Tree Search

Grows a search tree unevenly, spending rollouts where the value estimate is most uncertain. Needs a simulator it can query, which is exactly what mos… · 2006

World Model

Compresses observations into a latent with an autoencoder and learns a recurrent model of how that latent evolves. A controller can then be trained e… · 2018

Dreamer

Learns the latent dynamics and then optimises the policy by backpropagating through imagined rollouts inside it, never touching the environment durin… · 2019

MuZero

Learns a model that predicts only reward, value and policy, never observations, and plans with it. Reaches AlphaZero's play without being told the ru… · 2019

Genie

Trains on unlabelled internet video and infers a latent action space, producing a model that can be driven frame by frame. A playable environment lea… · 2024

Model-Based RL

Learns the environment's dynamics and plans against them, rather than caching what actions were worth. Buys sample efficiency and pays for it with wh…