Learning the dynamics of an environment so a policy can be trained inside the model.
Colour is the family; a dashed line is the second member of it.
Grows a search tree unevenly, spending rollouts where the value estimate is most uncertain. Needs a simulator it can query, which is exactly what mos… · 2006
Compresses observations into a latent with an autoencoder and learns a recurrent model of how that latent evolves. A controller can then be trained e… · 2018
Learns the latent dynamics and then optimises the policy by backpropagating through imagined rollouts inside it, never touching the environment durin… · 2019
Learns a model that predicts only reward, value and policy, never observations, and plans with it. Reaches AlphaZero's play without being told the ru… · 2019
Trains on unlabelled internet video and infers a latent action space, producing a model that can be driven frame by frame. A playable environment lea… · 2024
Learns the environment's dynamics and plans against them, rather than caching what actions were worth. Buys sample efficiency and pays for it with wh…