Learns a model that predicts only reward, value and policy, never observations, and plans with it. Reaches AlphaZero's play without being told the rules of the game.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in World Modelpredicting observations is capacity planning cannot use
requiresdoes not work without Monte Carlo Tree Searchit plans by searching inside the learned model
References
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model — Schrittwieser et al. — 2019 · link