MuZero

AI and Machine Learning · World Models · 2019 · muzero.yaml

Learns a model that predicts only reward, value and policy, never observations, and plans with it. Reaches AlphaZero's play without being told the rules of the game.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G monte-carlo-tree-search Monte Carlo Tree Search muzero MuZero muzero->monte-carlo-tree-search it plans by searching inside the learned model world-model World Model muzero->world-model predicting observations is capacity planning cannot use

This node

References