Learns the latent dynamics and then optimises the policy by backpropagating through imagined rollouts inside it, never touching the environment during that step.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
extendsadds capability to World Modeltrains the policy inside the model, not the environment
References
Dream to Control: Learning Behaviors by Latent Imagination — Hafner, Lillicrap, Ba, Norouzi — 2019 · link
Mastering Diverse Domains through World Models — Hafner, Pasukonis, Ba, Lillicrap — 2023 · link