Pay the agent for reaching states its own predictor gets wrong, so novelty becomes a reward it can follow. What Montezuma's Revenge needed and what undirected exploration was never going to supply.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in Epsilon-Greedyrandom actions never chain into the long sequence a sparse reward needs
References
Curiosity-driven Exploration by Self-supervised Prediction — Pathak, Agrawal, Efros, Darrell — 2017 · link
Exploration by Random Network Distillation — Burda, Edwards, Storkey, Klimov — 2018 · link