Predicts the whole distribution of returns as a histogram over fixed atoms, and applies the Bellman backup to that. Actions are still chosen by its mean, so the gain is not risk sensitivity: the richer target is simply easier for a network to fit.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
extendsadds capability to Q-Learningan expectation discards the shape of the returns that produced it
Referenced by
requiresRainbow does not work without thisit pays off only late in training, where shorter papers stopped looking
References
A Distributional Perspective on Reinforcement Learning — Bellemare, Dabney, Munos — 2017 · link