Distributional Reinforcement Learning

AI and Machine Learning · Reinforcement Learning · 2017 · also: C51, Categorical DQN · distributional-rl.yaml

Predicts the whole distribution of returns as a histogram over fixed atoms, and applies the Bellman backup to that. Actions are still chosen by its mean, so the gain is not risk sensitivity: the richer target is simply easier for a network to fit.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G distributional-rl Distributional Reinforcement Learning q-learning Q-Learning distributional-rl->q-learning an expectation discards the shape of the returns that produced it rainbow Rainbow rainbow->distributional-rl it pays off only late in training, where shorter papers stopped looking

This node

Referenced by

References