Splits the head into a state-value stream and a per-action advantage stream, then adds them back. How good a state is gets learned once instead of separately inside every action's estimate.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in Deep Q-Networkin most states the action barely matters, yet each is estimated alone
References
Dueling Network Architectures for Deep Reinforcement Learning — Wang, Schaul, Hessel, van Hasselt, Lanctot, de Freitas — 2016 · link