Choose the next action with the online network, score it with the target copy. Two estimators disagree about which action noise flattered, so the optimism the single max was compounding largely cancels.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in Deep Q-Networkone network both picks the next action and rates it, so noise reads as value
requiresdoes not work without Target Networkthe independent second estimator is the frozen copy already kept
Referenced by
extendsTwin Delayed DDPG adds capability to thiscarries the two-estimator fix into continuous actions, as a min over the pair
References
Deep Reinforcement Learning with Double Q-learning — van Hasselt, Guez, Silver — 2015 · link