Double DQN

AI and Machine Learning · Reinforcement Learning · 2015 · also: Double Q-learning, DDQN · double-dqn.yaml

Choose the next action with the online network, score it with the target copy. Two estimators disagree about which action noise flattered, so the optimism the single max was compounding largely cancels.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G deep-q-network Deep Q-Network double-dqn Double DQN double-dqn->deep-q-network one network both picks the next action and rates it, so noise reads as value target-network Target Network double-dqn->target-network the independent second estimator is the frozen copy already kept td3 Twin Delayed DDPG td3->double-dqn carries the two-estimator fix into continuous actions, as a min over the pair

This node

Referenced by

References