Target Network

AI and Machine Learning · Reinforcement Learning · 2015 · target-network.yaml

Compute the bootstrap target from a frozen copy of the Q-network, refreshed every few thousand steps. Two lines of code, and the difference between DQN converging and DQN diverging.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G ddpg Deep Deterministic Policy Gradient target-network Target Network ddpg->target-network the same bootstrap divergence, now in the actor as well deep-q-network Deep Q-Network double-dqn Double DQN double-dqn->target-network the independent second estimator is the frozen copy already kept target-network->deep-q-network a target computed from the weights being updated moves with every step

This node

Referenced by

References