Twin Delayed DDPG

AI and Machine Learning · Reinforcement Learning · 2018 · also: TD3 · td3.yaml

Take the smaller of two critics, update the actor every other step, and smooth the target over nearby actions. Three small changes aimed at one failure: the actor's gradient ascent finds whatever the critic overrates and goes there.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G ddpg Deep Deterministic Policy Gradient double-dqn Double DQN td3 Twin Delayed DDPG td3->ddpg the actor is trained to maximise exactly the critic's own errors td3->double-dqn carries the two-estimator fix into continuous actions, as a min over the pair

This node

References