Deep Deterministic Policy Gradient

AI and Machine Learning · Reinforcement Learning · 2015 · also: DDPG · ddpg.yaml

Trains a deterministic actor by climbing the critic's gradient with respect to the action, which is what stands in for the max DQN cannot take over a continuous action space. Explores by adding noise to the actor's output, and is notoriously touchy about how much.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G actor-critic Actor-Critic ddpg Deep Deterministic Policy Gradient ddpg->actor-critic a deterministic actor, so the critic differentiates straight into it deep-q-network Deep Q-Network ddpg->deep-q-network the max over actions has no closed form once an action is a vector target-network Target Network ddpg->target-network the same bootstrap divergence, now in the actor as well soft-actor-critic Soft Actor-Critic soft-actor-critic->ddpg a deterministic actor explores only as far as the noise someone tuned for it td3 Twin Delayed DDPG td3->ddpg the actor is trained to maximise exactly the critic's own errors

This node

Referenced by

References