Trains a deterministic actor by climbing the critic's gradient with respect to the action, which is what stands in for the max DQN cannot take over a continuous action space. Explores by adding noise to the actor's output, and is notoriously touchy about how much.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in Deep Q-Networkthe max over actions has no closed form once an action is a vector
specializesis a specific case of Actor-Critica deterministic actor, so the critic differentiates straight into it
requiresdoes not work without Target Networkthe same bootstrap divergence, now in the actor as well
Referenced by
correctsSoft Actor-Critic fixes a defect in thisa deterministic actor explores only as far as the noise someone tuned for it
correctsTwin Delayed DDPG fixes a defect in thisthe actor is trained to maximise exactly the critic's own errors
References
Continuous control with deep reinforcement learning — Lillicrap et al. — 2015 · link