Off-policy actor-critic on the maximum-entropy objective, with twin critics and a temperature tuned automatically against a target entropy. The default for continuous control: DDPG's sample efficiency without DDPG's habit of collapsing when a hyperparameter moves.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
specializesis a specific case of Maximum Entropy RLthe critic prices entropy alongside reward, so the actor stays stochastic
correctsfixes a defect in Deep Deterministic Policy Gradienta deterministic actor explores only as far as the noise someone tuned for it
alternative-tois a competing approach to Proximal Policy Optimizationoff-policy, so a transition is reused for the rest of training instead of dropped after one batch
requiresdoes not work without Experience Replayits updates come from the buffer, not from the latest rollout
References
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor — Haarnoja, Zhou, Abbeel, Levine — 2018 · link
Soft Actor-Critic Algorithms and Applications — Haarnoja et al. — 2018 · link