Soft Actor-Critic

AI and Machine Learning · Reinforcement Learning · 2018 · also: SAC · soft-actor-critic.yaml

Off-policy actor-critic on the maximum-entropy objective, with twin critics and a temperature tuned automatically against a target entropy. The default for continuous control: DDPG's sample efficiency without DDPG's habit of collapsing when a hyperparameter moves.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G ddpg Deep Deterministic Policy Gradient experience-replay Experience Replay maximum-entropy-rl Maximum Entropy RL ppo Proximal Policy Optimization soft-actor-critic Soft Actor-Critic soft-actor-critic->ddpg a deterministic actor explores only as far as the noise someone tuned for it soft-actor-critic->experience-replay its updates come from the buffer, not from the latest rollout soft-actor-critic->maximum-entropy-rl the critic prices entropy alongside reward, so the actor stays stochastic soft-actor-critic->ppo off-policy, so a transition is reused for the rest of training instead of dropped after one batch

This node

References