Runs many actors on their own copy of the environment and applies their gradients to one shared set of weights; A2C is the same idea with the actors stepped in lockstep. Cheap enough to train on CPUs, and the reason on-policy methods no longer needed a replay buffer.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
specializesis a specific case of Actor-Criticone actor and critic, updated from many environments at once
alternative-tois a competing approach to Experience Replayparallel environments decorrelate the updates instead of a buffer
Referenced by
correctsIMPALA fixes a defect in thisan actor that computes gradients stalls at every parameter sync
References
Asynchronous Methods for Deep Reinforcement Learning — Mnih et al. — 2016 · link