Asynchronous Advantage Actor-Critic

AI and Machine Learning · Reinforcement Learning · 2016 · also: A3C, A2C · a3c.yaml

Runs many actors on their own copy of the environment and applies their gradients to one shared set of weights; A2C is the same idea with the actors stepped in lockstep. Cheap enough to train on CPUs, and the reason on-policy methods no longer needed a replay buffer.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G a3c Asynchronous Advantage Actor- Critic actor-critic Actor-Critic a3c->actor-critic one actor and critic, updated from many environments at once experience-replay Experience Replay a3c->experience-replay parallel environments decorrelate the updates instead of a buffer impala IMPALA impala->a3c an actor that computes gradients stalls at every parameter sync

This node

Referenced by

References