IMPALA

AI and Machine Learning · Reinforcement Learning · 2018 · also: Importance Weighted Actor-Learner Architecture, V-trace · impala.yaml

Actors ship trajectories to a central learner instead of gradients, so they never wait on a parameter sync, and V-trace corrects for how far the policy has moved by the time a batch is used. The shape distributed RL settled into.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G a3c Asynchronous Advantage Actor- Critic impala IMPALA impala->a3c an actor that computes gradients stalls at every parameter sync importance-sampling Importance Sampling impala->importance-sampling the trajectory is off- policy by the time the learner gets to it

This node

References