Generalized Advantage Estimation

AI and Machine Learning · Reinforcement Learning · 2015 · also: GAE · generalized-advantage-estimation.yaml

Blends TD residuals over many horizons with an exponential weight, giving one knob that trades bias against variance in the advantage estimate.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G actor-critic Actor-Critic generalized-advantage-estimation Generalized Advantage Estimation generalized-advantage-estimation->actor-critic the value function it leans on is the critic policy-gradient Policy Gradient generalized-advantage-estimation->policy-gradient raw returns are too noisy to learn from directly

This node

References