Scale each action's log-probability gradient by the return that followed it. The whole policy-gradient family starts here, and so does its variance problem: one lucky episode is indistinguishable from a better action.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
specializesis a specific case of Policy Gradientthe coefficient is the episode's own return, with nothing subtracted
requiresdoes not work without Monte Carlo Integrationthe gradient is an expectation over trajectories it can only sample
References
Simple statistical gradient-following algorithms for connectionist reinforcement learning — Williams — 1992 · link