REINFORCE

AI and Machine Learning · Reinforcement Learning · 1992 · also: Score function estimator, Likelihood ratio method · reinforce.yaml

Scale each action's log-probability gradient by the return that followed it. The whole policy-gradient family starts here, and so does its variance problem: one lucky episode is indistinguishable from a better action.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G monte-carlo-integration Monte Carlo Integration policy-gradient Policy Gradient reinforce REINFORCE reinforce->monte-carlo-integration the gradient is an expectation over trajectories it can only sample reinforce->policy-gradient the coefficient is the episode's own return, with nothing subtracted

This node

References