Policy Gradient
AI and Machine Learning · Reinforcement Learning · policy-gradient.yaml
Differentiates expected return with respect to the policy parameters directly, so stochastic and continuous action spaces need no argmax over actions.
- supersedescorrects · extends
- classifiesspecializes · part-of
- substitutes forapproximates · alternative-to
- depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
This node
- alternative-tois a competing approach to Q-Learningoptimises the policy itself instead of a value function
Referenced by
- extendsActor-Critic adds capability to thisa learned baseline instead of sampled returns
- alternative-toBehavior Cloning is a competing approach to thisdemonstrations instead of a reward, so nothing has to be explored
- correctsGeneralized Advantage Estimation fixes a defect in thisraw returns are too noisy to learn from directly
- correctsProximal Policy Optimization fixes a defect in thisan unconstrained step can collapse the policy
- specializesREINFORCE is a specific case of thisthe coefficient is the episode's own return, with nothing subtracted
- correctsTrust Region Policy Optimization fixes a defect in thisa parameter-space step moves the policy unpredictably
References