Learns a policy and a value function together, using the value estimate as the baseline the policy gradient is measured against. Most modern policy methods are some form of this.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
extendsadds capability to Policy Gradienta learned baseline instead of sampled returns