Constrains each update so the KL divergence between the old and new policy stays inside a trust region, giving a monotonic improvement guarantee at the cost of a second-order solve every step.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in Policy Gradienta parameter-space step moves the policy unpredictably
Referenced by
approximatesProximal Policy Optimization is a cheaper stand-in for thisclipping instead of a constrained second-order solve
References
Trust Region Policy Optimization — Schulman, Levine, Abbeel, Jordan, Moritz — 2015 · link