Adds a term that pushes down the value of actions the dataset does not contain, so what comes out is a lower bound on the policy's true value rather than an optimistic guess.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in Offline RLthe critic overrates actions never seen, and the policy walks straight to them
References
Conservative Q-Learning for Offline Reinforcement Learning — Kumar, Zhou, Tucker, Levine — 2020 · link