Conservative Q-Learning

AI and Machine Learning · Reinforcement Learning · 2020 · also: CQL · conservative-q-learning.yaml

Adds a term that pushes down the value of actions the dataset does not contain, so what comes out is a lower bound on the policy's true value rather than an optimistic guess.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G conservative-q-learning Conservative Q-Learning offline-rl Offline RL conservative-q-learning->offline-rl the critic overrates actions never seen, and the policy walks straight to them

This node

References