Learn from a fixed dataset with no further interaction, which is the setting wherever a mistake is too expensive to make on purpose: driving, treatment, a live recommender. Its promise over copying the data is stitching: good fragments of separate mediocre episodes joined into one policy.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
alternative-tois a competing approach to Behavior Cloningstitches better trajectories out of mediocre ones instead of copying them
requiresdoes not work without Experience Replaythe buffer is the whole world, and nothing is ever added to it
Referenced by
correctsConservative Q-Learning fixes a defect in thisthe critic overrates actions never seen, and the policy walks straight to them
References
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems — Levine, Kumar, Tucker, Fu — 2020 · link