Offline RL

AI and Machine Learning · Reinforcement Learning · 2020 · also: Batch RL · offline-rl.yaml

Learn from a fixed dataset with no further interaction, which is the setting wherever a mistake is too expensive to make on purpose: driving, treatment, a live recommender. Its promise over copying the data is stitching: good fragments of separate mediocre episodes joined into one policy.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G behavior-cloning Behavior Cloning conservative-q-learning Conservative Q-Learning offline-rl Offline RL conservative-q-learning->offline-rl the critic overrates actions never seen, and the policy walks straight to them experience-replay Experience Replay offline-rl->behavior-cloning stitches better trajectories out of mediocre ones instead of copying them offline-rl->experience-replay the buffer is the whole world, and nothing is ever added to it

This node

Referenced by

References