Fit a policy to demonstrations with a supervised loss and skip the reward entirely. Its own small errors then take it off the states it was shown, where nothing in the data says what to do. That compounding drift is the failure DAgger names and patches.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
alternative-tois a competing approach to Policy Gradientdemonstrations instead of a reward, so nothing has to be explored
Referenced by
alternative-toOffline RL is a competing approach to thisstitches better trajectories out of mediocre ones instead of copying them
specializesSupervised Fine-Tuning is a specific case of thisthe demonstrations are text and the actions tokens, but the loss is the same
References
ALVINN: An Autonomous Land Vehicle in a Neural Network — Pomerleau — 1988 · link
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Ross, Gordon, Bagnell — 2011 · link