Behavior Cloning

AI and Machine Learning · Reinforcement Learning · 1988 · also: Behaviour cloning, Imitation learning · behavior-cloning.yaml

Fit a policy to demonstrations with a supervised loss and skip the reward entirely. Its own small errors then take it off the states it was shown, where nothing in the data says what to do. That compounding drift is the failure DAgger names and patches.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G behavior-cloning Behavior Cloning policy-gradient Policy Gradient behavior-cloning->policy-gradient demonstrations instead of a reward, so nothing has to be explored offline-rl Offline RL offline-rl->behavior-cloning stitches better trajectories out of mediocre ones instead of copying them supervised-fine-tuning Supervised Fine- Tuning supervised-fine-tuning->behavior-cloning the demonstrations are text and the actions tokens, but the loss is the same

This node

Referenced by

References