Q-learning's update with one substitution: the target uses the action the policy actually took, not the best one on offer. So the values include the cost of exploring, and the agent walks the safe path along a cliff where Q-learning walks the edge and occasionally falls off.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
specializesis a specific case of Temporal Difference Learningthe bootstrapped target is the action the behaviour policy chose
alternative-tois a competing approach to Q-Learningon-policy, so the risk of its own exploration is priced into the values
References
On-line Q-learning using connectionist systems — Rummery, Niranjan — 1994 · link