Q-Learning
AI and Machine Learning · Reinforcement Learning · 1989 · q-learning.yaml
Learns the value of each state-action pair directly, converging to the optimal policy without ever modelling the transitions.
- supersedescorrects · extends
- classifiesspecializes · part-of
- substitutes forapproximates · alternative-to
- depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
This node
Referenced by
- correctsDeep Q-Network fixes a defect in thisa table cannot cover a high-dimensional state space
- extendsDistributional Reinforcement Learning adds capability to thisan expectation discards the shape of the returns that produced it
- correctsEpsilon-Greedy fixes a defect in thisacting greedily never revisits an action its first sample underrated
- alternative-toModel-Based RL is a competing approach to thisplans with learned dynamics instead of caching values
- alternative-toPolicy Gradient is a competing approach to thisoptimises the policy itself instead of a value function
- alternative-toSARSA is a competing approach to thison-policy, so the risk of its own exploration is priced into the values
References