SARSA

AI and Machine Learning · Reinforcement Learning · 1994 · also: State-Action-Reward-State-Action · sarsa.yaml

Q-learning's update with one substitution: the target uses the action the policy actually took, not the best one on offer. So the values include the cost of exploring, and the agent walks the safe path along a cliff where Q-learning walks the edge and occasionally falls off.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G q-learning Q-Learning sarsa SARSA sarsa->q-learning on-policy, so the risk of its own exploration is priced into the values temporal-difference-learning Temporal Difference Learning sarsa->temporal-difference-learning the bootstrapped target is the action the behaviour policy chose

This node

References