Upper Confidence Bound

AI and Machine Learning · Reinforcement Learning · 2002 · also: UCB, UCB1 · upper-confidence-bound.yaml

Pick the arm with the highest optimistic estimate: its mean plus a term that grows the less it has been tried. Exploration follows from uncertainty rather than from a coin flip, and the regret grows with the logarithm of time where a fixed exploration rate makes it linear.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G epsilon-greedy Epsilon-Greedy monte-carlo-tree-search Monte Carlo Tree Search upper-confidence-bound Upper Confidence Bound monte-carlo-tree-search->upper-confidence-bound every node of the tree is its own bandit problem multi-armed-bandit Multi-Armed Bandit upper-confidence-bound->epsilon-greedy uniform exploration keeps re-trying arms already measured as bad upper-confidence-bound->multi-armed-bandit

This node

Referenced by

References