Multi-Armed Bandit

AI and Machine Learning · Reinforcement Learning · also: Bandit problem · multi-armed-bandit.yaml

One state, several arms, unknown payoffs. Strip an MDP of everything except the choice between gathering information and cashing it in, and this is what is left, which is why regret can be bounded here and almost nowhere above it.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G epsilon-greedy Epsilon-Greedy multi-armed-bandit Multi-Armed Bandit epsilon-greedy->multi-armed-bandit markov-decision-process Markov Decision Process multi-armed-bandit->markov-decision-process a single state, so no action changes what is faced next upper-confidence-bound Upper Confidence Bound upper-confidence-bound->multi-armed-bandit

This node

Referenced by

References