Epsilon-Greedy

AI and Machine Learning · Reinforcement Learning · also: ε-greedy · epsilon-greedy.yaml

Take the best action known, except on an ε fraction of steps where the action is uniformly random. The cheapest exploration that still visits every action infinitely often, and the baseline anything cleverer has to beat.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G epsilon-greedy Epsilon-Greedy multi-armed-bandit Multi-Armed Bandit epsilon-greedy->multi-armed-bandit q-learning Q-Learning epsilon-greedy->q-learning acting greedily never revisits an action its first sample underrated intrinsic-motivation Intrinsic Motivation intrinsic-motivation->epsilon-greedy random actions never chain into the long sequence a sparse reward needs upper-confidence-bound Upper Confidence Bound upper-confidence-bound->epsilon-greedy uniform exploration keeps re-trying arms already measured as bad

This node

Referenced by

References