AI and Machine Learning

Reinforcement Learning

Learning from a reward signal rather than from labelled answers.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it · hover one for the reasons on its edges, and to bring it forward
G a3c Asynchronous Advantage Actor- Critic actor-critic Actor-Critic a3c->actor-critic experience-replay Experience Replay a3c->experience-replay policy-gradient Policy Gradient actor-critic->policy-gradient behavior-cloning Behavior Cloning behavior-cloning->policy-gradient bellman-equation Bellman Equation markov-decision-process Markov Decision Process bellman-equation->markov-decision-process conservative-q-learning Conservative Q-Learning offline-rl Offline RL conservative-q-learning->offline-rl ddpg Deep Deterministic Policy Gradient ddpg->actor-critic deep-q-network Deep Q-Network ddpg->deep-q-network target-network Target Network ddpg->target-network q-learning Q-Learning deep-q-network->q-learning distributional-rl Distributional Reinforcement Learning distributional-rl->q-learning double-dqn Double DQN double-dqn->deep-q-network double-dqn->target-network dueling-network Dueling Network dueling-network->deep-q-network epsilon-greedy Epsilon-Greedy multi-armed-bandit Multi-Armed Bandit epsilon-greedy->multi-armed-bandit epsilon-greedy->q-learning experience-replay->deep-q-network generalized-advantage-estimation Generalized Advantage Estimation generalized-advantage-estimation->actor-critic generalized-advantage-estimation->policy-gradient grpo Group Relative Policy Optimization grpo->actor-critic ppo Proximal Policy Optimization grpo->ppo impala IMPALA impala->a3c importance-sampling Importance Sampling impala->importance-sampling intrinsic-motivation Intrinsic Motivation intrinsic-motivation->epsilon-greedy maximum-entropy-rl Maximum Entropy RL maximum-entropy-rl->markov-decision-process multi-armed-bandit->markov-decision-process offline-rl->behavior-cloning offline-rl->experience-replay policy-gradient->q-learning potential-based-reward-shaping Potential-Based Reward Shaping potential-based-reward-shaping->markov-decision-process ppo->policy-gradient trpo Trust Region Policy Optimization ppo->trpo ppo->importance-sampling prioritized-experience-replay Prioritized Experience Replay prioritized-experience-replay->experience-replay prioritized-experience-replay->importance-sampling temporal-difference-learning Temporal Difference Learning q-learning->temporal-difference-learning rainbow Rainbow rainbow->deep-q-network rainbow->distributional-rl rainbow->prioritized-experience-replay reinforce REINFORCE reinforce->policy-gradient monte-carlo-integration Monte Carlo Integration reinforce->monte-carlo-integration sarsa SARSA sarsa->q-learning sarsa->temporal-difference-learning soft-actor-critic Soft Actor-Critic soft-actor-critic->ddpg soft-actor-critic->experience-replay soft-actor-critic->maximum-entropy-rl soft-actor-critic->ppo target-network->deep-q-network td3 Twin Delayed DDPG td3->ddpg td3->double-dqn temporal-difference-learning->bellman-equation temporal-difference-learning->markov-decision-process trpo->policy-gradient upper-confidence-bound Upper Confidence Bound upper-confidence-bound->epsilon-greedy upper-confidence-bound->multi-armed-bandit
The 55 reasons on these edges, as text

34 nodes

Bellman Equation

The value of a state is the reward collected now plus the discounted value of wherever the next step lands. Everything else here is a way of solving… · 1957

Behavior Cloning

Fit a policy to demonstrations with a supervised loss and skip the reward entirely. Its own small errors then take it off the states it was shown, wh… · 1988

Q-Learning

Learns the value of each state-action pair directly, converging to the optimal policy without ever modelling the transitions. · 1989

REINFORCE

Scale each action's log-probability gradient by the return that followed it. The whole policy-gradient family starts here, and so does its variance p… · 1992

SARSA

Q-learning's update with one substitution: the target uses the action the policy actually took, not the best one on offer. So the values include the… · 1994

Potential-Based Reward Shaping

Add the difference of a potential function to the reward, and the optimal policy provably does not move. The result that turned reward engineering fr… · 1999

Upper Confidence Bound

Pick the arm with the highest optimistic estimate: its mean plus a term that grows the less it has been tried. Exploration follows from uncertainty r… · 2002

Deep Q-Network

Replaces the Q table with a convolutional network reading pixels. The result that made reinforcement learning on raw sensory input look tractable. · 2013

Experience Replay

Stores transitions in a buffer and samples them at random, so an update is not dominated by whatever just happened. · 2013

Deep Deterministic Policy Gradient

Trains a deterministic actor by climbing the critic's gradient with respect to the action, which is what stands in for the max DQN cannot take over a… · 2015

Double DQN

Choose the next action with the online network, score it with the target copy. Two estimators disagree about which action noise flattered, so the opt… · 2015

Generalized Advantage Estimation

Blends TD residuals over many horizons with an exponential weight, giving one knob that trades bias against variance in the advantage estimate. · 2015

Prioritized Experience Replay

Draw a transition in proportion to the size of its last TD error, so the buffer keeps returning the surprising ones and stops returning the solved on… · 2015

Target Network

Compute the bootstrap target from a frozen copy of the Q-network, refreshed every few thousand steps. Two lines of code, and the difference between D… · 2015

Trust Region Policy Optimization

Constrains each update so the KL divergence between the old and new policy stays inside a trust region, giving a monotonic improvement guarantee at t… · 2015

Asynchronous Advantage Actor-Critic

Runs many actors on their own copy of the environment and applies their gradients to one shared set of weights; A2C is the same idea with the actors… · 2016

Dueling Network

Splits the head into a state-value stream and a per-action advantage stream, then adds them back. How good a state is gets learned once instead of se… · 2016

Distributional Reinforcement Learning

Predicts the whole distribution of returns as a histogram over fixed atoms, and applies the Bellman backup to that. Actions are still chosen by its m… · 2017

Intrinsic Motivation

Pay the agent for reaching states its own predictor gets wrong, so novelty becomes a reward it can follow. What Montezuma's Revenge needed and what u… · 2017

Maximum Entropy RL

Add the policy's entropy to the objective, so it is paid to stay as random as the returns allow. Exploration stops being a schedule bolted on the out… · 2017

Proximal Policy Optimization

Clips the ratio between new and old policy so an update cannot move too far from the data it was collected under. Most of TRPO's stability with none… · 2017

Rainbow

Six DQN improvements in one agent: double, dueling, prioritized replay, multi-step returns, distributional values and noisy exploration, with an abla… · 2017

IMPALA

Actors ship trajectories to a central learner instead of gradients, so they never wait on a parameter sync, and V-trace corrects for how far the poli… · 2018

Soft Actor-Critic

Off-policy actor-critic on the maximum-entropy objective, with twin critics and a temperature tuned automatically against a target entropy. The defau… · 2018

Twin Delayed DDPG

Take the smaller of two critics, update the actor every other step, and smooth the target over nearby actions. Three small changes aimed at one failu… · 2018

Conservative Q-Learning

Adds a term that pushes down the value of actions the dataset does not contain, so what comes out is a lower bound on the policy's true value rather… · 2020

Offline RL

Learn from a fixed dataset with no further interaction, which is the setting wherever a mistake is too expensive to make on purpose: driving, treatme… · 2020

Group Relative Policy Optimization

Samples a group of completions per prompt and uses their mean reward as the baseline, so the advantage is relative within the group. Drops the value… · 2024

Actor-Critic

Learns a policy and a value function together, using the value estimate as the baseline the policy gradient is measured against. Most modern policy m…

Epsilon-Greedy

Take the best action known, except on an ε fraction of steps where the action is uniformly random. The cheapest exploration that still visits every a…

Markov Decision Process

States, actions, transition probabilities and rewards, where the next state depends only on the present one. The formalism every method below is defi…

Multi-Armed Bandit

One state, several arms, unknown payoffs. Strip an MDP of everything except the choice between gathering information and cashing it in, and this is w…

Policy Gradient

Differentiates expected return with respect to the policy parameters directly, so stochastic and continuous action spaces need no argmax over actions.

Temporal Difference Learning

Updates a value estimate toward a later estimate rather than waiting for the final return. Learns from incomplete episodes, which is what makes onlin…