Learning from a reward signal rather than from labelled answers.
Colour is the family; a dashed line is the second member of it.
The value of a state is the reward collected now plus the discounted value of wherever the next step lands. Everything else here is a way of solving… · 1957
Fit a policy to demonstrations with a supervised loss and skip the reward entirely. Its own small errors then take it off the states it was shown, wh… · 1988
Learns the value of each state-action pair directly, converging to the optimal policy without ever modelling the transitions. · 1989
Scale each action's log-probability gradient by the return that followed it. The whole policy-gradient family starts here, and so does its variance p… · 1992
Q-learning's update with one substitution: the target uses the action the policy actually took, not the best one on offer. So the values include the… · 1994
Add the difference of a potential function to the reward, and the optimal policy provably does not move. The result that turned reward engineering fr… · 1999
Pick the arm with the highest optimistic estimate: its mean plus a term that grows the less it has been tried. Exploration follows from uncertainty r… · 2002
Replaces the Q table with a convolutional network reading pixels. The result that made reinforcement learning on raw sensory input look tractable. · 2013
Stores transitions in a buffer and samples them at random, so an update is not dominated by whatever just happened. · 2013
Trains a deterministic actor by climbing the critic's gradient with respect to the action, which is what stands in for the max DQN cannot take over a… · 2015
Choose the next action with the online network, score it with the target copy. Two estimators disagree about which action noise flattered, so the opt… · 2015
Blends TD residuals over many horizons with an exponential weight, giving one knob that trades bias against variance in the advantage estimate. · 2015
Draw a transition in proportion to the size of its last TD error, so the buffer keeps returning the surprising ones and stops returning the solved on… · 2015
Compute the bootstrap target from a frozen copy of the Q-network, refreshed every few thousand steps. Two lines of code, and the difference between D… · 2015
Constrains each update so the KL divergence between the old and new policy stays inside a trust region, giving a monotonic improvement guarantee at t… · 2015
Runs many actors on their own copy of the environment and applies their gradients to one shared set of weights; A2C is the same idea with the actors… · 2016
Splits the head into a state-value stream and a per-action advantage stream, then adds them back. How good a state is gets learned once instead of se… · 2016
Predicts the whole distribution of returns as a histogram over fixed atoms, and applies the Bellman backup to that. Actions are still chosen by its m… · 2017
Pay the agent for reaching states its own predictor gets wrong, so novelty becomes a reward it can follow. What Montezuma's Revenge needed and what u… · 2017
Add the policy's entropy to the objective, so it is paid to stay as random as the returns allow. Exploration stops being a schedule bolted on the out… · 2017
Clips the ratio between new and old policy so an update cannot move too far from the data it was collected under. Most of TRPO's stability with none… · 2017
Six DQN improvements in one agent: double, dueling, prioritized replay, multi-step returns, distributional values and noisy exploration, with an abla… · 2017
Actors ship trajectories to a central learner instead of gradients, so they never wait on a parameter sync, and V-trace corrects for how far the poli… · 2018
Off-policy actor-critic on the maximum-entropy objective, with twin critics and a temperature tuned automatically against a target entropy. The defau… · 2018
Take the smaller of two critics, update the actor every other step, and smooth the target over nearby actions. Three small changes aimed at one failu… · 2018
Adds a term that pushes down the value of actions the dataset does not contain, so what comes out is a lower bound on the policy's true value rather… · 2020
Learn from a fixed dataset with no further interaction, which is the setting wherever a mistake is too expensive to make on purpose: driving, treatme… · 2020
Samples a group of completions per prompt and uses their mean reward as the baseline, so the advantage is relative within the group. Drops the value… · 2024
Learns a policy and a value function together, using the value estimate as the baseline the policy gradient is measured against. Most modern policy m…
Take the best action known, except on an ε fraction of steps where the action is uniformly random. The cheapest exploration that still visits every a…
States, actions, transition probabilities and rewards, where the next state depends only on the present one. The formalism every method below is defi…
One state, several arms, unknown payoffs. Strip an MDP of everything except the choice between gathering information and cashing it in, and this is w…
Differentiates expected return with respect to the policy parameters directly, so stochastic and continuous action spaces need no argmax over actions.
Updates a value estimate toward a later estimate rather than waiting for the final return. Learns from incomplete episodes, which is what makes onlin…