Bellman Equation

AI and Machine Learning · Reinforcement Learning · 1957 · bellman-equation.yaml

The value of a state is the reward collected now plus the discounted value of wherever the next step lands. Everything else here is a way of solving that recursion without being handed the transition probabilities it is written in terms of.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G bellman-equation Bellman Equation markov-decision-process Markov Decision Process bellman-equation->markov-decision-process temporal-difference-learning Temporal Difference Learning temporal-difference-learning->bellman-equation samples one transition where the equation averages over every next state

This node

Referenced by

References