Markov Decision Process

AI and Machine Learning · Reinforcement Learning · also: MDP · markov-decision-process.yaml

States, actions, transition probabilities and rewards, where the next state depends only on the present one. The formalism every method below is defined against.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G bellman-equation Bellman Equation markov-decision-process Markov Decision Process bellman-equation->markov-decision-process maximum-entropy-rl Maximum Entropy RL maximum-entropy-rl->markov-decision-process reward alone commits to one of several equally good actions and stops looking monte-carlo-tree-search Monte Carlo Tree Search monte-carlo-tree-search->markov-decision-process it searches the tree those transitions define multi-armed-bandit Multi-Armed Bandit multi-armed-bandit->markov-decision-process a single state, so no action changes what is faced next potential-based-reward-shaping Potential-Based Reward Shaping potential-based-reward-shaping->markov-decision-process dense guidance can be added to a sparse reward without moving its optimum temporal-difference-learning Temporal Difference Learning temporal-difference-learning->markov-decision-process estimates the value function the MDP defines

Referenced by

References