Maximum Entropy RL

AI and Machine Learning · Reinforcement Learning · 2017 · maximum-entropy-rl.yaml

Add the policy's entropy to the objective, so it is paid to stay as random as the returns allow. Exploration stops being a schedule bolted on the outside and becomes part of what is optimised.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G markov-decision-process Markov Decision Process maximum-entropy-rl Maximum Entropy RL maximum-entropy-rl->markov-decision-process reward alone commits to one of several equally good actions and stops looking soft-actor-critic Soft Actor-Critic soft-actor-critic->maximum-entropy-rl the critic prices entropy alongside reward, so the actor stays stochastic

This node

Referenced by

References