Add the policy's entropy to the objective, so it is paid to stay as random as the returns allow. Exploration stops being a schedule bolted on the outside and becomes part of what is optimised.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
extendsadds capability to Markov Decision Processreward alone commits to one of several equally good actions and stops looking
Referenced by
specializesSoft Actor-Critic is a specific case of thisthe critic prices entropy alongside reward, so the actor stays stochastic
References
Reinforcement Learning with Deep Energy-Based Policies — Haarnoja, Tang, Abbeel, Levine — 2017 · link
Maximum Entropy Inverse Reinforcement Learning — Ziebart, Maas, Bagnell, Dey — 2008 · link