Add the difference of a potential function to the reward, and the optimal policy provably does not move. The result that turned reward engineering from guesswork into a statement about which extra signal is safe to give.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
extendsadds capability to Markov Decision Processdense guidance can be added to a sparse reward without moving its optimum
References
Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping — Ng, Harada, Russell — 1999 · link