Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping¶
Ng, A. Y., Harada, D., & Russell, S. (1999). Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping: Theory and Application to Reward Shaping. Proceedings of the Sixteenth International Conference on Machine Learning.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Reference-Point Dependence
- The subtraction can leave the optimum untouched while transforming the learning dynamics — reference dependence in a system with no experience of gain or loss whatsoever.
This sourceProves that potential-based shaping leaves the optimal policy unchanged while altering the learning dynamics.
- The subtraction can leave the optimum untouched while transforming the learning dynamics — reference dependence in a system with no experience of gain or loss whatsoever.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:53210b5e3465 · see in the full table