Multi-agent Reinforcement Learning in Sequential Social Dilemmas.¶
Leibo, J. Z., Zambaldi, V., Lanctot, M., Marecki, J., & Graepel, T. (2017). Multi-agent Reinforcement Learning in Sequential Social Dilemmas. Proceedings of the 16th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), 464-473.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Other-Regarding Preferences
- In multi-agent reinforcement learning, a designer who wants cooperative rather than competitive convergence writes each agent's reward as its own task reward plus a weighted term in other agents' rewards; the sign and size of that cross-argument weight determine whether agents learn to share, ignore, or sabotage one another, and tuning it is the direct engineering analogue of the economist's "target the cross-argument weights" move.
This sourceShows reward shaping with cross-agent terms steers convergence toward cooperative versus competitive equilibria.
- In multi-agent reinforcement learning, a designer who wants cooperative rather than competitive convergence writes each agent's reward as its own task reward plus a weighted term in other agents' rewards; the sign and size of that cross-argument weight determine whether agents learn to share, ignore, or sabotage one another, and tuning it is the direct engineering analogue of the economist's "target the cross-argument weights" move.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:7fb1d9e41ccd · see in the full table