Emergent Prosociality in Multi-Agent Games Through Gifting

Emergent Prosociality in Multi-Agent Games Through Gifting
复制标题

DOI:
10.24963/ijcai.2021/61
复制
发表时间:
2021-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Woodrow Z. Wang;M. Beliaev;Erdem Biyik;Daniel A. Lazar;Ramtin Pedarsani;Dorsa Sadigh
Woodrow Z. Wang;M. Beliaev;Erdem Biyik;Daniel A. Lazar;Ramtin Pedarsani;Dorsa Sadigh
中科院分区:
其他
文献类型:
--
作者:
Woodrow Z. Wang;M. Beliaev;Erdem Biyik;Daniel A. Lazar;Ramtin Pedarsani;Dorsa Sadigh

文献摘要

被引文献

相似文献

协调对于形成亲社会行为通常至关重要,这种行为会增加多智能体游戏中所有智能体获得的奖励总额。然而,当存在多个均衡时,最先进的强化学习算法常常会收敛到社会上不太理想的均衡。之前的工作通过明确的奖励塑造来解决这一挑战,这需要强有力的假设,即代理人可以被迫变得亲社会。我们建议使用一种限制较少的同伴奖励机制,即礼物,引导代理人走向社会更理想的平衡,同时允许代理人保持自私和去中心化。礼物允许每个代理将他们的一些奖励送给其他代理。我们采用了一个理论框架,通过描述动态系统中平衡的吸引力盆地来捕捉礼物在收敛到亲社会平衡方面的好处。通过赠送,我们通过数值分析和实验证明了高风险、总和协调博弈与亲社会均衡的收敛性增强。
Coordination is often critical to forming prosocial behaviors -- behaviors that increase the overall sum of rewards received by all agents in a multi-agent game. However, state of the art reinforcement learning algorithms often suffer from converging to socially less desirable equilibria when multiple equilibria exist. Previous works address this challenge with explicit reward shaping, which requires the strong assumption that agents can be forced to be prosocial. We propose using a less restrictive peer-rewarding mechanism, gifting, that guides the agents toward more socially desirable equilibria while allowing agents to remain selfish and decentralized. Gifting allows each agent to give some of their reward to other agents. We employ a theoretical framework that captures the benefit of gifting in converging to the prosocial equilibrium by characterizing the equilibria's basins of attraction in a dynamical system. With gifting, we demonstrate increased convergence of high risk, general-sum coordination games to the prosocial equilibrium both via numerical analysis and experiments.