Stable and Efficient Shapley Value-Based Reward Reallocation for Multi-Agent Reinforcement Learning of Autonomous Vehicles

Stable and Efficient Shapley Value-Based Reward Reallocation for Multi-Agent Reinforcement Learning of Autonomous Vehicles
复制标题

DOI:
10.48550/arxiv.2203.06333
复制
发表时间:
2022-03
期刊:
2022 International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Songyang Han;He Wang;Sanbao Su;Yuanyuan Shi;Fei Miao
Songyang Han;He Wang;Sanbao Su;Yuanyuan Shi;Fei Miao
中科院分区:
其他
文献类型:
--
作者:
Songyang Han;He Wang;Sanbao Su;Yuanyuan Shi;Fei Miao

文献摘要

被引文献

相似文献

随着网络化信息物理系统(CPS)中传感和通信技术的发展,基于多智能体强化学习(MARL)的方法被集成到物理系统的控制过程中,并在广泛的CPS领域(如连接的自动驾驶车辆(CAV))中表现出突出的性能。然而,它仍然具有挑战性的数学表征的通信和合作能力的CAV的性能的改善。当每个自主车辆最初都是自我利益时,我们不能假设所有智能体在训练过程中都会自然合作。在这项工作中,我们建议有效地重新分配系统的总奖励,以激励自动驾驶汽车之间的稳定合作。我们正式定义和量化如何重新分配系统的总奖励给每个代理下提出的可转让的效用博弈,这样,基于通信的多代理之间的合作增加了系统的总奖励。证明了当可转移效用博弈是凸博弈时,基于Shapley值的MARL报酬再分配位于核心。因此,合作是稳定和有效的,代理人应该留在联盟或合作组。然后,我们提出了一个合作的策略学习算法与Shapley值奖励重新分配。在实验中,与几个文献算法相比,我们表明,使用我们的算法的CAV系统的平均情节系统奖励的改善。
With the development of sensing and communication technologies in networked cyber-physical systems (CPSs), multi-agent reinforcement learning (MARL)-based methodologies are integrated into the control process of physical systems and demonstrate prominent performance in a wide array of CPS domains, such as connected autonomous vehicles (CAVs). However, it remains challenging to mathematically characterize the improvement of the performance of CAVs with communication and cooperation capability. When each individual autonomous vehicle is originally self-interest, we can not assume that all agents would cooperate naturally during the training process. In this work, we propose to reallocate the system's total reward efficiently to motivate stable cooperation among autonomous vehicles. We formally define and quantify how to reallocate the system's total reward to each agent under the proposed transferable utility game, such that communication-based cooperation among multi-agents increases the system's total reward. We prove that Shapley value-based reward reallocation of MARL locates in the core if the transferable utility game is a convex game. Hence, the cooperation is stable and efficient and the agents should stay in the coalition or the cooperating group. We then propose a cooperative policy learning algorithm with Shapley value reward reallocation. In experiments, compared with several literature algorithms, we show the improvement of the mean episode system reward of CAV systems using our proposed algorithm.