Multi-Agent Reinforcement Learning via Double Averaging Primal-Dual Optimization

Multi-Agent Reinforcement Learning via Double Averaging Primal-Dual Optimization
复制标题

DOI:
--
复制
发表时间:
2018-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Hoi-To Wai;Zhuoran Yang;Zhaoran Wang;Mingyi Hong
Hoi-To Wai;Zhuoran Yang;Zhaoran Wang;Mingyi Hong
中科院分区:
其他
文献类型:
--
作者:
Hoi-To Wai;Zhuoran Yang;Zhaoran Wang;Mingyi Hong

文献摘要

被引文献

相似文献

尽管单智能体强化学习取得了成功,但由于智能体之间复杂的交互,多智能体强化学习(MARL)仍然具有挑战性。受传感器网络、群体机器人和电网等去中心化应用的推动,我们研究了 MARL 中的策略评估,其中具有共同观察的状态-动作对和私人本地奖励的代理协作以了解给定策略的价值。在本文中,我们提出了一种双重平均方案,其中每个代理在空间和时间上迭代执行平均,以分别合并相邻梯度信息和局部奖励信息。我们证明所提出的算法以全局几何速率收敛到最优解。特别是,这种算法建立在均方投影贝尔曼误差最小化问题的原对偶重构的基础上,这产生了分散的凸凹鞍点问题。据我们所知,所提出的双平均原对偶优化算法是第一个在分散凸凹鞍点问题上实现快速有限时间收敛的算法。
Despite the success of single-agent reinforcement learning, multi-agent reinforcement learning (MARL) remains challenging due to complex interactions between agents. Motivated by decentralized applications such as sensor networks, swarm robotics, and power grids, we study policy evaluation in MARL, where agents with jointly observed state-action pairs and private local rewards collaborate to learn the value of a given policy. In this paper, we propose a double averaging scheme, where each agent iteratively performs averaging over both space and time to incorporate neighboring gradient information and local reward information, respectively. We prove that the proposed algorithm converges to the optimal solution at a global geometric rate. In particular, such an algorithm is built upon a primal-dual reformulation of the mean squared projected Bellman error minimization problem, which gives rise to a decentralized convex-concave saddle-point problem. To the best of our knowledge, the proposed double averaging primal-dual optimization algorithm is the first to achieve fast finite-time convergence on decentralized convex-concave saddle-point problems.