Context-Aware Bayesian Network Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning

Context-Aware Bayesian Network Actor-Critic Methods for Cooperative Multi-Agent Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2306.01920
复制
发表时间:
2023-06
期刊:
--
影响因子:
--
通讯作者:
Dingyang Chen;Qi Zhang
Dingyang Chen;Qi Zhang
中科院分区:
其他
文献类型:
--
作者:
Dingyang Chen;Qi Zhang

文献摘要

相似文献

以关联方式执行动作是人类协调的常见策略,通常会带来更好的合作,这也可能有利于合作多智能体强化学习(MARL)。然而,MARL 最近的成功在很大程度上依赖于纯粹去中心化执行的便捷范式,出于可扩展性的考虑,代理之间没有动作关联。在这项工作中,我们引入了贝叶斯网络来建立代理在其联合策略中的行动选择之间的相关性。从理论上讲,我们通过推导这种贝叶斯网络联合策略下的多智能体策略梯度公式,并在合作马尔可夫博弈中的表格 Softmax 策略参数化下证明其全局收敛于纳什均衡,为动作依赖性为何有益提供了理论依据。此外,通过为现有的 MARL 算法配备最新的可微有向无环图(DAG)方法,我们开发了实用的算法来在具有部分可观测性和各种难度的场景中学习上下文感知贝叶斯网络策略。我们还在整个训练过程中动态降低学习到的 DAG 的稀疏性,这导致去中心化执行的策略较弱甚至完全独立。一系列 MARL 基准的实证结果显示了我们方法的优点。
Executing actions in a correlated manner is a common strategy for human coordination that often leads to better cooperation, which is also potentially beneficial for cooperative multi-agent reinforcement learning (MARL). However, the recent success of MARL relies heavily on the convenient paradigm of purely decentralized execution, where there is no action correlation among agents for scalability considerations. In this work, we introduce a Bayesian network to inaugurate correlations between agents' action selections in their joint policy. Theoretically, we establish a theoretical justification for why action dependencies are beneficial by deriving the multi-agent policy gradient formula under such a Bayesian network joint policy and proving its global convergence to Nash equilibria under tabular softmax policy parameterization in cooperative Markov games. Further, by equipping existing MARL algorithms with a recent method of differentiable directed acyclic graphs (DAGs), we develop practical algorithms to learn the context-aware Bayesian network policies in scenarios with partial observability and various difficulty. We also dynamically decrease the sparsity of the learned DAG throughout the training process, which leads to weakly or even purely independent policies for decentralized execution. Empirical results on a range of MARL benchmarks show the benefits of our approach.