Decentralized Graph-Based Multi-Agent Reinforcement Learning Using Reward Machines

Decentralized Graph-Based Multi-Agent Reinforcement Learning Using Reward Machines
复制标题

DOI:
10.1016/j.neucom.2023.126974
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Jueming Hu;Zhe Xu;Weichang Wang;Guannan Qu;Yutian Pang;Yongming Liu
Jueming Hu;Zhe Xu;Weichang Wang;Guannan Qu;Yutian Pang;Yongming Liu
中科院分区:
其他
文献类型:
--
作者:
Jueming Hu;Zhe Xu;Weichang Wang;Guannan Qu;Yutian Pang;Yongming Liu

文献摘要

相似文献

在多智能体强化学习(MARL)中,智能体集合学习复杂的时间扩展任务具有挑战性。困难在于计算复杂性以及如何学习奖励函数背后的高级思想。我们研究基于图的马尔可夫决策过程(MDP),其中相邻代理的动态是耦合的。为了学习复杂的时间扩展任务,我们使用奖励机(RM)对每个代理的任务进行编码并公开奖励函数的内部结构。 RM 具有描述高级知识和编码非马尔可夫奖励函数的能力。我们提出了一种去中心化学习算法来解决计算复杂性,称为使用奖励机器的去中心化基于图的强化学习(DGRM),该算法为每个智能体配备本地化策略,允许智能体根据可用信息独立做出决策。 DGRM 使用 actor-critic 结构,并且我们引入了离散状态问题的表格 Q 函数。我们表明,随着其他智能体之间距离的增加,Q 函数对其他智能体的依赖性呈指数下降。为了进一步提高效率,我们还提出了深度DGRM算法,使用深度神经网络来近似Q函数和策略函数来解决大规模或连续状态问题。所提出的 DGRM 算法的有效性通过三个案例研究、两个分别具有独立和相关奖励函数的无线通信案例研究以及 COVID-19 大流行缓解来评估。实验结果表明,本地信息对于 DGRM 来说是足够的,并且智能体可以在 RM 的帮助下完成复杂的任务。与缓解 COVID-19 大流行的情况下的基线相比,DGRM 将全球累积奖励提高了 119%。
In multi-agent reinforcement learning (MARL), it is challenging for a collection of agents to learn complex temporally extended tasks. The difficulties lie in computational complexity and how to learn the high-level ideas behind reward functions. We study the graph-based Markov Decision Process (MDP), where the dynamics of neighboring agents are coupled. To learn complex temporally extended tasks, we use a reward machine (RM) to encode each agent’s task and expose reward function internal structures. RM has the capacity to describe high-level knowledge and encode non-Markovian reward functions. We propose a decentralized learning algorithm to tackle computational complexity, called decentralized graph-based reinforcement learning using reward machines (DGRM), that equips each agent with a localized policy, allowing agents to make decisions independently based on the information available to the agents. DGRM uses the actor-critic structure, and we introduce the tabular Q-function for discrete state problems. We show that the dependency of the Q-function on other agents decreases exponentially as the distance between them increases. To further improve efficiency, we also propose the deep DGRM algorithm, using deep neural networks to approximate the Q-function and policy function to solve large-scale or continuous state problems. The effectiveness of the proposed DGRM algorithm is evaluated by three case studies, two wireless communication case studies with independent and dependent reward functions, respectively, and COVID-19 pandemic mitigation. Experimental results show that local information is sufficient for DGRM and agents can accomplish complex tasks with the help of RM. DGRM improves the global accumulated reward by 119% compared to the baseline in the case of COVID-19 pandemic mitigation.