Coordinated Multi-Agent Learning for Decentralized POMDPs

Coordinated Multi-Agent Learning for Decentralized POMDPs
复制标题

去中心化 POMDP 的协调多智能体学习

DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
V. Lesser
V. Lesser
中科院分区:
--
文献类型:
--
作者:
Chongjie Zhang;V. Lesser

文献摘要

被引文献

相似文献

在许多多智能体应用中,如分布式传感器网络,智能体网络在不确定性和局部交互下协同工作。网络化分布式POMDP(ND-POMDP)提供了一个框架,这样的合作多智能体决策建模。现有的ND-POMDPs的研究主要集中在需要精确模型的离线学习技术上,而这些离线学习技术在实际应用中通常很难获得.本文提出了一种无模型的、可扩展的学习方法,该方法综合了多智能体强化学习(MARL)和分布式约束优化(DCOP)技术,利用ND-POMDPs中的结构化交互,将联合策略的学习分布化,并利用DCOP技术协调分布式学习,以保证全局学习性能.我们的方法可以学习一个全局最优的政策ND-POMDPs的属性称为groupwise可观性。实验结果表明,在学习和执行过程中进行通信,我们的方法显着优于离线计算的接近最优的非通信策略。
In many multi-agent applications such as distributed sensor nets, a network of agents act collaboratively under uncertainty and local interactions. Networked Distributed POMDP (ND-POMDP) provides a framework to model such cooperative multi-agent decision making. Existing work on ND-POMDPs has focused on offline techniques that require accurate models, which areusuallycostlytoobtaininpractice.Thispaperpresents a model-free, scalable learning approach that synthesizes multi-agent reinforcement learning (MARL) and distributedconstraintoptimization(DCOP).ByexploitingstructuredinteractioninND-POMDPs,ourapproach distributes the learning of the joint policy and employs DCOP techniques to coordinate distributed learning to ensure the global learning performance. Our approach can learn a globally optimal policy for ND-POMDPs with a property called groupwise observability. Experimental results show that, with communication during learning and execution, our approach significantly outperforms the nearly-optimal non-communication policies computed offline.