Coordinated Multi-Agent Learning for Decentralized POMDPs
Coordinated Multi-Agent Learning for Decentralized POMDPs
复制标题
去中心化 POMDP 的协调多智能体学习
DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
V. Lesser
中科院分区:
文献类型:
--
作者:
Chongjie Zhang;V. Lesser
In many multi-agent applications such as distributed sensor nets, a network of agents act collaboratively under uncertainty and local interactions. Networked Distributed POMDP (ND-POMDP) provides a framework to model such cooperative multi-agent decision making. Existing work on ND-POMDPs has focused on offline techniques that require accurate models, which areusuallycostlytoobtaininpractice.Thispaperpresents a model-free, scalable learning approach that synthesizes multi-agent reinforcement learning (MARL) and distributedconstraintoptimization(DCOP).ByexploitingstructuredinteractioninND-POMDPs,ourapproach distributes the learning of the joint policy and employs DCOP techniques to coordinate distributed learning to ensure the global learning performance. Our approach can learn a globally optimal policy for ND-POMDPs with a property called groupwise observability. Experimental results show that, with communication during learning and execution, our approach significantly outperforms the nearly-optimal non-communication policies computed offline.