Decentralized MCTS via Learned Teammate Models

Decentralized MCTS via Learned Teammate Models
复制标题

DOI:
10.24963/ijcai.2020/12
复制
发表时间:
2020-03
期刊:
--
影响因子:
--
通讯作者:
A. Czechowski;F. Oliehoek
A. Czechowski;F. Oliehoek
中科院分区:
其他
文献类型:
--
作者:
A. Czechowski;F. Oliehoek

文献摘要

被引文献

相似文献

由于改进的可扩展性和健壮性,分散的在线计划可能是协作多智能体系统的一个有吸引力的范例。这种方法的一个关键困难在于对其他代理人的决定做出准确的预测。在本文中,我们提出了一种基于分散蒙特卡罗树搜索的可训练在线分散计划算法,并结合从以前的情节运行中学习的队友模型。通过一次只允许一个代理调整其模型,在理想策略近似的假设下,我们的方法的连续迭代保证改进联合策略,并最终导致收敛到纳什均衡。我们通过在[Claes等人,2015]中介绍的空间任务分配环境的几个场景中进行实验,测试了算法的效率。我们表明,深度学习和卷积神经网络可以用来产生利用问题的空间特征的精确策略逼近器,并且对于特别具有挑战性的域配置,所提出的算法比基线规划性能有所改善。
Decentralized online planning can be an attractive paradigm for cooperative multi-agent systems, due to improved scalability and robustness. A key difficulty of such approach lies in making accurate predictions about the decisions of other agents. In this paper, we present a trainable online decentralized planning algorithm based on decentralized Monte Carlo Tree Search, combined with models of teammates learned from previous episodic runs. By only allowing one agent to adapt its models at a time, under the assumption of ideal policy approximation, successive iterations of our method are guaranteed to improve joint policies, and eventually lead to convergence to a Nash equilibrium. We test the efficiency of the algorithm by performing experiments in several scenarios of the spatial task allocation environment introduced in [Claes et al., 2015]. We show that deep learning and convolutional neural networks can be employed to produce accurate policy approximators which exploit the spatial features of the problem, and that the proposed algorithm improves over the baseline planning performance for particularly challenging domain configurations.