Scalable Decision-Theoretic Planning in Open and Typed Multiagent Systems

Scalable Decision-Theoretic Planning in Open and Typed Multiagent Systems
复制标题

DOI:
10.1609/aaai.v34i05.6200
复制
发表时间:
2019-11
期刊:
--
影响因子:
--
通讯作者:
A. Eck;Maulik Shah;Prashant Doshi;Leen-Kiat Soh
A. Eck;Maulik Shah;Prashant Doshi;Leen-Kiat Soh
中科院分区:
其他
文献类型:
--
作者:
A. Eck;Maulik Shah;Prashant Doshi;Leen-Kiat Soh

文献摘要

相似文献

在开放式代理系统中,合作或竞争的代理集会随着时间的推移而变化,并且变化的方式是不可预测的。例如,如果协作机器人的任务是扑灭野火,它们可能会耗尽抑制剂,暂时无法帮助它们的同伴。我们认为,在这些背景下的规划问题与代理人无法相互沟通,有很多人的额外挑战。因为一个代理的最佳行动取决于其他代理的行动,每个代理不仅要预测其对等体的行动,而且在此之前,还要推理它们是否存在以执行某个行动。因此,解决开放性问题需要代理人对彼此的存在进行建模,这在大量代理人的情况下变得难以计算。我们提出了一种新颖的,原则性的,可扩展的方法,在这种情况下,使代理人的原因,在其共享的环境和他们的行动有关的其他人的存在。我们的方法外推模型的几个同行的多代理系统的整体行为,并结合它与推广的蒙特卡洛树搜索,在多代理的开放环境中执行个人代理推理。理论分析建立代理模型的数量,以实现可接受的最坏情况下的外推误差的界限,以及遗憾的界限,代理的效用,从建模只有一些邻居。多代理野火抑制问题的模拟表明,我们的方法的有效性相比,替代基线。
In open agent systems, the set of agents that are cooperating or competing changes over time and in ways that are nontrivial to predict. For example, if collaborative robots were tasked with fighting wildfires, they may run out of suppressants and be temporarily unavailable to assist their peers. We consider the problem of planning in these contexts with the additional challenges that the agents are unable to communicate with each other and that there are many of them. Because an agent's optimal action depends on the actions of others, each agent must not only predict the actions of its peers, but, before that, reason whether they are even present to perform an action. Addressing openness thus requires agents to model each other's presence, which becomes computationally intractable with high numbers of agents. We present a novel, principled, and scalable method in this context that enables an agent to reason about others' presence in its shared environment and their actions. Our method extrapolates models of a few peers to the overall behavior of the many-agent system, and combines it with a generalization of Monte Carlo tree search to perform individual agent reasoning in many-agent open environments. Theoretical analyses establish the number of agents to model in order to achieve acceptable worst case bounds on extrapolation error, as well as regret bounds on the agent's utility from modeling only some neighbors. Simulations of multiagent wildfire suppression problems demonstrate our approach's efficacy compared with alternative baselines.