Learning Robot Swarm Tactics over Complex Adversarial Environments

Learning Robot Swarm Tactics over Complex Adversarial Environments
复制标题

DOI:
10.1109/mrs50823.2021.9620707
复制
发表时间:
2021-09
期刊:
2021 International Symposium on Multi-Robot and Multi-Agent Systems (MRS)
影响因子:
--
通讯作者:
A. Behjat;H. Manjunatha;Prajit KrisshnaKumar;Apurv Jani;Leighton Collins;P. Ghassemi;Joseph P. Distefano-Joseph
A. Behjat;H. Manjunatha;Prajit KrisshnaKumar;Apurv Jani;Leighton Collins;P. Ghassemi;Joseph P. Distefano-Joseph
中科院分区:
其他
文献类型:
--
作者:
A. Behjat;H. Manjunatha;Prajit KrisshnaKumar;Apurv Jani;Leighton Collins;P. Ghassemi;Joseph P. Distefano-Joseph

文献摘要

相似文献

为了在现实世界中完成复杂的群体机器人任务,需要规划和执行单个机器人行为、任务分配、路径规划和编队控制等群体原语以及目标搜索和群体覆盖等任务特定目标的组合。大多数此类任务都是由机器人专家团队手动设计的。最近在学习群体行为的自动化方法方面的工作仅限于学习完成任务的稀疏工作的单个原语。本文提出了一种系统的方法,利用具有特殊输入和输出编码的神经网络,学习由群元组成的战术任务特定策略,以有效地完成任务。为了在对抗环境中学习群体策略,我们采用了以下方法的组合:1)地图到图的抽象;2)通过兴趣点的帕累托滤波和机器人聚类进行输入/输出编码;3)通过神经进化和策略梯度方法进行学习。我们说明这种组合对于提供可处理的学习至关重要,特别是考虑到模拟这种规模和复杂性的群体任务的计算成本。通过多达60个机器人演示成功完成任务的结果。此外,训练和测试场景中性能统计数据的密切匹配显示了所提出框架的潜在推广性。
To accomplish complex swarm robotic missions in the real world, one needs to plan and execute a combination of single robot behaviors, group primitives such as task allocation, path planning, and formation control, and mission-specific objectives such as target search and group coverage. Most such missions are designed manually by teams of robotics experts. Recent work in automated approaches to learning swarm behavior has been limited to individual primitives with sparse work on learning complete missions. This paper presents a systematic approach to learn tactical mission-specific policies that compose primitives in a swarm to accomplish the mission efficiently using neural networks with special input and output encoding. To learn swarm tactics in an adversarial environment, we employ a combination of 1) map-to-graph abstraction, 2) input/output encoding via Pareto filtering of points of interest and clustering of robots, and 3) learning via neuroevolution and policy gradient approaches. We illustrate this combination as critical to providing tractable learning, especially given the computational cost of simulating swarm missions of this scale and complexity. Successful mission completion outcomes are demonstrated with up to 60 robots. In addition, a close match in the performance statistics in training and testing scenarios shows the potential generalizability of the proposed framework.