Efficient Planning of Multi-Robot Collective Transport using Graph Reinforcement Learning with Higher Order Topological Abstraction

Efficient Planning of Multi-Robot Collective Transport using Graph Reinforcement Learning with Higher Order Topological Abstraction
复制标题

DOI:
10.1109/icra48891.2023.10161517
复制
发表时间:
2023-03
期刊:
2023 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Steve Paul;Wenyuan Li;B. Smyth;Yuzhou Chen;Y. Gel;Souma Chowdhury
Steve Paul;Wenyuan Li;B. Smyth;Yuzhou Chen;Y. Gel;Souma Chowdhury
中科院分区:
其他
文献类型:
--
作者:
Steve Paul;Wenyuan Li;B. Smyth;Yuzhou Chen;Y. Gel;Souma Chowdhury

文献摘要

被引文献

相似文献

高效的多机器人任务分配(MRTA)是各种时间敏感型应用程序(如灾难响应、仓库操作和建筑)的基础。本文解决了一类特殊的问题,我们称之为mrta -集体运输或MRTA-CT -这里的任务呈现不同的工作量和截止日期,机器人受到飞行范围,通信范围和有效载荷的限制。对于这些涉及100 -1000个任务和10 -100个机器人的问题的大型实例,传统的非学习解决方案通常是时间效率低下的,并且新兴的基于学习的策略在没有昂贵的再培训的情况下不能很好地扩展到更大规模的问题。为了解决这一差距,我们使用了最近提出的包含Capsule网络和多头注意机制的编码器-解码器图神经网络,并创新地添加了拓扑描述符(TD)作为新特征,以提高对类似和更大规模的未见问题的可转移性。采用持久同调法推导TD,采用近端策略优化方法训练TD增广图神经网络。由此产生的策略模型比最先进的非学习基线更有优势,同时速度更快。当扩展到测试比训练中使用的问题更大的问题时,使用TD的好处是显而易见的。
Efficient multi-robot task allocation (MRTA) is fundamental to various time-sensitive applications such as disaster response, warehouse operations, and construction. This paper tackles a particular class of these problems that we call MRTA-collective transport or MRTA-CT - here tasks present varying workloads and deadlines, and robots are subject to flight range, communication range, and payload constraints. For large instances of these problems involving 100s-1000's of tasks and 10s-100s of robots, traditional non-learning solvers are often time-inefficient, and emerging learning-based policies do not scale well to larger-sized problems without costly retraining. To address this gap, we use a recently proposed encoder-decoder graph neural network involving Capsule networks and multi-head attention mechanism, and innovatively add topological descriptors (TD) as new features to improve transferability to unseen problems of similar and larger size. Persistent homology is used to derive the TD, and proximal policy optimization is used to train our TD-augmented graph neural network. The resulting policy model compares favorably to state-of-the-art non-learning baselines while being much faster. The benefit of using TD is readily evident when scaling to test problems of size larger than those used in training.