Impact of Heterogeneity and Risk Aversion on Task Allocation in Multi-Agent Teams

Impact of Heterogeneity and Risk Aversion on Task Allocation in Multi-Agent Teams
复制标题

DOI:
10.1109/lra.2021.3097259
复制
发表时间:
2021-10
影响因子:
5.2
通讯作者:
Haochen Wu;Amin Ghadami;A. E. Bayrak;J. Smereka;B. Epureanu
Haochen Wu;Amin Ghadami;A. E. Bayrak;J. Smereka;B. Epureanu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Haochen Wu;Amin Ghadami;A. E. Bayrak;J. Smereka;B. Epureanu

文献摘要

被引文献

相似文献

多智能体协作决策是一个普遍存在的问题,有着广泛的应用。在许多实际应用中,人们希望设计一个多代理团队与异构组成的代理可以有不同的能力和水平的风险承受能力,以满足不同的要求。虽然在多代理团队的异质性提供了好处,新的挑战出现,包括如何找到最佳的异质团队组合,以及如何在复杂的操作代理之间动态分配任务。在这项工作中,我们开发了一个人工智能框架,多代理异构团队动态学习代理之间的任务分配,通过强化学习。该框架扩展了分散式部分可观测马尔可夫决策过程(Dec-POMDP),使其能够兼容各种类型的异质性模型。我们展示了我们的方法与救灾方案的基准问题。分析了在不确定环境中,Agent能力和决策策略的异质性和风险规避对多Agent团队性能的影响。研究结果表明,在不确定环境中,设计良好的异质团队的表现优于同质团队,具有更高的适应性。
Cooperative multi-agent decision-making is a ubiquitous problem with many real-world applications. In many practical applications, it is desirable to design a multi-agent team with a heterogeneous composition where the agents can have different capabilities and levels of risk tolerance to address diverse requirements. While heterogeneity in multi-agent teams offers benefits, new challenges arise including how to find optimal heterogeneous team compositions and how to dynamically distribute tasks among agents in complex operations. In this work, we develop an artificial intelligence framework for multi-agent heterogeneous teams to dynamically learn task distributions among agents through reinforcement learning. The framework extends Decentralized Partially Observable Markov Decision Processes (Dec-POMDP) to be compatible to model various types of heterogeneity. We demonstrate our approach with a benchmark problem on a disaster relief scenario. The effect of heterogeneity and risk aversion in agent capabilities and decision-making strategies on the performance of multi-agent teams in uncertain environments is analyzed. Results show that a well-designed heterogeneous team outperforms its homogeneous counterpart and possesses higher adaptivity in uncertain environments.