Hybrid BDI-POMDP Framework for Multiagent Teaming

Hybrid BDI-POMDP Framework for Multiagent Teaming
复制标题

用于多代理分组的混合 BDI-POMDP 框架

DOI:
--
复制
发表时间:
2011
影响因子:
5
通讯作者:
Milind Tambe
Milind Tambe
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ranjit R. Nair;Milind Tambe

文献摘要

被引文献

相似文献

目前许多大规模的多智能体团队实现的特点是遵循“信念-愿望-意图”(BDI)范式,明确表示团队计划。尽管他们的承诺,目前的BDI团队方法缺乏工具,定量性能分析的不确定性。分布式部分可观测马尔可夫决策问题(POMDPs)非常适合这样的分析,但在这样的模型中找到最优策略的复杂性是非常棘手的。这篇文章的主要贡献是一个混合BDI-POMDP的方法,其中BDI团队计划被利用来提高POMDP的易处理性和POMDP分析提高BDI团队计划的性能。 具体来说,我们专注于角色分配,在BDI团队的一个基本问题:代理分配到团队中的不同角色。本文提供了三个关键贡献。首先,我们描述了一个角色分配技术,考虑到未来的不确定性,在域中,以前的工作在多智能体角色分配未能解决这些不确定性。为此,我们引入RMTDP(基于角色的马尔可夫团队决策问题),一个新的分布式POMDP模型的角色分配分析。我们的技术收益的易处理性显着减少RMTDP政策搜索,特别是,BDI团队计划提供不完整的RMTDP政策,和RMTDP政策搜索填补了空白,在这种不完整的政策,通过寻找最佳的角色分配。我们的第二个关键贡献是一种新的分解技术,以进一步提高RMTDP策略搜索效率。即使仅限于搜索角色分配,仍然存在组合的许多角色分配,并且在RMTDP中评估每个角色以确定最佳分配是极其困难的。我们的分解技术利用BDI团队计划中的结构来显著修剪角色分配的搜索空间。我们的第三个关键贡献是一个显着更快的政策评估算法适合我们的BDI-POMDP混合方法。最后,我们也提出了实验结果,从两个领域:使命演习模拟和RoboCupRescue灾难救援模拟。
Many current large-scale multiagent team implementations can be characterized as following the "belief-desire-intention" (BDI) paradigm, with explicit representation of team plans. Despite their promise, current BDI team approaches lack tools for quantitative performance analysis under uncertainty. Distributed partially observable Markov decision problems (POMDPs) are well suited for such analysis, but the complexity of finding optimal policies in such models is highly intractable. The key contribution of this article is a hybrid BDI-POMDP approach, where BDI team plans are exploited to improve POMDP tractability and POMDP analysis improves BDI team plan performance. Concretely, we focus on role allocation, a fundamental problem in BDI teams: which agents to allocate to the different roles in the team. The article provides three key contributions. First, we describe a role allocation technique that takes into account future uncertainties in the domain; prior work in multiagent role allocation has failed to address such uncertainties. To that end, we introduce RMTDP (Role-based Markov Team Decision Problem), a new distributed POMDP model for analysis of role allocations. Our technique gains in tractability by significantly curtailing RMTDP policy search; in particular, BDI team plans provide incomplete RMTDP policies, and the RMTDP policy search fills the gaps in such incomplete policies by searching for the best role allocation. Our second key contribution is a novel decomposition technique to further improve RMTDP policy search efficiency. Even though limited to searching role allocations, there are still combinatorially many role allocations, and evaluating each in RMTDP to identify the best is extremely difficult. Our decomposition technique exploits the structure in the BDI team plans to significantly prune the search space of role allocations. Our third key contribution is a significantly faster policy evaluation algorithm suited for our BDI-POMDP hybrid approach. Finally, we also present experimental results from two domains: mission rehearsal simulation and RoboCupRescue disaster rescue simulation.