Experimental Demonstration of Adaptive MDP-Based Planning with Model Uncertainty

Experimental Demonstration of Adaptive MDP-Based Planning with Model Uncertainty
复制标题

具有模型不确定性的基于自适应 MDP 规划的实验演示

DOI:
--
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
J. How
J. How
中科院分区:
--
文献类型:
--
作者:
Brett Bethke;L. Bertuccelli;J. How

文献摘要

被引文献

相似文献

马尔可夫决策过程(mdp)是解决多智能体规划问题的自然框架,因为它可以模拟随机系统动力学和智能体之间的相互依赖性。在这些方法中,系统的精确建模非常重要,因为错误建模可能导致性能严重下降(即车辆损失)。此外,在许多感兴趣的问题中,在系统开始运行之前可能很难或不可能获得准确的模型;相反,该模型必须在线估计。因此,一种能够在线估计系统模型和调整系统控制策略的自适应机制可以比静态(无线)方法提高性能。本文提出了一个多智能体持续监视问题的MDP公式,并在仿真中说明了系统精确建模的重要性。然后讨论了一种由贝叶斯模型估计器和连续运行的MDP求解器组成的自适应机制。最后,我们给出了麻省理工学院RAVEN测试平台的硬件飞行结果,清楚地证明了这种自适应方法在持续监视问题中的性能优势。
Markov decision processes (MDPs) are a natural framework for solving multiagent planning problems since they can model stochastic system dynamics and interdependencies between agents. In these approaches, accurate modeling of the system in question is important, since mismodeling may lead to severely degraded performance (i.e. loss of vehicles). Furthermore, in many problems of interest, it may be dicult or impossible to obtain an accurate model before the system begins operating; rather, the model must be estimated online. Therefore, an adaptation mechanism that can estimate the system model and adjust the system control policy online can improve performance over a static (o-line) approach. This paper presents an MDP formulation of a multi-agent persistent surveillance problem and shows, in simulation, the importance of accurate modeling of the system. An adaptation mechanism, consisting of a Bayesian model estimator and a continuouslyrunning MDP solver, is then discussed. Finally, we present hardware flight results from the MIT RAVEN testbed that clearly demonstrate the performance benefits of this adaptive approach in the persistent surveillance problem.