Information-Guided Robotic Maximum Seek-and-Sample in Partially Observable Continuous Environments

Information-Guided Robotic Maximum Seek-and-Sample in Partially Observable Continuous Environments
复制标题

DOI:
10.1109/lra.2019.2929997
复制
发表时间:
2019-10-01
影响因子:
5.2
通讯作者:
Roy, Nicholas
Roy, Nicholas
中科院分区:
计算机科学2区
文献类型:
--
作者:
Flaspohler, Genevieve;Preston, Victoria;Roy, Nicholas

文献摘要

被引文献

相似文献

本文提出了一种基于最大值信息搜索(maximum - value information and Search, PLUMES)的不确定羽流定位方法,该方法用于在先验未知和部分可观测连续环境的全局最大值下定位和收集样本。这种“最大搜索和样本”(MSS)问题在环境和地球科学中普遍存在。专家们希望在环境极值处(例如,石油泄漏源)收集有科学价值的样本,但对这种现象的分布没有事先的了解。我们将MSS问题表述为具有连续状态和观察空间的部分可观察马尔可夫决策过程(POMDP),以及稀疏奖励信号。为了解决MSS POMDP问题,PLUMES采用了信息论奖励启发式的连续观测蒙特卡罗树搜索来有效地定位和采样全局最大值。在模拟和现场实验中,与最先进的计划人员相比,PLUMES在不同的环境中收集了更多具有科学价值的样本,这些环境包括各种平台、传感器和具有挑战性的现实世界条件。
We present Plume Localization under Uncertainty using Maximum-ValuE information and Search (PLUMES), a planner for localizing and collecting samples at the global maximum of an a priori unknown and partially observable continuous environment. This "maximum seek-and-sample" (MSS) problem is pervasive in the environmental and earth sciences. Experts want to collect scientifically valuable samples at an environmental maximum (e.g., an oil-spill source), but do not have prior knowledge about the phenomenon's distribution. We formulate the MSS problem as a partially-observable Markov decision process (POMDP) with continuous state and observation spaces, and a sparse reward signal. To solve the MSS POMDP, PLUMES uses an information-theoretic reward heuristic with continuous-observation Monte Carlo Tree Search to efficiently localize and sample from the global maximum. In simulation and field experiments, PLUMES collects more scientifically valuable samples than state-of-the-art planners in a diverse set of environments, with various platforms, sensors, and challenging real-world conditions.