Combining a POMDP Abstraction with Replanning to Solve Complex, Position-Dependent Sensing Tasks

Combining a POMDP Abstraction with Replanning to Solve Complex, Position-Dependent Sensing Tasks
复制标题

将 POMDP 抽象与重新规划相结合来解决复杂的、位置相关的传感任务

DOI:
--
复制
发表时间:
2013
期刊:
AAAI Fall Symposia
影响因子:
--
通讯作者:
L. Kavraki
L. Kavraki
中科院分区:
--
文献类型:
--
作者:
D. Grady;Mark Moll;L. Kavraki

文献摘要

被引文献

相似文献

部分可观察马尔可夫决策过程(POMDP)是一个通用框架,用于在嘈杂的行动和感知条件下确定奖励最大化的行动策略。然而,由于所需计算的 PSPACE 完备性,确定 POMDP 的最佳策略对于机器人任务来说通常很棘手。最近推出的几种求解器扩大了可以考虑的问题的规模。尽管这些 POMDP 求解器在理论上可以遵循复杂的运动约束,但我们表明,与依赖于忽略某些运动约束的策略的替代方法相比,计算成本并没有在最终的在线执行中带来好处。我们提倡在关键的地方使用 POMDP 框架——找到一种策略,在给定所有过去的噪声传感器观测的情况下提供最佳动作,同时抽象一些运动约束以减少求解时间。然而,抽象机器人的动作通常无法在其真实运动约束下执行。该问题通过约束较少的 POMDP 离线解决,并通过重新规划在线处理完整系统约束下的导航。我们凭经验证明,与直接在 POMDP 实验中使用的解决类汽车机器人的运动约束相比,使用这种抽象运动模型生成的策略计算速度更快,并且获得类似或更高的奖励。
The Partially-Observable Markov Decision Process (POMDP) is a general framework to determine reward-maximizing action policies under noisy action and sensing conditions. However, determining an optimal policy for POMDPs is often intractable for robotic tasks due to the PSPACE-complete nature of the computation required. Several recent solvers have been introduced that expand the size of problems that can be considered. Although these POMDP solvers can respect complex motion constraints in theory, we show that the computational cost does not provide a benefit in the eventual online execution, compared to our alternative approach that relies on a policy that ignores some of the motion constraints. We advocate using the POMDP framework where it is critical ‐ to find a policy that provides the optimal action given all past noisy sensor observations, while abstracting some of the motion constraints to reduce solution time. However, the actions of an abstract robot are generally not executable under its true motion constraints. The problem is addressed offline with a less-constrained POMDP, and navigation under the full system constraints is handled online with replanning. We empirically demonstrate that the policy generated using this abstracted motion model is faster to compute and achieves similar or higher reward than addressing the motion constraints for a car-like robot as used in our experiments directly in the POMDP.