Optimal Continuous State POMDP Planning With Semantic Observations: A Variational Approach

Optimal Continuous State POMDP Planning With Semantic Observations: A Variational Approach
复制标题

DOI:
10.1109/tro.2019.2933720
复制
发表时间:
2018-07
影响因子:
7.8
通讯作者:
Luke Burks;Ian Loefgren;N. Ahmed
Luke Burks;Ian Loefgren;N. Ahmed
中科院分区:
计算机科学1区
文献类型:
--
作者:
Luke Burks;Ian Loefgren;N. Ahmed

文献摘要

被引文献

相似文献

本文使用连续状态部分可观察马尔可夫决策过程 (CPOMDP) 开发了通过语义观察进行优化规划的新颖策略。提出了与高斯混合 (GM) CPOMDP 策略逼近方法相关的两项主要创新。虽然现有方法具有许多理想的理论特性,但它们无法有效地表示和推理混合连续离散概率模型。第一个重大创新是使用连续离散语义观察概率的 Softmax 模型推导基于点的值迭代贝尔曼策略备份的封闭式变分贝叶斯 GM 近似。这种方法的一个主要好处是可以在复杂的非高斯不确定性下执行动态决策任务,同时还利用连续动态状态空间模型(从而避免繁琐且昂贵的离散化)。第二个重大创新是一种基于聚类的新混合压缩技术,可以很好地扩展到非常大的 GM 政策函数和信念函数。具有语义观察的目标搜索和拦截任务的模拟结果表明,这些创新产生的 GM 策略比其他最先进的策略近似产生的策略更有效,但需要的建模开销和在线运行时成本显着减少。其他结果显示了这种方法对模型误差和扩展到更高维度的鲁棒性。
This article develops novel strategies for optimal planning with semantic observations using continuous state partially observable Markov decision processes (CPOMDPs). Two major innovations are presented in relation to Gaussian mixture (GM) CPOMDP policy approximation methods. While existing methods have many desirable theoretical properties, they are unable to efficiently represent and reason over hybrid continuous-discrete probabilistic models. The first major innovation is the derivation of closed-form variational Bayes GM approximations of point-based value iteration Bellman policy backups, using softmax models of continuous-discrete semantic observation probabilities. A key benefit of this approach is that dynamic decision making tasks can be performed with complex non-Gaussian uncertainties, while also exploiting continuous dynamic state-space models (thus avoiding cumbersome and costly discretization). The second major innovation is a new clustering-based technique for mixture condensation that scales well to very large GM policy functions and belief functions. Simulation results for a target search and interception task with semantic observations show that the GM policies resulting from these innovations are more effective than those produced by other state-of-the-art policy approximations, but require significantly less modeling overhead and online runtime cost. Additional results show the robustness of this approach to model errors and scaling to higher dimensions.