Observation-Based Optimization for POMDPs With Continuous State, Observation, and Action Spaces

Observation-Based Optimization for POMDPs With Continuous State, Observation, and Action Spaces
复制标题

具有连续状态、观测和动作空间的 POMDP 基于观测的优化

DOI:
10.1109/tac.2018.2861910
复制
发表时间:
2019-05-01
影响因子:
6.8
通讯作者:
Xi, Hongsheng
Xi, Hongsheng
中科院分区:
计算机科学2区
文献类型:
--
作者:
Jiang, Xiaofeng;Yang, Jian;Xi, Hongsheng

文献摘要

被引文献

相似文献

研究了状态、观测和动作空间连续的部分可观测马尔可夫决策过程的优化问题。基于离散空间的POMDPs是解决不完全状态信息决策系统的一种有效方法。然而,在最近的POMDPs的应用中,有许多问题具有连续的状态,观察和行动。对于这类问题,由于信念空间的无穷维性,现有的研究通常将具有充分或不充分统计量的连续空间离散化,这可能导致维数灾难和性能下降。在本文中,基于性能指标的灵敏度分析,我们已经开发了一个基于模拟的策略迭代算法,以找到局部最优的POMDPs的基于观测的政策与连续空间。该算法不需要任何特定的假设条件和先验信息,计算复杂度低。复杂多输入多输出波束形成问题的数值例子表明,该算法具有显着的性能改善。
This paper considers the optimization problem for partially observable Markov decision processes (POMDPs) with the continuous state, observation, and action spaces. POMDPs with the discrete spaces have emerged as a promising approach to the decision systems with imperfect state information. However, in recent applications of POMDPs, there are many problems that have continuous states, observations, and actions. For such problems, due to the infinite dimensionality of the belief space, the existing studies usually discretize the continuous spaces with the sufficient or nonsufficient statistics, which may cause the curse of dimensionality and performance degradation. In this paper, based on the sensitivity analysis of the performance criteria, we have developed a simulation-based policy iteration algorithm to find the local optimal observation-based policy for POMDPs with the continuous spaces. The proposed algorithm needs none of the specific assumptions and prior information, and has a low computational complexity. One numerical example of the complicated multiple-input multiple-output beamforming problem shows that the algorithm has a significant performance improvement.