Indefinite-Horizon POMDPs with Action-Based Termination

Indefinite-Horizon POMDPs with Action-Based Termination
复制标题

具有基于操作终止的无限期 POMDP

DOI:
--
复制
发表时间:
2007
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
E. Hansen
E. Hansen
中科院分区:
--
文献类型:
--
作者:
E. Hansen

文献摘要

被引文献

相似文献

对于具有不确定范围的决策理论规划问题,计划执行在有限数量的步骤之后以概率1终止,但是直到终止的步骤的数量(即,地平线)是不确定的和无界的。在对此类问题建模的传统方法(称为随机最短路径问题)中,当达到特定状态(通常是目标状态)时,计划执行就会终止。我们考虑一个模型,其中计划执行终止时,采取停止行动。我们表明,一个基于行动的终止模型有几个优点,部分可观察的规划问题。它不要求目标状态是完全可观察的;它不要求保证目标状态的实现;它允许更容易地找到适当的策略。这个框架允许许多部分可观察的规划问题,以更现实的方式,不需要一个人为的折扣因子建模。
For decision-theoretic planning problems with an indefinite horizon, plan execution terminates after a finite number of steps with probability one, but the number of steps until termination (i.e., the horizon) is uncertain and unbounded. In the traditional approach to modeling such problems, called a stochastic shortest-path problem, plan execution terminates when a particular state is reached, typically a goal state. We consider a model in which plan execution terminates when a stopping action is taken. We show that an action-based model of termination has several advantages for partially observable planning problems. It does not require a goal state to be fully observable; it does not require achievement of a goal state to be guaranteed; and it allows a proper policy to be found more easily. This framework allows many partially observable planning problems to be modeled in a more realistic way that does not require an artificial discount factor.