Risk-sensitive planning in partially observable environments

Risk-sensitive planning in partially observable environments
复制标题

部分可观察环境中的风险敏感规划

DOI:
--
复制
发表时间:
2010
期刊:
Adaptive Agents and Multi-Agent Systems
影响因子:
--
通讯作者:
Pradeep Varakantham
Pradeep Varakantham
中科院分区:
--
文献类型:
--
作者:
J. Marecki;Pradeep Varakantham

文献摘要

被引文献

相似文献

部分可观察马尔可夫决策过程(POMDP)是一种用于部分可观察域不确定性规划的常用框架。然而,POMDP模型是风险中性的,因为它假设代理最大化其行为的预期回报。相反,在财务规划等领域,通常要求代理决策对风险敏感(对于非线性效用函数,代理行为的效用最大化)。不幸的是,现有的POMDP求解器不能精确地解决这样的规划问题。通过考虑效用函数的分段线性逼近,本文在三个方面解决了这一缺点:(i)定义了风险敏感的POMDP模型;(ii)推导了底层价值函数的基本性质,并提供了一种函数值迭代技术来精确计算它们;(c)提出了一种有效的程序来确定主导价值函数,以加快算法的速度。我们的实验表明,所提出的方法是可行的,适用于现实的财务规划领域。
Partially Observable Markov Decision Process (POMDP) is a popular framework for planning under uncertainty in partially observable domains. Yet, the POMDP model is risk-neutral in that it assumes that the agent is maximizing the expected reward of its actions. In contrast, in domains like financial planning, it is often required that the agent decisions are risk-sensitive (maximize the utility of agent actions, for non-linear utility functions). Unfortunately, existing POMDP solvers cannot solve such planning problems exactly. By considering piecewise linear approximations of utility functions, this paper addresses this shortcoming in three contributions: (i) It defines the Risk-Sensitive POMDP model; (ii) It derives the fundamental properties of the underlying value functions and provides a functional value iteration technique to compute them exactly and (c) It proposes an efficient procedure to determine the dominated value functions, to speed up the algorithm. Our experiments show that the proposed approach is feasible and applicable to realistic financial planning domains.