Probabilistic planning with non-linear utility functions and worst-case guarantees

Probabilistic planning with non-linear utility functions and worst-case guarantees
复制标题

具有非线性效用函数和最坏情况保证的概率规划

DOI:
--
复制
发表时间:
2012
期刊:
Adaptive Agents and Multi-Agent Systems
影响因子:
--
通讯作者:
A. Vladimirsky
A. Vladimirsky
中科院分区:
--
文献类型:
--
作者:
Stefano Ermon;C. Gomes;B. Selman;A. Vladimirsky

文献摘要

被引文献

相似文献

马尔可夫决策过程是最广泛使用的框架之一,制定概率规划问题。由于规划者在高风险的情况下往往是风险敏感的,因此经常引入非线性效用函数来描述他们在所有可能结果中的偏好。另外,对风险敏感的决策者通常要求他们的计划满足某些最坏情况的保证。 我们展示了如何通过考虑最坏情况下的约束,我们最大化总奖励的预期效用的问题,将这两种方法联合收割机。我们推广现有的几个结果的结构的最优政策的约束情况下,无论是有限和无限的地平线问题。我们提供了一个动态规划算法来计算最优策略,我们引入了一个允许的启发式有效地修剪搜索空间。最后,我们使用一个随机最短路径问题的大型现实世界的道路网络,以证明我们的方法的实用性。
Markov Decision Processes are one of the most widely used frameworks to formulate probabilistic planning problems. Since planners are often risk-sensitive in high-stake situations, non-linear utility functions are often introduced to describe their preferences among all possible outcomes. Alternatively, risk-sensitive decision makers often require their plans to satisfy certain worst-case guarantees. We show how to combine these two approaches by considering problems where we maximize the expected utility of the total reward subject to worst-case constraints. We generalize several existing results on the structure of optimal policies to the constrained case, both for finite and infinite horizon problems. We provide a Dynamic Programming algorithm to compute the optimal policy, and we introduce an admissible heuristic to effectively prune the search space. Finally, we use a stochastic shortest path problem on large real-world road networks to demonstrate the practical applicability of our method.