More Risk-Sensitive Markov Decision Processes

More Risk-Sensitive Markov Decision Processes
复制标题

DOI:
10.1287/moor.2013.0601
复制
发表时间:
2014-02
期刊:
Math. Oper. Res.
影响因子:
--
通讯作者:
N. Bäuerle;U. Rieder
N. Bäuerle;U. Rieder
中科院分区:
其他
文献类型:
--
作者:
N. Bäuerle;U. Rieder

文献摘要

被引文献

相似文献

我们研究了马尔可夫决策过程MDP在有限和无限范围内最小化总费用或折扣费用的确定性等价问题。与风险中性的决策者不同,这种优化标准考虑了成本的可变性。作为特例,它包含了具有指数效用的经典风险敏感优化准则。我们证明了这个优化问题可以用一个具有扩展状态空间的普通MDP来求解,并给出了存在最优策略的条件。在无限时间域的情况下,我们证明了最小折扣代价可以通过值迭代得到,并且可以刻画为使用“三明治”引理的不动点方程的唯一解。有趣的是,在幂效用的情况下,这个问题简化了,并且与指数效用的情形具有相似的复杂性,然而到目前为止文献中还没有处理过。我们还证明了政策改进方法的有效性和收敛性。通过一个简单的数值例子,即经典的重复赌场博弈,说明了确定性当量及其参数的影响。最后,还对平均成本问题进行了研究。令人惊讶的是,对于凸幂效用,在适当的重现条件下,最小平均成本不依赖于效用函数的参数,并且等于风险中性的平均成本。这与具有指数效用的经典风险敏感标准形成了鲜明对比。
We investigate the problem of minimizing a certainty equivalent of the total or discounted cost over a finite and an infinite horizon that is generated by a Markov decision process MDP. In contrast to a risk-neutral decision maker this optimization criterion takes the variability of the cost into account. It contains as a special case the classical risk-sensitive optimization criterion with an exponential utility. We show that this optimization problem can be solved by an ordinary MDP with extended state space and give conditions under which an optimal policy exists. In the case of an infinite time horizon we show that the minimal discounted cost can be obtained by value iteration and can be characterized as the unique solution of a fixed-point equation using a “sandwich” argument. Interestingly, it turns out that in the case of a power utility, the problem simplifies and is of similar complexity than the exponential utility case, however has not been treated in the literature so far. We also establish the validity and convergence of the policy improvement method. A simple numerical example, namely, the classical repeated casino game, is considered to illustrate the influence of the certainty equivalent and its parameters. Finally, the average cost problem is also investigated. Surprisingly, it turns out that under suitable recurrence conditions on the MDP for convex power utility, the minimal average cost does not depend on the parameter of the utility function and is equal to the risk-neutral average cost. This is in contrast to the classical risk-sensitive criterion with exponential utility.