A Convex Analytic Approach to Risk-Aware Markov Decision Processes

A Convex Analytic Approach to Risk-Aware Markov Decision Processes
复制标题

DOI:
10.1137/140969221
复制
发表时间:
2015-06
期刊:
SIAM J. Control. Optim.
影响因子:
--
通讯作者:
W. Haskell;R. Jain
W. Haskell;R. Jain
中科院分区:
其他
文献类型:
--
作者:
W. Haskell;R. Jain

文献摘要

被引文献

相似文献

在经典马尔可夫决策过程(MDP)理论中,我们寻找一种策略,例如最小化预期无限范围贴现成本。当然,预期是一种风险中性的衡量标准,但在许多应用中,尤其是在金融领域,这还不够。我们用一般风险函数代替期望,并将此类模型称为风险感知 MDP 模型。我们考虑在两种情况下最小化此类风险函数:预期效用框架和条件风险价值(一种流行的连贯风险度量)。随后,我们考虑风险感知 MDP,其中风险用约束来表示。这包括随机优势约束和经典的机会约束优化问题。在每种情况下,我们都开发了一种凸分析方法来解决此类风险意识 MDP。在大多数情况下,我们表明,当我们扩大状态空间时,问题可以表示为占用度量中的无限维线性程序(LP)。我们提供离散化方法和有限近似...
In classical Markov decision process (MDP) theory, we search for a policy that, say, minimizes the expected infinite horizon discounted cost. Expectation is, of course, a risk neutral measure, which does not suffice in many applications, particularly in finance. We replace the expectation with a general risk functional, and call such models risk-aware MDP models. We consider minimization of such risk functionals in two cases, the expected utility framework, and conditional value-at-risk, a popular coherent risk measure. Later, we consider risk-aware MDPs wherein the risk is expressed in the constraints. This includes stochastic dominance constraints, and the classical chance-constrained optimization problems. In each case, we develop a convex analytic approach to solve such risk-aware MDPs. In most cases, we show that the problem can be formulated as an infinite-dimensional linear program (LP) in occupation measures when we augment the state space. We provide a discretization method and finite approximati...