Risk-Averse Decision Making Under Uncertainty

Risk-Averse Decision Making Under Uncertainty
复制标题

DOI:
10.1109/tac.2023.3264178
复制
发表时间:
2021-09
影响因子:
6.8
通讯作者:
M. Ahmadi;Ugo Rosolia;M. Ingham;R. Murray;A. Ames
M. Ahmadi;Ugo Rosolia;M. Ingham;R. Murray;A. Ames
中科院分区:
计算机科学2区
文献类型:
--
作者:
M. Ahmadi;Ugo Rosolia;M. Ingham;R. Murray;A. Ames

文献摘要

被引文献

相似文献

在不确定性问题下的一大类决策可以通过马尔可夫决策过程(MDP)或部分可观测MDP(POMDP)来描述,并应用于人工智能和运筹学等。在这篇文章中,我们考虑的问题,设计政策的MDPs和POMDPs的目标和约束的动态一致的风险措施,而不是传统的总期望,我们称之为约束风险厌恶问题。我们的贡献可以描述如下:首先,对于MDP,在一些温和的假设下,我们提出了一个基于优化的方法来合成马尔可夫策略。然后,我们证明,这样的政策可以通过求解差分凸规划(DCP)。我们表明,我们的配方推广线性规划的总折扣预期成本和约束的约束MDPs;第二,POMDPs,我们表明,如果一致的风险措施可以被定义为一个马尔可夫风险转移映射,无限维优化可以用来设计马尔可夫信念为基础的政策。对于随机有限状态控制器(FSC)的POMDPs,我们表明,后者的优化简化为(有限维)DCP。我们将这些DCP的政策迭代算法设计风险规避FSC POMDPs。我们证明了所提出的方法的有效性与数值实验,涉及条件价值的风险和熵值的风险措施。
A large class of decision making under uncertainty problems can be described via Markov decision processes (MDPs) or partially observable MDPs (POMDPs), with application to artificial intelligence and operations research, among others. In this article, we consider the problem of designing policies for MDPs and POMDPs with objectives and constraints in terms of dynamic coherent risk measures rather than the traditional total expectation, which we refer to as the constrained risk-averse problem. Our contributions can be described as follows: first, for MDPs, under some mild assumptions, we propose an optimization-based method to synthesize Markovian policies. We then demonstrate that such policies can be found by solving difference convex programs (DCPs). We show that our formulation generalize linear programs for constrained MDPs with total discounted expected costs and constraints; second, for POMDPs, we show that, if the coherent risk measures can be defined as a Markov risk transition mapping, an infinite-dimensional optimization can be used to design Markovian belief-based policies. For POMDPs with stochastic finite-state controllers (FSCs), we show that the latter optimization simplifies to a (finite dimensional) DCP. We incorporate these DCPs in a policy iteration algorithm to design risk-averse FSCs for POMDPs. We demonstrate the efficacy of the proposed method with numerical experiments involving conditional-value-at-risk and entropic-value-at-risk risk measures.