Entropic Risk Optimization in Discounted MDPs

Entropic Risk Optimization in Discounted MDPs
复制标题

DOI:
--
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
J. Hau;Marek Petrik;M. Ghavamzadeh
J. Hau;Marek Petrik;M. Ghavamzadeh
中科院分区:
其他
文献类型:
--
作者:
J. Hau;Marek Petrik;M. Ghavamzadeh

文献摘要

被引文献

相似文献

风险厌恶马尔可夫决策过程(MDP)具有最优策略,可以在低可变性的情况下获得高收益,但这些MDP通常很难求解。只有少数风险厌恶的目标承认动态规划(DP)制定,这是大多数MDP和RL算法的支柱。我们推导出一个新的DP公式贴现风险厌恶MDP熵风险度量(ERM)和熵值风险(EVaR)的目标。我们的DP制定的ERM,这是可能的,因为我们的新定义的价值函数与时间相关的风险水平,可以近似最优政策的时间是多项式的近似误差。然后,我们使用的ERM算法优化的EVaR目标在多项式时间内使用优化的离散化方案。我们的数值结果表明,我们的配方和算法在折扣MDPs的可行性。
Risk-averse Markov Decision Processes (MDPs) have optimal policies that achieve high returns with low variability, but these MDPs are often difficult to solve. Only a few risk-averse objectives admit a dynamic programming (DP) formulation, which is the mainstay of most MDP and RL algorithms. We derive a new DP formulation for discounted risk-averse MDPs with En-tropic Risk Measure (ERM) and Entropic Value at Risk (EVaR) objectives. Our DP formulation for ERM, which is possible because of our novel definition of value function with time-dependent risk levels, can approximate optimal policies in a time that is polynomial in the approximation error. We then use the ERM algorithm to optimize the EVaR objective in polynomial time using an optimized discretization scheme. Our numerical results show the viability of our formulations and algorithms in discounted MDPs.