Mean-variance optimization of discrete time discounted Markov decision processes

Mean-variance optimization of discrete time discounted Markov decision processes
复制标题

离散时间贴现马尔可夫决策过程的均值-方差优化

DOI:
10.1016/j.automatica.2017.11.012
复制
发表时间:
2017-08
期刊:
影响因子:
6.4
通讯作者:
Li Xia
Li Xia
中科院分区:
计算机科学2区
文献类型:
--
作者:
Li Xia

文献摘要

参考文献

相似文献

本文研究无限时间离散折扣马尔可夫决策过程(MDP)中的均值-方差优化问题。目标是在平均性能约束下使系统报酬的方差最小。与文献中大多数要求平均性能已经达到最优的工作不同,我们可以使折现性能等于任何常数。这个问题的困难是由于方差函数的二次型,使得方差最小化问题不是标准的MDP问题。通过证明可行策略空间的可分解结构,在新的折扣准则和新的奖励函数下,我们将约束方差最小化问题转化为等价的无约束MDP问题。在任意两个可行策略下,马尔可夫链的方差的差异由一个差分公式来量化。基于方差差公式,提出了一种寻找最优策略的策略迭代算法。我们还证明了确定性策略相对于均值约束策略空间中生成的随机策略的最优性。数值实验证明了该方法的有效性。
In this paper, we study a mean–variance optimization problem in an infinite horizon discrete time discounted Markov decision process (MDP). The objective is to minimize the variance of system rewards with the constraint of mean performance. Different from most of works in the literature which require the mean performance already achieve optimum, we can let the discounted performance equal any constant. The difficulty of this problem is caused by the quadratic form of the variance function which makes the variance minimization problem not a standard MDP. By proving the decomposable structure of the feasible policy space, we transform this constrained variance minimization problem to an equivalent unconstrained MDP under a new discounted criterion and a new reward function. The difference of the variances of Markov chains under any two feasible policies is quantified by a difference formula. Based on the variance difference formula, a policy iteration algorithm is developed to find the optimal policy. We also prove the optimality of deterministic policy over the randomized policy generated in the mean-constrained policy space. Numerical experiments demonstrate the effectiveness of our approach.
DOI: --
发表时间: 2012-11
期刊: 2012 2nd Australian Control Conference
影响因子: --
作者:
Yonghao Huang;Xi Chen
通讯作者: Yonghao Huang;Xi Chen
DOI: 10.2307/3213832
发表时间: 1982-12
影响因子: 1
作者:
M. J. Sobel
通讯作者: M. J. Sobel
DOI: 10.1007/978-3-658-27956-1_2
发表时间: 2019
期刊: Finanzwirtschaft, Banken und Bankmanagement I Finance, Banks and Bank Management
影响因子: --
作者:
Gevorg Hunanyan
通讯作者: Gevorg Hunanyan
DOI: 10.1287/opre.42.1.184
发表时间: 1994-02
期刊: Oper. Res.
影响因子: --
作者:
Kun-Jen Chung
通讯作者: Kun-Jen Chung
贴现连续时间马尔可夫决策过程的风险概率准则
DOI: 10.1007/s10626-017-0257-6
发表时间: 2017-08
期刊: Discrete Event Dyn. Syst.
影响因子: --
作者:
HuoHaifeng;ZouXiaolong;GuoXianping
通讯作者: GuoXianping