Optimization of parametric policies of Markov decision processes under a variance criterion

Optimization of parametric policies of Markov decision processes under a variance criterion
复制标题

DOI:
10.1109/wodes.2016.7497868
复制
发表时间:
2016-05
期刊:
2016 13th International Workshop on Discrete Event Systems (WODES)
影响因子:
--
通讯作者:
L. Xia
L. Xia
中科院分区:
其他
文献类型:
--
作者:
L. Xia

文献摘要

被引文献

相似文献

方差准则是马尔可夫决策过程中一个不常见但重要的准则。方差函数的非线性(二次)结构引起的非马尔可夫性质使得传统的MDP方法对此问题无效。在本文中,我们研究了方差准则下MDP参数策略的优化,其中优化参数是在每个状态下选择动作的概率。利用基于灵敏度的优化的基本思想,我们推导了奖励方差相对于系统参数的差分公式和导数公式。方差差分公式是该问题的基础,它通过非负项部分解决了方差函数非线性特性的难题。通过这些敏感性公式,我们证明可以在确定性策略空间中找到具有最小方差的最优策略。还导出了最优策略的必要条件。与文献中基于梯度的方法相比,我们的方法可以为这种方差优化问题提供清晰的观点。
The variance criterion is an uncommon while important criterion in Markov decision processes. The non-Markovian property caused by the nonlinear (quadratic) structure of variance function makes the traditional MDP approaches invalid for this problem. In this paper, we study the optimization of parametric policies of MDPs under the variance criterion, where the optimization parameters are the probabilities of selecting actions at each state. With the basic idea of sensitivity-based optimization, we derive a difference formula and a derivative formula of the reward variance with respect to the system parameter. The variance difference formula is fundamental for this problem and it partly handles the difficulty of nonlinear property of variance function through a nonnegative term. With these sensitivity formulas, we prove that the optimal policy with the minimal variance can be found in the deterministic policy space. A necessary condition of the optimal policy is also derived. Compared with the counterpart of gradient-based approaches in the literature, our approach can provide a clear viewpoint for this variance optimization problem.