The variance of discounted Markov decision processes

The variance of discounted Markov decision processes
复制标题

DOI:
10.2307/3213832
复制
发表时间:
1982-12
影响因子:
1
通讯作者:
M. J. Sobel
M. J. Sobel
中科院分区:
数学4区
文献类型:
--
作者:
M. J. Sobel

文献摘要

被引文献

相似文献

给出了有限马尔可夫决策过程中单阶段报酬现值的方差和高阶矩的计算公式。对于半马尔可夫决策过程,也给出了类似的公式。有一个简短的讨论,在算法中使用方差公式来最大化平均值减去标准差的倍数的障碍。
Formulae are presented for the variance and higher moments of the present value of single-stage rewards in a finite Markov decision process. Similar formulae are exhibited for a semi-Markov decision process. There is a short discussion of the obstacles to using the variance formula in algorithms to maximize the mean minus a multiple of the standard deviation.