The variance of discounted Markov decision processes
The variance of discounted Markov decision processes
复制标题
DOI:
10.2307/3213832
复制
发表时间:
1982-12
影响因子:
1
通讯作者:
M. J. Sobel
中科院分区:
文献类型:
--
作者:
M. J. Sobel
Formulae are presented for the variance and higher moments of the present value of single-stage rewards in a finite Markov decision process. Similar formulae are exhibited for a semi-Markov decision process. There is a short discussion of the obstacles to using the variance formula in algorithms to maximize the mean minus a multiple of the standard deviation.