TD algorithm for the variance of return and mean-variance reinforcement learning
TD algorithm for the variance of return and mean-variance reinforcement learning
复制标题
用于收益方差和均值方差强化学习的 TD 算法
DOI:
10.1527/tjsai.16.353
复制
发表时间:
2001
影响因子:
--
通讯作者:
S. Kobayashi
中科院分区:
文献类型:
--
作者:
Makoto Sato;H. Kimura;S. Kobayashi
Estimating probability distributions on returns provides various sophisticated decision making schemes for control problems in Markov environments, including risk-sensitive control, efficient exploration of environments and so on. Many reinforcement learning algorithms, however, have simply relied on the expected return. This paper provides a scheme of decision making using mean and variance of returndistributions. This paper presents a TD algorithm for estimating the variance of return in MDP(Markov decision processes) environments and a gradient-based reinforcement learning algorithm on the variance penalized criterion, which is a typical criterion in risk-avoiding control. Empirical results demonstrate behaviors of the algorithms and validates of the criterion for risk-avoiding sequential decision tasks.