TD algorithm for the variance of return and mean-variance reinforcement learning

TD algorithm for the variance of return and mean-variance reinforcement learning
复制标题

用于收益方差和均值方差强化学习的 TD 算法

DOI:
10.1527/tjsai.16.353
复制
发表时间:
2001
影响因子:
--
通讯作者:
S. Kobayashi
S. Kobayashi
中科院分区:
--
文献类型:
--
作者:
Makoto Sato;H. Kimura;S. Kobayashi

文献摘要

被引文献

相似文献

估计收益的概率分布为马尔可夫环境下的控制问题提供了各种复杂的决策方案,包括风险敏感控制、环境的有效探索等。然而,许多强化学习算法仅仅依赖于预期回报。本文提出了一种利用收益分布均值和方差进行决策的方案。本文提出了一种用于估计MDP(Markov决策过程)环境下收益方差的TD算法和一种基于梯度的方差惩罚准则的强化学习算法,方差惩罚准则是风险规避控制中的典型准则。实证结果证明了算法的行为,并验证了风险规避顺序决策任务的准则。
Estimating probability distributions on returns provides various sophisticated decision making schemes for control problems in Markov environments, including risk-sensitive control, efficient exploration of environments and so on. Many reinforcement learning algorithms, however, have simply relied on the expected return. This paper provides a scheme of decision making using mean and variance of returndistributions. This paper presents a TD algorithm for estimating the variance of return in MDP(Markov decision processes) environments and a gradient-based reinforcement learning algorithm on the variance penalized criterion, which is a typical criterion in risk-avoiding control. Empirical results demonstrate behaviors of the algorithms and validates of the criterion for risk-avoiding sequential decision tasks.