Ensemble Usage for More Reliable Policy Identification in Reinforcement Learning
Ensemble Usage for More Reliable Policy Identification in Reinforcement Learning
复制标题
使用集成在强化学习中实现更可靠的策略识别
DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
S. Udluft
中科院分区:
文献类型:
--
作者:
A. Hans;S. Udluft
Reinforcement learning (RL) methods employing powerful function approximators like neural networks have become an interesting approach for optimal control. Since they learn a policy from observations, they are also applicable when no analytical description of the system is available. Although impressive results have been reported, their handling in practice is still hard, as they can fail at reliably determining a good policy. In previous work, we used ensembles of policies from independent runs of neural fitted Q-iteration (NFQ) to produce successful policies more reliably. In this paper we evaluate the approach on more problems and propose to form ensembles from successive iterations of a single NFQ run as a computationally cheap alternative to completely independent runs.
DOI:
10.1007/3-540-45014-9
发表时间:
2000-06
期刊:
--
影响因子:
--
作者:
Thomas G. Dietterich
通讯作者:
Thomas G. Dietterich