Ensemble Usage for More Reliable Policy Identification in Reinforcement Learning

Ensemble Usage for More Reliable Policy Identification in Reinforcement Learning
复制标题

使用集成在强化学习中实现更可靠的策略识别

DOI:
--
复制
发表时间:
2011
期刊:
The European Symposium on Artificial Neural Networks
影响因子:
--
通讯作者:
S. Udluft
S. Udluft
中科院分区:
--
文献类型:
--
作者:
A. Hans;S. Udluft

文献摘要

参考文献

被引文献

相似文献

采用神经网络等强大函数逼近器的强化学习(RL)方法已成为一种有趣的最优控制方法。由于它们从观察中学习策略,因此当系统没有可用的分析描述时它们也适用。尽管已经报告了令人印象深刻的结果,但它们在实践中的处理仍然很困难,因为它们可能无法可靠地确定良好的政策。在之前的工作中,我们使用神经拟合 Q 迭代 (NFQ) 独立运行的策略集合来更可靠地生成成功的策略。在本文中,我们在更多问题上评估了该方法,并建议通过单个 NFQ 运行的连续迭代形成集成,作为完全独立运行的计算成本低廉的替代方案。
Reinforcement learning (RL) methods employing powerful function approximators like neural networks have become an interesting approach for optimal control. Since they learn a policy from observations, they are also applicable when no analytical description of the system is available. Although impressive results have been reported, their handling in practice is still hard, as they can fail at reliably determining a good policy. In previous work, we used ensembles of policies from independent runs of neural fitted Q-iteration (NFQ) to produce successful policies more reliably. In this paper we evaluate the approach on more problems and propose to form ensembles from successive iterations of a single NFQ run as a computationally cheap alternative to completely independent runs.
DOI: 10.1007/3-540-45014-9
发表时间: 2000-06
期刊: --
影响因子: --
作者:
Thomas G. Dietterich
通讯作者: Thomas G. Dietterich