Acceleration of Reinforcement Learning by Policy Evaluation Using Nonstationary Iterative Method

Acceleration of Reinforcement Learning by Policy Evaluation Using Nonstationary Iterative Method
复制标题

DOI:
10.1109/tcyb.2014.2313655
复制
发表时间:
2014-04
影响因子:
11.8
通讯作者:
K. Senda;Suguru Hattori;T. Hishinuma;T. Kohda
K. Senda;Suguru Hattori;T. Hishinuma;T. Kohda
中科院分区:
计算机科学1区
文献类型:
--
作者:
K. Senda;Suguru Hattori;T. Hishinuma;T. Kohda

文献摘要

被引文献

相似文献

解决强化学习问题的典型方法迭代两个步骤,策略评估和策略改进。本文提出了策略评估算法以提高学习效率。所提出的算法基于 Krylov 子空间方法(KSM),这是一种非平稳迭代方法。基于KSM的算法比现有的基于平稳迭代方法的算法效率提高数十至数百倍。基于 KSM 的算法比人们普遍预期的效率要高得多。本文通过数值示例和理论讨论阐明了基于 KSM 的算法为何更加高效。
Typical methods for solving reinforcement learning problems iterate two steps, policy evaluation and policy improvement. This paper proposes algorithms for the policy evaluation to improve learning efficiency. The proposed algorithms are based on the Krylov Subspace Method (KSM), which is a nonstationary iterative method. The algorithms based on KSM are tens to hundreds times more efficient than existing algorithms based on the stationary iterative methods. Algorithms based on KSM are far more efficient than they have been generally expected. This paper clarifies what makes algorithms based on KSM makes more efficient with numerical examples and theoretical discussions.