Acceleration of Reinforcement Learning by Controlled Use of Options Given as Prior Information

Acceleration of Reinforcement Learning by Controlled Use of Options Given as Prior Information
复制标题

通过控制使用作为先验信息给出的选项来加速强化学习

DOI:
10.9746/jcmsi.6.252
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
J. Murata
J. Murata
中科院分区:
--
文献类型:
--
作者:
Kento Terashima;H. Takano;J. Murata

文献摘要

被引文献

相似文献

强化学习是一种智能体通过试错法学习适当的行为策略来解决问题的方法。其优点是强化学习可以应用于未知或不确定的问题。但该方法存在一个缺点,即存在试错问题,求解时间较长。如果存在关于环境的先验信息,则可以省去一些试错,并且学习可以花费更短的时间。先验信息可以由人类设计者以选项的形式提供。但由于问题的不确定性,这些选择可能是错误的。如果使用错误的选项,可能会产生不良影响,例如无法获得最佳策略和减慢强化学习。本文提出控制期权的使用,以抑制不良影响。智能体在学习更好的策略时逐渐忘记给定的选项。所提出的方法应用于三种测试台环境和两种类型的先验信息。该方法在学习速度和获得的政策的质量方面表现出良好的效果。
Reinforcement learning is a method with which an agent learns an appropriate action policy for solving problems by the trial-and-error. The advantage is that reinforcement learning can be applied to unknown or uncertain problems. But instead, there is a drawback that this method needs a long time to solve the problem because of the trialand-error. If there is prior information about the environment, some of trial-and-error can be spared and the learning can take a shorter time. The prior information can be provided in the form of options by a human designer. But the options can be wrong because of uncertainties in the problems. If the wrong options are used, there can be bad effects such as failure to get the optimal policy and slowing down of reinforcement learning. This paper proposes to control use of the options to suppress the bad effects. The agent forgets the given options gradually while it learns the better policy. The proposed method is applied to three testbed environments and two types of prior information. The method shows good results in terms of both the learning speed and the quality of obtained policies.