Acceleration of Reinforcement Learning by Controlled Use of Options Given as Prior Information
Acceleration of Reinforcement Learning by Controlled Use of Options Given as Prior Information
复制标题
通过控制使用作为先验信息给出的选项来加速强化学习
DOI:
10.9746/jcmsi.6.252
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
J. Murata
中科院分区:
文献类型:
--
作者:
Kento Terashima;H. Takano;J. Murata
Reinforcement learning is a method with which an agent learns an appropriate action policy for solving problems by the trial-and-error. The advantage is that reinforcement learning can be applied to unknown or uncertain problems. But instead, there is a drawback that this method needs a long time to solve the problem because of the trialand-error. If there is prior information about the environment, some of trial-and-error can be spared and the learning can take a shorter time. The prior information can be provided in the form of options by a human designer. But the options can be wrong because of uncertainties in the problems. If the wrong options are used, there can be bad effects such as failure to get the optimal policy and slowing down of reinforcement learning. This paper proposes to control use of the options to suppress the bad effects. The agent forgets the given options gradually while it learns the better policy. The proposed method is applied to three testbed environments and two types of prior information. The method shows good results in terms of both the learning speed and the quality of obtained policies.