Analysis of a Method Improving Reinforcement Learning Agents' Policies
Analysis of a Method Improving Reinforcement Learning Agents' Policies
复制标题
一种改进强化学习代理策略的方法分析
DOI:
10.20965/jaciii.2003.p0276
复制
发表时间:
2003
期刊:
影响因子:
--
通讯作者:
M. Kurihara
中科院分区:
文献类型:
--
作者:
D. Kitakoshi;H. Shioya;M. Kurihara
Reinforcement learning (RL) is a kind of machine learning. It aims to optimize agents’ policies by adapting the agents to an environment according to rewards. In this paper, we propose a method for improving policies by using stochastic knowledge, in which reinforcement learning agents obtain. We use a Bayesian Network (BN), which is a stochastic model, as knowledge of an agent. Its structure is decided by minimum description length criterion using series of an agent's input-output and rewards as sample data. A BN constructed in our study represents stochastic dependences between input-output and rewards. In our proposed method, policies are improved by supervised learning using the structure of BN (i.e. stochastic knowledge). The proposed improvement mechanism makes RL agents acquire more effective policies. We carry out simulations in the pursuit problem in order to show the effectiveness of our proposed method.