Analysis of a Method Improving Reinforcement Learning Agents' Policies

Analysis of a Method Improving Reinforcement Learning Agents' Policies
复制标题

一种改进强化学习代理策略的方法分析

DOI:
10.20965/jaciii.2003.p0276
复制
发表时间:
2003
期刊:
J. Adv. Comput. Intell. Intell. Informatics
影响因子:
--
通讯作者:
M. Kurihara
M. Kurihara
中科院分区:
--
文献类型:
--
作者:
D. Kitakoshi;H. Shioya;M. Kurihara

文献摘要

被引文献

相似文献

强化学习(RL)是机器学习的一种。它的目的是通过根据奖励使智能体适应环境来优化智能体的策略。在本文中,我们提出了一种利用强化学习代理获得的随机知识来改进策略的方法。我们使用贝叶斯网络(BN),它是一种随机模型,作为代理的知识。其结构由最小描述长度准则决定,使用一系列代理的输入输出和奖励作为样本数据。我们研究中构建的 BN 代表了输入输出和奖励之间的随机依赖性。在我们提出的方法中,通过使用 BN 结构(即随机知识)的监督学习来改进策略。所提出的改进机制使强化学习代理获得更有效的策略。我们对追踪问题进行了模拟,以证明我们提出的方法的有效性。
Reinforcement learning (RL) is a kind of machine learning. It aims to optimize agents’ policies by adapting the agents to an environment according to rewards. In this paper, we propose a method for improving policies by using stochastic knowledge, in which reinforcement learning agents obtain. We use a Bayesian Network (BN), which is a stochastic model, as knowledge of an agent. Its structure is decided by minimum description length criterion using series of an agent's input-output and rewards as sample data. A BN constructed in our study represents stochastic dependences between input-output and rewards. In our proposed method, policies are improved by supervised learning using the structure of BN (i.e. stochastic knowledge). The proposed improvement mechanism makes RL agents acquire more effective policies. We carry out simulations in the pursuit problem in order to show the effectiveness of our proposed method.