Learning Action-Value Functions Using Neural Networks with Incremental Learning Ability

Learning Action-Value Functions Using Neural Networks with Incremental Learning Ability
复制标题

使用具有增量学习能力的神经网络学习行动价值函数

DOI:
--
复制
发表时间:
2001
期刊:
--
影响因子:
--
通讯作者:
N. Shiraga
N. Shiraga
中科院分区:
--
文献类型:
--
作者:
N. Shiraga

文献摘要

被引文献

相似文献

当给定的训练数据的分布是有偏的和随时间变化的,这是众所周知的,神经网络的学习变得困难,一般。在强化学习(RL)问题中,经常会出现这种情况。在本文中,一个增量学习系统,这已经被设计用于监督学习,被实现为一个RL代理,即使在上述困难的情况下,也可以正确地获得一个动作值函数。建议的RL代理应用于一个扩展的山地车任务,其中学习域在时间上扩展。通过计算机模拟,我们证明了所提出的代理可以获得一个正确的政策,在这个任务。
When the distribution of given training data is biased and temporally varied, it is well known that the learning of neural networks becomes difficult in general. In Reinforcement Learning (RL) problems, such situations often arise. In this paper, an incremental learning system, which has been devised for supervised learning, is implemented as an RL agent that can acquire an action-value function properly even in the above difficult situations. The proposed RL agent is applied to an extended mountain-car task in which learning domains are temporally expanded. Through computer simulations, we demonstrate that the proposed agent can acquire a right policy in this task.