Recurrent neural networks for reinforcement learning: architecture, learning algorithms and internal representation

Recurrent neural networks for reinforcement learning: architecture, learning algorithms and internal representation
复制标题

用于强化学习的循环神经网络:架构、学习算法和内部表示

DOI:
10.1109/ijcnn.1998.687168
复制
发表时间:
1998
期刊:
1998 IEEE International Joint Conference on Neural Networks Proceedings. IEEE World Congress on Computational Intelligence (Cat. No.98CH36227)
影响因子:
--
通讯作者:
Y. Nishikawa
Y. Nishikawa
中科院分区:
--
文献类型:
--
作者:
Ahmet Onat;Hajime Kita;Y. Nishikawa

文献摘要

被引文献

相似文献

强化学习是一种自主代理的学习方案,允许代理找到采取行动的最佳策略,从而在未知环境中最大化标量强化信号。如果代理可以访问环境的整个状态,则将感官输入映射到操作的反应策略就足够了。然而,如果环境状态是部分可观察的,则需要使用过去的观察结果来创建动态策略的特殊方法。为了克服这个问题,作者提出了一种使用带有 Q 学习的循环神经网络作为学习代理的方法。论文通过计算机模拟比较了该方法的几种网络架构和学习算法。此外,使用聚类技术检查训练网络中的内部表示。它表明环境状态的表征在网络中得到了很好的发展。
Reinforcement learning is a learning scheme for an autonomous agent that allows the agent to find the optimal policy of taking actions which maximize a scalar reinforcement signal in unknown environments. If the agent has access to the whole state of the environment, a reactive policy which maps the sensory input to the action is sufficient. However, if the state of the environment is partially observable, special methods for creating a dynamic policy that utilizes the past observations are necessary. To overcome this problem, the authors have proposed a method using recurrent neural networks with Q-learning, as a learning agent. The paper compares several types of network architecture and learning algorithms for this method through computer simulation. Further, the internal representation in the trained networks is examined using a clustering technique. It shows that the representation of the environmental state is developed well in the networks.