LEARNING FINITE-STATE MACHINES WITH SELF-CLUSTERING RECURRENT NETWORKS

LEARNING FINITE-STATE MACHINES WITH SELF-CLUSTERING RECURRENT NETWORKS
复制标题

DOI:
10.1162/neco.1993.5.6.976
复制
发表时间:
1993-11-01
期刊:
影响因子:
2.9
通讯作者:
SMYTH, P
SMYTH, P
中科院分区:
计算机科学4区
文献类型:
--
作者:
ZENG, Z;GOODMAN, RM;SMYTH, P

文献摘要

被引文献

相似文献

最近的工作表明,循环神经网络具有从示例中学习有限状态自动机的能力。特别是,使用二阶单元的网络已经成功地完成了这项任务。在研究此类网络的性能和学习行为时,我们发现二阶网络模型试图在激活空间中形成簇作为其状态的内部表示。然而,随着越来越长的测试输入字符串呈现给网络,这些学习状态变得不稳定。本质上,网络“忘记”了各个状态在激活空间中的位置。在本文中,我们提出了一种新方法,通过在网络中引入离散化并使用伪梯度学习规则进行训练来迫使此类网络学习稳定状态。学习规则的本质是,在进行梯度下降时,它利用 sigmoid 函数的梯度作为启发式提示来代替硬限制函数的梯度,同时在反馈更新路径中仍然使用离散化值。新结构使用激活空间中的孤立点而不是模糊簇作为其状态的内部表示。它被证明具有与原始网络相似的学习有限状态自动机的能力,但不存在不稳定问题。所提出的伪梯度学习规则也可以用作训练具有硬限制阈值激活函数的其他类型网络的基础。
Recent work has shown that recurrent neural networks have the ability to learn finite state automata from examples. In particular, networks using second-order units have been successful at this task. In studying the performance and learning behavior of such networks we have found that the second-order network model attempts to form clusters in activation space as its internal representation of states. However, these learned states become unstable as longer and longer test input strings are presented to the network. In essence, the network ''forgets'' where the individual states are in activation space. In this paper we propose a new method to force such a network to learn stable states by introducing discretization into the network and using a pseudo-gradient learning rule to perform training. The essence of the learning rule is that in doing gradient descent, it makes use of the gradient of a sigmoid function as a heuristic hint in place of that of the hard-limiting function, while still using the discretized value in the feedback update path. The new structure uses isolated points in activation space instead of vague clusters as its internal representation of states. It is shown to have similar capabilities in learning finite state automata as the original network, but without the instability problem. The proposed pseudo-gradient learning rule may also be used as a basis for training other types of networks that have hard-limiting threshold activation functions.