Gradient Based Learning Methods

Gradient Based Learning Methods
复制标题

基于梯度的学习方法

DOI:
10.1007/bfb0053994
复制
发表时间:
1997
期刊:
--
影响因子:
--
通讯作者:
A. Tsoi
A. Tsoi
中科院分区:
--
文献类型:
--
作者:
A. Tsoi

文献摘要

被引文献

相似文献

在本文中,我们将考虑一类神经网络(即循环神经网络)的基于梯度的学习算法问题。此类神经网络可用于处理时间信息或建模时间序列。此外,此类网络还可用于非线性设备的自适应控制。关于基于一阶梯度的学习已经有很多论述,参见例如,[5]。[5]通过使用基于优化的方法来解决问题。结果表明,伴随变量可以用作处理由此产生的约束优化问题的一种方法。在本文中,我们将重新审视基于一阶梯度的方法,使用稍微不同的方法,即基于变分的方法。虽然伴随方法(庞特里亚金极大值原理)和变分方法给出了相同的结果,但人们认为变分方法更直观,并且它可能是读者理解递归神经网络类学习算法背后的一些问题的更容易的工具。正如我们在其他地方所观察到的[26],结构化循环神经网络有两个主要子类,即全连接隐藏层网络类(也称为 Elman 网络)和动态多层感知器类。我们将分别为每个类别推导基于一阶梯度的学习方法。众所周知,一阶梯度学习方法收敛速度慢。为了加速收敛,我们将为这两个主要的体系结构子类派生一类基于二阶梯度的方法,即扩展卡尔曼滤波器方法和递归最小二乘方法。最近,有一些研究考虑网络的输出对权重扰动的敏感性[4]。如果仔细检查[4]中指出的方法,很容易观察到它们本质上是基于误差准则的一阶梯度信息的评估。如果网络具有高输出敏感性,[4]中表明可以使用通常称为替代离散时间算子(ADTO)的方法来降低敏感性。[4]表明这种方法可以降低网络子类(即动态多层感知器)的灵敏度。他们在讨论该方法时使用了“极点”和“零点”的概念。然而,目前尚不清楚这种方法如何应用于完全连接的隐藏层循环神经网络架构的情况
In this paper, we will consider the issues of gradient based learning algorithms for a class of neural networks, viz., recurrent neural networks. This class of neural networks is useful in processing temporal information, or for modelling time series. In addition, this class of networks can be used for adaptive control of nonlinear plants. There have been a number of exposition on first order gradient based learning, see eg,[5].[5] approaches the problem by using an optimization based methods. It was shown that the adjoint variable can be used as a way of dealing with the resulting constrained optimization problem. In this paper, we will re-visit the first order gradient based methods, using a slightly different approach, viz., a variational based approach. While the adjoint approach (Pontryagin maximum principle) and the variational approach give the same results, it is felt that the variational approach is more intuitive, and it may be an easier vehicle for readers to appreciate some of the problems underlying learning algorithms for the class of recurrent neural networks. As we have observed elsewhere [26], there are two main subclasses of structured recurrent neural networks, viz., the class of fully connected hidden layer networks (alternatively known as the Elman network), and the class of dynamic multilayer perceptrons. We will derive the first order gradient based learning methods for each class respectively. It is well known that first order gradient learning methods are slow in convergence. In order to speed up the convergence, we will derive a class of second order gradient based methods, viz., extended Kalman filter approach, and recursive least squared approach, for these two main subclasses of architectures. Recently, there have been some studies in considering the output sensitivity of the network with respect to weight perturbations [4]. If one examines the approach indicated in [4] carefully, it is simple to observe that they are essentially based on the evaluation of first order gradient information of the error criterion, If a network suffers from high output sensitivity, it was shown in [4] that the sensitivity can be reduced using what is commonly known as alternative discrete time operators (ADTOs).[4] showed that such an approach can reduce the sensitivity of the subclass of networks, viz., the dynamic multilayer perceptrons. They used the concepts of" poles" and" zeros" in their discussion of the approach. It is however, not clear how such an approach can be applied to the situation of fully connected hidden layer recurrent neural network architectures as the