Gradient Based Learning Methods
Gradient Based Learning Methods
复制标题
基于梯度的学习方法
DOI:
10.1007/bfb0053994
复制
发表时间:
1997
期刊:
影响因子:
--
通讯作者:
A. Tsoi
中科院分区:
文献类型:
--
作者:
A. Tsoi
In this paper, we will consider the issues of gradient based learning algorithms for a class of neural networks, viz., recurrent neural networks. This class of neural networks is useful in processing temporal information, or for modelling time series. In addition, this class of networks can be used for adaptive control of nonlinear plants. There have been a number of exposition on first order gradient based learning, see eg,[5].[5] approaches the problem by using an optimization based methods. It was shown that the adjoint variable can be used as a way of dealing with the resulting constrained optimization problem. In this paper, we will re-visit the first order gradient based methods, using a slightly different approach, viz., a variational based approach. While the adjoint approach (Pontryagin maximum principle) and the variational approach give the same results, it is felt that the variational approach is more intuitive, and it may be an easier vehicle for readers to appreciate some of the problems underlying learning algorithms for the class of recurrent neural networks. As we have observed elsewhere [26], there are two main subclasses of structured recurrent neural networks, viz., the class of fully connected hidden layer networks (alternatively known as the Elman network), and the class of dynamic multilayer perceptrons. We will derive the first order gradient based learning methods for each class respectively. It is well known that first order gradient learning methods are slow in convergence. In order to speed up the convergence, we will derive a class of second order gradient based methods, viz., extended Kalman filter approach, and recursive least squared approach, for these two main subclasses of architectures. Recently, there have been some studies in considering the output sensitivity of the network with respect to weight perturbations [4]. If one examines the approach indicated in [4] carefully, it is simple to observe that they are essentially based on the evaluation of first order gradient information of the error criterion, If a network suffers from high output sensitivity, it was shown in [4] that the sensitivity can be reduced using what is commonly known as alternative discrete time operators (ADTOs).[4] showed that such an approach can reduce the sensitivity of the subclass of networks, viz., the dynamic multilayer perceptrons. They used the concepts of" poles" and" zeros" in their discussion of the approach. It is however, not clear how such an approach can be applied to the situation of fully connected hidden layer recurrent neural network architectures as the