Natural gradient works efficiently in learning

Natural gradient works efficiently in learning
复制标题

DOI:
10.1162/089976698300017746
复制
发表时间:
1998-02-15
期刊:
影响因子:
2.9
通讯作者:
Amari, S
Amari, S
中科院分区:
计算机科学4区
文献类型:
--
作者:
Amari, S

文献摘要

被引文献

相似文献

当参数空间具有一定的基础结构时,函数的普通梯度并不代表其最陡峭的方向,而是自然梯度。信息几何形状用于计算感知器的参数空间中的自然梯度,矩阵的空间(用于盲源分离)和线性动力学系统的空间(用于盲源反卷积)。分析了自然梯度在线学习的动态行为,并被证明是有效的,这意味着它具有与最佳参数估计的渐近性能。这表明,在使用自然梯度时,出现在多层感知的返回传播学习算法中的高原现象可能会消失或可能不会那么严重。提出和分析了一种更新学习率的自适应方法。
When a parameter space has a certain underlying structure, the ordinary gradient of a function does not represent its steepest direction, but the natural gradient does. Information geometry is used for calculating the natural gradients in the parameter space of perceptrons, the space of matrices (for blind source separation), and the space of linear dynamical systems (for blind source deconvolution). The dynamical behavior of natural gradient online learning is analyzed and is proved to be Fisher efficient, implying that it has asymptotically the same performance as the optimal batch estimation of parameters. This suggests that the plateau phenomenon, which appears in the backpropagation learning algorithm of multilayer perceptrons, might disappear or might not be so serious when the natural gradient is used. An adaptive method of updating the learning rate is proposed and analyzed.