Dynamics of Learning in MLP: Natural Gradient and Singularity Revisited

Dynamics of Learning in MLP: Natural Gradient and Singularity Revisited
复制标题

MLP 中的学习动态:重新审视自然梯度和奇点

DOI:
10.1162/neco_a_01029
复制
发表时间:
2018
期刊:
影响因子:
2.9
通讯作者:
Okada Masato
Okada Masato
中科院分区:
计算机科学4区
文献类型:
--
作者:
Amari Shun-ichi;Ozeki Tomoko;Karakida Ryo;Yoshida Yuki;Okada Masato

文献摘要

参考文献

被引文献

相似文献

监督学习的动态性在深度学习中起着主要作用,它发生在多层感知器(MLP)的参数空间中。我们回顾了监督随机梯度学习的历史,重点是它的奇异结构和自然梯度。参数空间包括参数不可识别的奇异区域。我们的结果之一是在一个基本的奇异网络的随机梯度学习的动力学行为的充分探索。坏消息是它的病态性质,其中奇异区域的一部分成为吸引子,另一部分同时成为排斥子,形成米尔诺吸引子。学习轨迹被吸引子区域吸引,在吸引子区域中停留很长时间,然后通过排斥子区域逃离奇异区域。这是典型的学习高原现象。通过引入可用于分析自然梯度动力学的吹落坐标,我们证明了奇异区域的奇异拓扑。我们确认,自然梯度动力学是免费的临界减速。第二个主要结果是好消息:基本奇异网络的相互作用消除了吸引子部分,Milnor-type吸引子消失了。这就解释了为什么大规模网络不会因为奇点而遭受严重的临界减速。我们最后表明,单位明智的自然梯度是有效的学习,尽管它的计算成本低。
The dynamics of supervised learning play a main role in deep learning, which takes place in the parameter space of a multilayer perceptron (MLP). We review the history of supervised stochastic gradient learning, focusing on its singular structure and natural gradient. The parameter space includes singular regions in which parameters are not identifiable. One of our results is a full exploration of the dynamical behaviors of stochastic gradient learning in an elementary singular network. The bad news is its pathological nature, in which part of the singular region becomes an attractor and another part a repulser at the same time, forming a Milnor attractor. A learning trajectory is attracted by the attractor region, staying in it for a long time, before it escapes the singular region through the repulser region. This is typical of plateau phenomena in learning. We demonstrate the strange topology of a singular region by introducing blow-down coordinates, which are useful for analyzing the natural gradient dynamics. We confirm that the natural gradient dynamics are free of critical slowdown. The second main result is the good news: the interactions of elementary singular networks eliminate the attractor part and the Milnor-type attractors disappear. This explains why large-scale networks do not suffer from serious critical slowdowns due to singularities. We finally show that the unit-wise natural gradient is effective for learning in spite of its low computational cost.
DOI: 10.1093/imaiai/iav006
发表时间: 2013-03
期刊: arXiv: Neural and Evolutionary Computing
影响因子: --
作者:
Y. Ollivier
通讯作者: Y. Ollivier
DOI: --
发表时间: 2012
期刊:
影响因子: --
作者:
J. J. Burns;N. Cantor
通讯作者: N. Cantor
神经网络的黎曼度量 II:循环网络和学习符号数据序列
DOI: --
发表时间: 2013
期刊:
影响因子: --
作者:
Y. Ollivier
通讯作者: Y. Ollivier
DOI: 10.4135/9781412983907.n140
发表时间: 2020
期刊: Definitions
影响因子: --
作者:
Duane C. McBride;John M. Berecz;Sharon A. Gillespie;Herbert W. Helm;James H. Hopkins;Øystein S. LaBianca;Lionel N. A. Matthews;Susan E. Murray;Derrick L. Proctor;Larry S. Ulery
通讯作者: Larry S. Ulery
DOI: 10.1073/pnas.1603583113
发表时间: 2016-12-20
影响因子: 11.1
作者:
Oizumi, Masafumi;Tsuchiya, Naotsugu;Amari, Shun-ichi
通讯作者: Amari, Shun-ichi