Adaptive natural gradient learning algorithms for unnormalized statistical models

Adaptive natural gradient learning algorithms for unnormalized statistical models
复制标题

非归一化统计模型的自适应自然梯度学习算法

DOI:
10.1007/978-3-319-44778-0_50
复制
发表时间:
2016
期刊:
Lecture Notes in Computer Science (ICANN2016)
影响因子:
--
通讯作者:
Shun-ichi Amari
Shun-ichi Amari
中科院分区:
--
文献类型:
--
作者:
Ryo Karakida;Masato Okada;Shun-ichi Amari

文献摘要

相似文献

自然梯度是一种利用参数空间的几何结构来改善学习瞬态动力学的有效方法。许多自然梯度方法已经被开发用于最大似然学习,其基于Kullback-Leibler(KL)散度及其Fisher度量。然而,它们需要计算归一化常数,并且不适用于具有分析上难以处理的归一化常数的统计模型。在这项研究中,我们将自然梯度框架扩展到非归一化统计模型的分歧:分数匹配和比率匹配。此外,我们推导出新的自适应自然梯度算法,不需要计算要求反演的度量,并在一些数值实验中显示其有效性。特别是,在一个多层神经网络模型的实验结果表明,该方法可以摆脱高原现象比传统的随机梯度下降方法快得多。
The natural gradient is a powerful method to improve the transient dynamics of learning by utilizing the geometric structure of the parameter space. Many natural gradient methods have been developed for maximum likelihood learning, which is based on Kullback-Leibler (KL) divergence and its Fisher metric. However, they require the computation of the normalization constant and are not applicable to statistical models with an analytically intractable normalization constant. In this study, we extend the natural gradient framework to divergences for the unnormalized statistical models: score matching and ratio matching. In addition, we derive novel adaptive natural gradient algorithms that do not require computationally demanding inversion of the metric and show their effectiveness in some numerical experiments. In particular, experimental results in a multi-layer neural network model demonstrate that the proposed method can escape from the plateau phenomena much faster than the conventional stochastic gradient descent method.