Learning in Gated Neural Networks

Learning in Gated Neural Networks
复制标题

DOI:
--
复制
发表时间:
2019-06
期刊:
--
影响因子:
--
通讯作者:
Ashok Vardhan Makkuva;Sewoong Oh;Sreeram Kannan;P. Viswanath
Ashok Vardhan Makkuva;Sewoong Oh;Sreeram Kannan;P. Viswanath
中科院分区:
其他
文献类型:
--
作者:
Ashok Vardhan Makkuva;Sewoong Oh;Sreeram Kannan;P. Viswanath

文献摘要

相似文献

选通是现代神经网络的一个重要特征,包括LSTM、GRUS和稀疏选通的深度神经网络。这种门控网络的主干是一个专家混合层,在那里几个专家做出回归决策,门控控制如何以依赖于输入的方式权衡决策。尽管在现代和经典机器学习中都有如此重要的作用,但人们对混合专家的参数恢复了解很少,因为众所周知,梯度下降和EM算法在这类模型中陷入局部最优。在本文中,我们对优化景观进行了仔细的分析,并证明了在适当设计损失函数的情况下,梯度下降确实可以准确地学习参数。支持我们结果的一个关键思想是设计了两个不同的损失函数,一个用于恢复专家参数,另一个用于恢复门控参数。对于任何算法,我们都证明了该模型中参数恢复的第一个样本复杂性结果,并在数值实验中证明了相对于标准损失函数的显著性能改进。
Gating is a key feature in modern neural networks including LSTMs, GRUs and sparsely-gated deep neural networks. The backbone of such gated networks is a mixture-of-experts layer, where several experts make regression decisions and gating controls how to weigh the decisions in an input-dependent manner. Despite having such a prominent role in both modern and classical machine learning, very little is understood about parameter recovery of mixture-of-experts since gradient descent and EM algorithms are known to be stuck in local optima in such models. In this paper, we perform a careful analysis of the optimization landscape and show that with appropriately designed loss functions, gradient descent can indeed learn the parameters accurately. A key idea underpinning our results is the design of two {\em distinct} loss functions, one for recovering the expert parameters and another for recovering the gating parameters. We demonstrate the first sample complexity results for parameter recovery in this model for any algorithm and demonstrate significant performance gains over standard loss functions in numerical experiments.