MSE-Optimal Neural Network Initialization via Layer Fusion

MSE-Optimal Neural Network Initialization via Layer Fusion
复制标题

DOI:
10.1109/ciss48834.2020.1570617381
复制
发表时间:
2020-01
期刊:
2020 54th Annual Conference on Information Sciences and Systems (CISS)
影响因子:
--
通讯作者:
Ramina Ghods;Andrew S. Lan;T. Goldstein;Christoph Studer
Ramina Ghods;Andrew S. Lan;T. Goldstein;Christoph Studer
中科院分区:
其他
文献类型:
--
作者:
Ramina Ghods;Andrew S. Lan;T. Goldstein;Christoph Studer

文献摘要

相似文献

深度神经网络在一系列分类和推理任务中实现了最先进的性能。然而,使用随机梯度下降结合非凸性的基本优化问题,使参数学习容易初始化。为了解决这个问题,过去已经提出了依赖于随机参数初始化或知识蒸馏的各种方法。在本文中,我们提出了FuseInit,这是一种通过融合随机初始化训练的深层网络的相邻层来初始化浅层网络的新方法。我们开发的理论结果和有效的算法的均方误差(MSE)的最佳融合相邻的密集,卷积密集,卷积卷积层。我们展示了一系列分类和回归数据集的实验,这些实验表明,如果使用FuseInit初始化,较深的神经网络对初始化不太敏感,较浅的网络可以表现得更好(有时与较深的网络一样)。
Deep neural networks achieve state-of-the-art performance for a range of classification and inference tasks. However, the use of stochastic gradient descent combined with the nonconvexity of the underlying optimization problems renders parameter learning susceptible to initialization. To address this issue, a variety of methods that rely on random parameter initialization or knowledge distillation have been proposed in the past. In this paper, we propose FuseInit, a novel method to initialize shallower networks by fusing neighboring layers of deeper networks that are trained with random initialization. We develop theoretical results and efficient algorithms for mean-square error (MSE)- optimal fusion of neighboring dense-dense, convolutional-dense, and convolutional-convolutional layers. We show experiments for a range of classification and regression datasets, which suggest that deeper neural networks are less sensitive to initialization and shallower networks can perform better (sometimes as well as their deeper counterparts) if initialized with FuseInit.