The Goldilocks zone: Towards better understanding of neural network loss landscapes

The Goldilocks zone: Towards better understanding of neural network loss landscapes
复制标题

金发姑娘区:更好地理解神经网络损失景观

DOI:
--
复制
发表时间:
2018
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
Adam Scherlis
Adam Scherlis
中科院分区:
--
文献类型:
--
作者:
Stanislav Fort;Adam Scherlis

文献摘要

被引文献

相似文献

我们使用随机的低维超平面和超球体来探索全连接和卷积神经网络的损失情况。通过计算这些超曲面上损失函数的Hessian H,我们观察到:1)H的正特征值的数目异常地多; 2)Tr(H)/||H||在一个明确定义的配置空间半径范围内,对应于一个厚的,中空的,球壳,我们称之为金发区。我们在具有ReLU和tanh非线性的MNIST和CIFAR-10数据集上观察到了一系列网络宽度和深度的全连接神经网络的这种效应,以及卷积网络的类似效应。使用我们的观察,我们证明了一个密切的联系,金发区,措施的局部凸性/患病率的正曲率,和网络初始化的适用性。我们表明,高和稳定的精度达到随机优化时,低维超曲面直接相关的超曲面和Goldilocks区之间的重叠,并作为一个推论表明,内在尺寸的概念是初始化依赖。我们注意到,常见的初始化技术在这个异常高凸性/正曲率流行的特定区域初始化神经网络,并为它们的成功提供几何直观。此外,我们证明了在多个点初始化神经网络并选择局部凸性的高测度,如Tr(H)/||H||,H的正特征值的数量,或低初始损失,导致统计上显著更快的MNIST训练。根据我们的观察,我们假设,适居带包含一个异常高密度的合适的初始化配置。
We explore the loss landscape of fully-connected and convolutional neural networks using random, low-dimensional hyperplanes and hyperspheres. Evaluating the Hessian, H, of the loss function on these hypersurfaces, we observe 1) an unusual excess of the number of positive eigenvalues of H, and 2) a large value of Tr(H)/||H|| at a well defined range of configuration space radii, corresponding to a thick, hollow, spherical shell we refer to as the Goldilocks zone. We observe this effect for fully-connected neural networks over a range of network widths and depths on MNIST and CIFAR-10 datasets with the ReLU and tanh non-linearities, and a similar effect for convolutional networks. Using our observations, we demonstrate a close connection between the Goldilocks zone, measures of local convexity/prevalence of positive curvature, and the suitability of a network initialization. We show that the high and stable accuracy reached when optimizing on random, low-dimensional hypersurfaces is directly related to the overlap between the hypersurface and the Goldilocks zone, and as a corollary demonstrate that the notion of intrinsic dimension is initialization-dependent. We note that common initialization techniques initialize neural networks in this particular region of unusually high convexity/prevalence of positive curvature, and offer a geometric intuition for their success. Furthermore, we demonstrate that initializing a neural network at a number of points and selecting for high measures of local convexity such as Tr(H)/||H||, number of positive eigenvalues of H, or low initial loss, leads to statistically significantly faster training on MNIST. Based on our observations, we hypothesize that the Goldilocks zone contains an unusually high density of suitable initialization configurations.