Perspective: new insights from loss function landscapes of neural networks

Perspective: new insights from loss function landscapes of neural networks
复制标题

观点:神经网络损失函数景观的新见解

DOI:
10.1088/2632-2153/ab7aef
复制
发表时间:
2020
期刊:
Science and Technology
影响因子:
--
通讯作者:
Chitturi S
Chitturi S
中科院分区:
--
文献类型:
--
作者:
Chitturi S

文献摘要

参考文献

被引文献

相似文献

我们使用为能源景观探索开发的各种技术,研究了在数据集标签错误、训练集多样性增加和节点连通性降低的情况下,神经网络的损失函数景观的结构。基准模型是原子几何优化和手写数字预测的分类问题。我们考虑了改变用于生成初始几何的原子组态空间的大小的影响,发现驻点的数量随着训练组态空间的大小而迅速增加。我们引入了节点局部性的度量来限制网络连通性和扰动权重对称性,并研究了该参数如何影响结果景观。我们发现,高度精简的系统具有低容量,并且呈现出极小值很少的景观。另一方面,少量减少的连接可以增强网络的表现力,并可能产生更复杂的情况。研究了刻意分类错误对训练数据的影响,我们发现,在最小值样本上计算的AUC测试中的方差随着训练错误的增加而显著增加,这为在噪声下进行训练时方差-偏差权衡的作用提供了新的见解。最后,我们说明了具有两个和三个隐含层的网络的局部极小值的数目是如何随着层数的增加和训练数据的减少而显著增加的。这项工作有助于进一步阐明神经网络的损失情况,并为未来的神经网络训练和优化工作提供指导。
We investigate the structure of the loss function landscape for neural networks subject to dataset mislabelling, increased training set diversity, and reduced node connectivity, using various techniques developed for energy landscape exploration. The benchmarking models are classification problems for atomic geometry optimisation and hand-written digit prediction. We consider the effect of varying the size of the atomic configuration space used to generate initial geometries and find that the number of stationary points increases rapidly with the size of the training configuration space. We introduce a measure of node locality to limit network connectivity and perturb permutational weight symmetry, and examine how this parameter affects the resulting landscapes. We find that highly-reduced systems have low capacity and exhibit landscapes with very few minima. On the other hand, small amounts of reduced connectivity can enhance network expressibility and can yield more complex landscapes. Investigating the effect of deliberate classification errors in the training data, we find that the variance in testing AUC, computed over a sample of minima, grows significantly with the training error, providing new insight into the role of the variance-bias trade-off when training under noise. Finally, we illustrate how the number of local minima for networks with two and three hidden layers, but a comparable number of variable edge weights, increases significantly with the number of layers, and as the number of training data decreases. This work helps shed further light on neural network loss landscapes and provides guidance for future work on neural network training and optimisation.
DOI: 10.1098/rspa.1925.0047
发表时间: 1925
期刊: Proceedings of The Royal Society A: Mathematical, Physical and Engineering Sciences
影响因子: --
作者:
Janet E. Jones;A. E. Ingham
通讯作者: A. E. Ingham
原型能源景观:动态诊断。
DOI: 10.1063/1.1829633
发表时间: 2005
期刊: The Journal of chemical physics
影响因子: --
作者:
F. Despa;D. Wales;R. Berry
通讯作者: R. Berry
DOI: --
发表时间: 1997
期刊:
影响因子: --
作者:
J. Doye;D. Wales
通讯作者: D. Wales
活化配合物的对称性
DOI: --
发表时间: 1968
期刊:
影响因子: --
作者:
J. Murrell;K. Laidler
通讯作者: K. Laidler
DOI: 10.1016/0166-1280(88)80133-7
发表时间: 1988-10-01
影响因子: --
作者:
LI, ZQ;SCHERAGA, HA
通讯作者: SCHERAGA, HA