On the Optimization Landscape of Neural Collapse under MSE Loss: Global Optimality with Unconstrained Features

On the Optimization Landscape of Neural Collapse under MSE Loss: Global Optimality with Unconstrained Features
复制标题

DOI:
10.48550/arxiv.2203.01238
复制
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Jinxin Zhou-;Xiao Li;Tian Ding;Chong You;Qing Qu;Zhihui Zhu
Jinxin Zhou-;Xiao Li;Tian Ding;Chong You;Qing Qu;Zhihui Zhu
中科院分区:
其他
文献类型:
--
作者:
Jinxin Zhou-;Xiao Li;Tian Ding;Chong You;Qing Qu;Zhihui Zhu

文献摘要

被引文献

相似文献

在训练用于分类任务的深度神经网络时,在最后一层分类器和特征中广泛观察到一个有趣的经验现象,其中(i)类均值和最后一层分类器都塌陷到单纯形等角紧框架(ETF)的顶点,直到缩放,以及(ii)最后一层激活的跨示例类内变异性塌陷到零。这种现象被称为神经崩溃(NC),它似乎发生在损失函数的选择无关。在这项工作中,我们证明NC下的均方误差(MSE)的损失,最近的经验证据表明,它的性能比事实上的交叉熵损失,甚至更好。在一个简化的无约束特征模型,我们提供了第一个全球景观分析香草非凸MSE损失,并表明(只有!)全局极小点是神经崩溃解,而所有其他临界点都是严格鞍点,其Hessian具有负曲率方向。此外,我们通过探测NC解决方案周围的优化景观来证明重新缩放MSE损失的使用,表明可以通过调整重新缩放超参数来改善景观。最后,我们的理论发现在实际网络架构上得到了实验验证。
When training deep neural networks for classification tasks, an intriguing empirical phenomenon has been widely observed in the last-layer classifiers and features, where (i) the class means and the last-layer classifiers all collapse to the vertices of a Simplex Equiangular Tight Frame (ETF) up to scaling, and (ii) cross-example within-class variability of last-layer activations collapses to zero. This phenomenon is called Neural Collapse (NC), which seems to take place regardless of the choice of loss functions. In this work, we justify NC under the mean squared error (MSE) loss, where recent empirical evidence shows that it performs comparably or even better than the de-facto cross-entropy loss. Under a simplified unconstrained feature model, we provide the first global landscape analysis for vanilla nonconvex MSE loss and show that the (only!) global minimizers are neural collapse solutions, while all other critical points are strict saddles whose Hessian exhibit negative curvature directions. Furthermore, we justify the usage of rescaled MSE loss by probing the optimization landscape around the NC solutions, showing that the landscape can be improved by tuning the rescaling hyperparameters. Finally, our theoretical findings are experimentally verified on practical network architectures.