A Geometric Analysis of Neural Collapse with Unconstrained Features

A Geometric Analysis of Neural Collapse with Unconstrained Features
复制标题

DOI:
--
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
Zhihui Zhu;Tianyu Ding;Jinxin Zhou;Xiao Li;Chong You;Jeremias Sulam;Qing Qu
Zhihui Zhu;Tianyu Ding;Jinxin Zhou;Xiao Li;Chong You;Jeremias Sulam;Qing Qu
中科院分区:
其他
文献类型:
--
作者:
Zhihui Zhu;Tianyu Ding;Jinxin Zhou;Xiao Li;Chong You;Jeremias Sulam;Qing Qu

文献摘要

相似文献

我们提供了$Neural\;Collapse$的第一个全局优化景观分析-这是一种有趣的经验现象,在训练的最终阶段出现在神经网络的最后一层分类器和特征中。正如Papyan等人最近报道的那样,这种现象意味着($i$)类平均值和最后一层分类器都坍缩到单纯形等角紧框架(ETF)的顶点,直到缩放,并且($ii$)最后一层激活的跨示例类内可变性坍缩到零。我们研究的问题的基础上,简化的$无约束\;特征\;模型$,隔离的最上层的分类器的神经网络。在这种情况下,我们表明,经典的交叉熵损失与重量衰减有一个良性的全球景观,在这个意义上说,唯一的全球最小值是单纯形ETF,而所有其他的临界点是严格的鞍,其Hessian表现出负曲率方向。与深度神经网络的现有景观分析(通常与实践脱节)相比,我们对简化模型的分析不仅解释了在最后一层学习了什么样的特征,而且还表明了为什么它们可以在简化设置中有效优化,与实际深度网络架构中的经验观察相匹配。这些发现可能对广泛兴趣的优化,泛化和鲁棒性产生深远的影响。例如,我们的实验表明,可以将特征维度设置为等于类的数量,并将最后一层分类器固定为用于网络训练的Simplex ETF,这在ResNet 18上减少了超过20美元的内存成本,而不会牺牲泛化性能。
We provide the first global optimization landscape analysis of $Neural\;Collapse$ -- an intriguing empirical phenomenon that arises in the last-layer classifiers and features of neural networks during the terminal phase of training. As recently reported by Papyan et al., this phenomenon implies that ($i$) the class means and the last-layer classifiers all collapse to the vertices of a Simplex Equiangular Tight Frame (ETF) up to scaling, and ($ii$) cross-example within-class variability of last-layer activations collapses to zero. We study the problem based on a simplified $unconstrained\;feature\;model$, which isolates the topmost layers from the classifier of the neural network. In this context, we show that the classical cross-entropy loss with weight decay has a benign global landscape, in the sense that the only global minimizers are the Simplex ETFs while all other critical points are strict saddles whose Hessian exhibit negative curvature directions. In contrast to existing landscape analysis for deep neural networks which is often disconnected from practice, our analysis of the simplified model not only does it explain what kind of features are learned in the last layer, but it also shows why they can be efficiently optimized in the simplified settings, matching the empirical observations in practical deep network architectures. These findings could have profound implications for optimization, generalization, and robustness of broad interests. For example, our experiments demonstrate that one may set the feature dimension equal to the number of classes and fix the last-layer classifier to be a Simplex ETF for network training, which reduces memory cost by over $20\%$ on ResNet18 without sacrificing the generalization performance.