Neural Collapse with Normalized Features: A Geometric Analysis over the Riemannian Manifold

Neural Collapse with Normalized Features: A Geometric Analysis over the Riemannian Manifold
复制标题

DOI:
10.48550/arxiv.2209.09211
复制
发表时间:
2022-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Can Yaras;Peng Wang;Zhihui Zhu;L. Balzano;Qing Qu
Can Yaras;Peng Wang;Zhihui Zhu;L. Balzano;Qing Qu
中科院分区:
其他
文献类型:
--
作者:
Can Yaras;Peng Wang;Zhihui Zhu;L. Balzano;Qing Qu

文献摘要

相似文献

当训练过度参数化的深度网络进行分类任务时,人们广泛观察到学习到的特征表现出所谓的“神经崩溃”现象。更具体地说,对于倒数第二层的输出特征,对于每个类,类内特征收敛到它们的均值,并且不同类的均值表现出某种紧密的框架结构,该结构也与最后一层的分类器对齐。由于最后一层的特征归一化成为现代表示学习中的常见做法,因此在这项工作中,我们从理论上证明了归一化特征的神经崩溃现象。基于一个无约束的特征模型,我们简化了经验损失函数的多类分类任务的黎曼流形上的非凸优化问题,约束所有的功能和分类球。在这种情况下,我们分析的非凸景观的黎曼优化问题的产品领域,显示一个良性的全球景观的意义上说,唯一的全球最小值是神经崩溃的解决方案,而所有其他的临界点是严格的鞍负曲率。在实际深度网络上的实验结果证实了我们的理论,并证明通过特征归一化可以更快地学习更好的表示。
When training overparameterized deep networks for classification tasks, it has been widely observed that the learned features exhibit a so-called"neural collapse"phenomenon. More specifically, for the output features of the penultimate layer, for each class the within-class features converge to their means, and the means of different classes exhibit a certain tight frame structure, which is also aligned with the last layer's classifier. As feature normalization in the last layer becomes a common practice in modern representation learning, in this work we theoretically justify the neural collapse phenomenon for normalized features. Based on an unconstrained feature model, we simplify the empirical loss function in a multi-class classification task into a nonconvex optimization problem over the Riemannian manifold by constraining all features and classifiers over the sphere. In this context, we analyze the nonconvex landscape of the Riemannian optimization problem over the product of spheres, showing a benign global landscape in the sense that the only global minimizers are the neural collapse solutions while all other critical points are strict saddles with negative curvature. Experimental results on practical deep networks corroborate our theory and demonstrate that better representations can be learned faster via feature normalization.