Are All Losses Created Equal: A Neural Collapse Perspective

Are All Losses Created Equal: A Neural Collapse Perspective
复制标题

DOI:
10.48550/arxiv.2210.02192
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Jinxin Zhou-;Chong You;Xiao Li;Kangning Liu;Sheng Liu;Qing Qu;Zhihui Zhu
Jinxin Zhou-;Chong You;Xiao Li;Kangning Liu;Sheng Liu;Qing Qu;Zhihui Zhu
中科院分区:
其他
文献类型:
--
作者:
Jinxin Zhou-;Chong You;Xiao Li;Kangning Liu;Sheng Liu;Qing Qu;Zhihui Zhu

文献摘要

被引文献

相似文献

虽然交叉熵(CE)是训练用于分类任务的深度神经网络最常用的损失,但已经开发了许多替代损失以获得更好的经验性能。其中,哪一个是最好的使用仍然是一个谜,因为似乎有多种因素影响答案,如数据集的属性,网络架构的选择等等,本文通过考察深度网络的最后一层特征来研究损失函数的选择,从最近的一项工作中得到启发,该工作表明CE和均方误差(MSE)损失的全局最优解表现出神经崩溃现象。也就是说,对于训练到收敛的足够大的网络,(i)同一类的所有特征都坍缩到相应的类均值,以及(ii)与不同类相关联的均值处于它们的成对距离都相等且最大化的配置中。我们扩展了这些结果,并通过全局解和景观分析表明,一个广泛的损失函数家族,包括常用的标签平滑(LS)和焦点损失(FL)表现出神经崩溃。因此,所有相关损失(即,CE、LS、FL、MSE)在训练数据上产生等效特征。基于无约束特征模型假设,我们提供了LS损失的全局景观分析或FL损失的局部景观分析,并表明(唯一!)全局最小值是神经崩溃解,而所有其他临界点是严格鞍,其Hessian在LS损失的全局范围内或在最优解附近的FL损失的局部范围内表现出负曲率方向。实验进一步表明,从所有相关损失中获得的神经崩溃特征在测试数据上也会导致基本相同的性能,前提是网络足够大并经过训练直到收敛。
While cross entropy (CE) is the most commonly used loss to train deep neural networks for classification tasks, many alternative losses have been developed to obtain better empirical performance. Among them, which one is the best to use is still a mystery, because there seem to be multiple factors affecting the answer, such as properties of the dataset, the choice of network architecture, and so on. This paper studies the choice of loss function by examining the last-layer features of deep networks, drawing inspiration from a recent line work showing that the global optimal solution of CE and mean-square-error (MSE) losses exhibits a Neural Collapse phenomenon. That is, for sufficiently large networks trained until convergence, (i) all features of the same class collapse to the corresponding class mean and (ii) the means associated with different classes are in a configuration where their pairwise distances are all equal and maximized. We extend such results and show through global solution and landscape analyses that a broad family of loss functions including commonly used label smoothing (LS) and focal loss (FL) exhibits Neural Collapse. Hence, all relevant losses(i.e., CE, LS, FL, MSE) produce equivalent features on training data. Based on the unconstrained feature model assumption, we provide either the global landscape analysis for LS loss or the local landscape analysis for FL loss and show that the (only!) global minimizers are neural collapse solutions, while all other critical points are strict saddles whose Hessian exhibit negative curvature directions either in the global scope for LS loss or in the local scope for FL loss near the optimal solution. The experiments further show that Neural Collapse features obtained from all relevant losses lead to largely identical performance on test data as well, provided that the network is sufficiently large and trained until convergence.