Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels

Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels
复制标题

DOI:
--
复制
发表时间:
2018-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhilu Zhang;M. Sabuncu
Zhilu Zhang;M. Sabuncu
中科院分区:
其他
文献类型:
--
作者:
Zhilu Zhang;M. Sabuncu

文献摘要

被引文献

相似文献

深度神经网络(DNN)在许多学科的各种应用中取得了巨大的成功。然而,它们的上级性能伴随着需要正确注释的大规模数据集的昂贵成本。此外,由于DNN的丰富容量,训练标签中的错误可能会影响性能。为了解决这个问题,最近提出了平均绝对误差(MAE)作为常用的分类交叉熵(CCE)损失的噪声鲁棒替代方案。然而,正如我们在本文中所展示的那样,MAE在DNN和具有挑战性的数据集上表现不佳。在这里,我们提出了一个理论上接地的噪声鲁棒损失函数,可以看作是一个泛化的MAE和CCE。提出的损失函数可以很容易地应用于任何现有的DNN架构和算法,同时在各种噪声标签场景中产生良好的性能。我们报告了CIFAR-10,CIFAR-100和FASHION-MNIST数据集和合成生成的噪声标签进行的实验结果。
Deep neural networks (DNNs) have achieved tremendous success in a variety of applications across many disciplines. Yet, their superior performance comes with the expensive cost of requiring correctly annotated large-scale datasets. Moreover, due to DNNs' rich capacity, errors in training labels can hamper performance. To combat this problem, mean absolute error (MAE) has recently been proposed as a noise-robust alternative to the commonly-used categorical cross entropy (CCE) loss. However, as we show in this paper, MAE can perform poorly with DNNs and challenging datasets. Here, we present a theoretically grounded set of noise-robust loss functions that can be seen as a generalization of MAE and CCE. Proposed loss functions can be readily applied with any existing DNN architecture and algorithm, while yielding good performance in a wide range of noisy label scenarios. We report results from experiments conducted with CIFAR-10, CIFAR-100 and FASHION-MNIST datasets and synthetically generated noisy labels.