Error-Bounded Correction of Noisy Labels

Error-Bounded Correction of Noisy Labels
复制标题

DOI:
--
复制
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Songzhu Zheng;Pengxiang Wu;A. Goswami;Mayank Goswami;Dimitris N. Metaxas;Chao Chen
Songzhu Zheng;Pengxiang Wu;A. Goswami;Mayank Goswami;Dimitris N. Metaxas;Chao Chen
中科院分区:
其他
文献类型:
--
作者:
Songzhu Zheng;Pengxiang Wu;A. Goswami;Mayank Goswami;Dimitris N. Metaxas;Chao Chen

文献摘要

被引文献

相似文献

为了收集大规模注释数据,不可避免地引入标签噪声,即,不正确的类标签。为了对标签噪声具有鲁棒性,许多成功的方法依赖于噪声分类器(即,在噪声训练数据上训练的模型)来确定标签是否可信。然而,它仍然不知道为什么这种启发式在实践中运作良好。在本文中,我们提供了这些方法的第一个理论解释。我们证明了噪声分类器的预测确实可以很好地指示训练数据的标签是否干净。基于理论结果,我们提出了一种新的算法,纠正标签的基础上嘈杂的分类器预测。修正后的标签与真实的贝叶斯最优分类器具有很高的概率一致。我们将我们的标签校正算法纳入深度神经网络的训练中,并训练在多个公共数据集上实现上级测试性能的模型。
To collect large scale annotated data, it is inevitable to introduce label noise, i.e., incorrect class labels. To be robust against label noise, many successful methods rely on the noisy classifiers (i.e., models trained on the noisy training data) to determine whether a label is trustworthy. However, it remains unknown why this heuristic works well in practice. In this paper, we provide the first theoretical explanation for these methods. We prove that the prediction of a noisy classifier can indeed be a good indicator of whether the label of a training data is clean. Based on the theoretical result, we propose a novel algorithm that corrects the labels based on the noisy classifier prediction. The corrected labels are consistent with the true Bayesian optimal classifier with high probability. We incorporate our label correction algorithm into the training of deep neural networks and train models that achieve superior testing performance on multiple public datasets.