Learning from Noisy Labels with No Change to the Training Process

Learning from Noisy Labels with No Change to the Training Process
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Mingyuan Zhang;Jane Lee;S. Agarwal
Mingyuan Zhang;Jane Lee;S. Agarwal
中科院分区:
其他
文献类型:
--
作者:
Mingyuan Zhang;Jane Lee;S. Agarwal

文献摘要

相似文献

近年来,人们对开发能够从带有噪声标签的数据中学习准确分类器的学习算法非常感兴趣。一个被广泛研究的噪声模型是类别条件噪声(CCN)模型,其中一个标签y被翻转到一个标签(cid:101) y,并带有一些相关的噪声概率,该噪声概率取决于y和(cid:101) y。在多类设置中,CCN模型下的所有先前提出的算法都涉及改变训练过程,通过对代理损失引入“噪声校正”来最小化噪声训练示例。在本文中,我们表明这真的是不必要的:人们可以简单地对有噪声的例子进行类概率估计(CPE),例如使用标准(多类)逻辑回归算法,然后仅在最后的预测步骤中应用噪声校正。这意味着训练算法本身不需要任何更改,并且可以简单地使用标准的现成实现,而无需修改训练代码。我们的方法可以处理一般的多类损失矩阵,包括通常的0-1损失,也可以处理其他损失,例如用于有序回归问题的损失。我们还提供了一个定量的后悔转移界,它根据CPE后悔在噪声分布上的界限来限制目标后悔在真实分布上的界限;在此过程中,我们将Agarwal(2014)为二元损失引入的强适当性概念扩展到多类情况。我们的界表明,CCN下学习的样本复杂度随着噪声矩阵接近奇点而增加。我们还提供了涉及计算锚点的噪声估计方法的修复和潜在改进。我们的实验证实了我们的理论发现。
There has been much interest in recent years in developing learning algorithms that can learn accurate classifiers from data with noisy labels. A widely-studied noise model is that of class-conditional noise (CCN), wherein a label y is flipped to a label (cid:101) y with some associated noise probability that depends on both y and (cid:101) y . In the multiclass setting, all previously proposed algorithms under the CCN model involve changing the training process, by introducing a ‘noise-correction’ to the surrogate loss to be minimized over the noisy training examples. In this paper, we show that this is really unnecessary: one can simply perform class probability estimation (CPE) on the noisy examples, e.g. using a standard (multiclass) logistic regression algorithm, and then apply noise-correction only in the final prediction step. This means that the training algorithm itself does not need any change, and one can simply use standard off-the-shelf implementations with no modification to the code for training. Our approach can handle general multiclass loss matrices, including the usual 0-1 loss but also other losses such as those used for ordinal regression problems. We also provide a quantitative regret transfer bound, which bounds the target regret on the true distribution in terms of the CPE regret on the noisy distribution; in doing so, we extend the notion of strong properness introduced for binary losses by Agarwal (2014) to the multiclass case. Our bound suggests that the sample complexity of learning under CCN increases as the noise matrix approaches singularity. We also provide fixes and potential improvements for noise estimation meth-ods that involve computing anchor points. Our experiments confirm our theoretical findings.