Noisy Label Detection and Counterfactual Correction

Noisy Label Detection and Counterfactual Correction
复制标题

DOI:
10.1109/tai.2023.3271963
复制
发表时间:
2024-02
期刊:
IEEE Transactions on Artificial Intelligence
影响因子:
--
通讯作者:
Wenting Qi;C. Chelmis
Wenting Qi;C. Chelmis
中科院分区:
其他
文献类型:
--
作者:
Wenting Qi;C. Chelmis

文献摘要

相似文献

数据质量对于任何机器学习模型的训练都是至关重要的。最近提出的噪声学习方法侧重于通过使用固定的丢失值阈值来检测噪声标记的数据实例,并在后续的训练步骤中排除检测到的噪声数据实例。然而,预定义的固定损失值阈值对于检测噪声标签数据可能不是最佳的,而排除检测到的噪声数据实例可以将训练集的大小减小到可以负面影响精度的程度。在本文中,我们提出了一种新的方法--噪声标签检测和反事实校正(NDCC),该方法自动选择一个损失阈值来识别噪声标签数据实例,并使用反事实学习来校正噪声标签。据我们所知,NDCC是第一个探索反事实学习在噪声学习领域中使用的工作。在不同的标签噪声环境下,我们展示了NDCC在几个数据集上的性能。实验结果表明,与现有技术相比,特别是在存在严重的标签噪声的情况下,该方法具有更好的性能。
Data quality is of paramount importance to the training of any machine learning model. Recently proposed approaches for noisy learning focus on detecting noisy labeled data instances by using a fixed loss value threshold and excluding detected noisy data instances in subsequent training steps. However, a predefined, fixed loss value threshold may not be optimal for detecting noisy labeled data, whereas excluding the detected noisy data instances can reduce the size of the training set to such an extent that accuracy can be negatively affected. In this article, we propose Noisy label Detection and Counterfactual Correction (NDCC), a new approach that automatically selects a loss value threshold to identify noisy labeled data instances, and uses counterfactual learning to correct the noisy labels. To the best of our knowledge, NDCC is the first work to explore the use of counterfactual learning in the noisy learning domain. We demonstrate the performance of NDCC on several datasets under a variety of label noise environments. Experimental results show the superiority of the proposed approach compared to the state of the art, especially in the presence of severe label noise.