Robust Learning with Noisy Label Detection and Counterfactual Correction

Robust Learning with Noisy Label Detection and Counterfactual Correction
复制标题

DOI:
10.1109/bigdata55660.2022.10020228
复制
发表时间:
2022-12
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Wenting Qi;C. Chelmis
Wenting Qi;C. Chelmis
中科院分区:
其他
文献类型:
--
作者:
Wenting Qi;C. Chelmis

文献摘要

相似文献

数据质量在任何机器学习模型的训练过程中都是至关重要的。最近提出的噪声学习方法主要是通过使用固定的损失值阈值来检测噪声标记的数据实例,并在随后的训练步骤中排除检测到的噪声数据实例。然而,预定义的固定损失值阈值可能并不总是最优的,并且排除检测到的噪声数据实例可能会损害训练集的大小。在本文中,我们提出了一种新的方法,NDCC,自动选择一个损失阈值来识别噪声标记的数据实例,并使用反事实学习来修复它们。据我们所知,NDCC是第一个探索在噪声学习领域使用反事实学习的可行性的工作。我们在各种标签噪声环境下演示了NDCC在Fashion-MNIST和CIFAR-10数据集上的性能。实验结果表明,与现有方法相比,该方法具有优越性,特别是在存在严重标签噪声的情况下。
Data quality is of paramount importance in the training process of any machine learning model. Recently proposed methods for noisy learning focus on detecting noisy labeled data instances by using a fixed loss value threshold, and exclude detected noisy data instances in subsequent training steps. However, a predefined, fixed loss value threshold may not always be optimal, and excluding the detected noisy data instances can hurt the size of the training set. In this paper, we propose a new method, NDCC, that automatically selects a loss threshold to identify noisy labeled data instance, and uses counterfactual learning to repair them. To the best of our knowledge, NDCC is the first work to explore the feasibility of using counterfactual learning in the noisy learning domain. We demonstrate the performance of NDCC on Fashion–MNIST and CIFAR–10 datasets under a variety of label noise environments. Experimental results show the superiority of the proposed method compared to the state–of–the–art, especially in the presence of severe label noise.