Augmentation Strategies for Learning with Noisy Labels

Augmentation Strategies for Learning with Noisy Labels
复制标题

DOI:
10.1109/cvpr46437.2021.00793
复制
发表时间:
2021-03
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Kento Nishi;Yi Ding;Alex Rich;Tobias Höllerer
Kento Nishi;Yi Ding;Alex Rich;Tobias Höllerer
中科院分区:
其他
文献类型:
--
作者:
Kento Nishi;Yi Ding;Alex Rich;Tobias Höllerer

文献摘要

被引文献

相似文献

不完美的标签在现实世界的数据集中无处不在。最近几种成功的训练对标签噪声鲁棒的深度神经网络(DNN)的方法使用了两种主要技术:在预热阶段基于损失过滤样本,以管理初始的干净标记样本集,并使用网络的输出作为后续损失计算的伪标签。在本文中,我们评估了不同的增强策略的算法解决“带噪声标签的学习”的问题。我们提出并研究了多种增强策略,并使用基于CIFAR-10和CIFAR-100的合成数据集以及真实世界数据集Clothing 1 M对其进行评估。由于这些算法中的几个共性,我们发现使用一组增强用于损失建模任务,另一组用于学习是最有效的,改进了最先进的和其他以前的方法的结果。此外,我们发现在预热期间应用增强会对正确标记样本与错误标记样本的损失收敛行为产生负面影响。我们将这种增强策略引入到最先进的技术中,并证明我们可以在所有评估的噪声水平上提高性能。特别是,我们在90%对称噪声下将CIFAR-10基准测试的准确性提高了15%以上,并且我们还提高了Clothing 1 M数据集的性能。
Imperfect labels are ubiquitous in real-world datasets. Several recent successful methods for training deep neural networks (DNNs) robust to label noise have used two primary techniques: filtering samples based on loss during a warm-up phase to curate an initial set of cleanly labeled samples, and using the output of a network as a pseudo-label for subsequent loss calculations. In this paper, we evaluate different augmentation strategies for algorithms tackling the "learning with noisy labels" problem. We propose and examine multiple augmentation strategies and evaluate them using synthetic datasets based on CIFAR-10 and CIFAR-100, as well as on the real-world dataset Clothing1M. Due to several commonalities in these algorithms, we find that using one set of augmentations for loss modeling tasks and another set for learning is the most effective, improving results on the state-of-the-art and other previous methods. Furthermore, we find that applying augmentation during the warm-up period can negatively impact the loss convergence behavior of correctly versus incorrectly labeled samples. We introduce this augmentation strategy to the state-of-the-art technique and demonstrate that we can improve performance across all evaluated noise levels. In particular, we improve accuracy on the CIFAR-10 benchmark at 90% symmetric noise by more than 15% in absolute accuracy, and we also improve performance on the Clothing1M dataset.