Image classification with deep learning in the presence of noisy labels: A survey

Image classification with deep learning in the presence of noisy labels: A survey
复制标题

DOI:
10.1016/j.knosys.2021.106771
复制
发表时间:
2021-01-18
影响因子:
8.8
通讯作者:
Ulusoy, Ilkay
Ulusoy, Ilkay
中科院分区:
计算机科学1区
文献类型:
--
作者:
Algan, Gorkem;Ulusoy, Ilkay

文献摘要

被引文献

相似文献

随着深度神经网络的发展,图像分类系统最近取得了巨大的飞跃。然而,这些系统需要大量的标记数据来进行充分的训练。收集一个正确注释的数据集并不总是可行的,这是由于几个因素,例如标签过程的昂贵或正确分类数据的困难,即使是专家。由于这些实际挑战,标签噪声是现实世界数据集中的一个常见问题,文献中提出了许多使用标签噪声训练深度神经网络的方法。尽管深度神经网络对标签噪声相对稳健,但它们过度拟合数据的倾向使它们很容易记住随机噪声。因此,考虑标签噪声的存在并开发对抗算法来消除其不利影响以有效地训练深度神经网络至关重要。尽管对标签噪声下的机器学习技术进行了广泛的调查,但文献缺乏对存在噪声标签的深度学习方法的全面调查。本文的目的是提出这些算法,同时将它们分为两个小组之一:基于噪声模型和噪声模型的方法。第一组算法旨在估计噪声结构,并使用此信息来避免噪声标签的不利影响。因此,第二组中的方法试图通过使用鲁棒损失,正则化器或其他学习范式等方法来提出固有的噪声鲁棒算法。(C)2021爱思唯尔有限公司版权所有。
Image classification systems recently made a giant leap with the advancement of deep neural networks. However, these systems require an excessive amount of labeled data to be adequately trained. Gathering a correctly annotated dataset is not always feasible due to several factors, such as the expensiveness of the labeling process or difficulty of correctly classifying data, even for the experts. Because of these practical challenges, label noise is a common problem in real-world datasets, and numerous methods to train deep neural networks with label noise are proposed in the literature. Although deep neural networks are known to be relatively robust to label noise, their tendency to overfit data makes them vulnerable to memorizing even random noise. Therefore, it is crucial to consider the existence of label noise and develop counter algorithms to fade away its adverse effects to train deep neural networks efficiently. Even though an extensive survey of machine learning techniques under label noise exists, the literature lacks a comprehensive survey of methodologies centered explicitly around deep learning in the presence of noisy labels. This paper aims to present these algorithms while categorizing them into one of the two subgroups: noise model based and noise model free methods. Algorithms in the first group aim to estimate the noise structure and use this information to avoid the adverse effects of noisy labels. Differently, methods in the second group try to come up with inherently noise robust algorithms by using approaches like robust losses, regularizers or other learning paradigms. (C) 2021 Elsevier B.V. All rights reserved.