Survey on deep learning with class imbalance

Survey on deep learning with class imbalance
复制标题

DOI:
10.1186/s40537-019-0192-5
复制
发表时间:
2019-03-19
影响因子:
8.1
通讯作者:
Khoshgoftaar, Taghi M.
Khoshgoftaar, Taghi M.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Johnson, Justin M.;Khoshgoftaar, Taghi M.

文献摘要

被引文献

相似文献

本研究的目的是检查现有的深度学习技术,以解决类别不平衡数据。对不平衡数据的有效分类是一个重要的研究领域,因为高类别不平衡在许多现实世界的应用中是自然固有的,例如,欺诈检测和癌症检测。此外,高度不平衡的数据带来了额外的困难,因为大多数学习者会表现出对多数阶层的偏见,在极端情况下,可能会完全忽略少数阶层。在过去的二十年里,使用传统的机器学习模型(即非深度学习)对类不平衡进行了深入的研究。尽管最近在深度学习方面取得了进展,沿着其越来越受欢迎,但在具有类不平衡的深度学习领域中的实证工作很少。在几个复杂领域取得了破纪录的性能结果之后,研究深度神经网络在包含高级别类不平衡问题中的应用具有极大的意义。调查了关于类不平衡和深度学习的现有研究,以便更好地了解深度学习在应用于类不平衡数据时的有效性。本调查讨论了每项研究的实施细节和实验结果,并提供了对其优势和劣势的更多见解。重点关注的几个领域包括:数据复杂性,测试的架构,性能解释,易用性,大数据应用程序和其他领域的推广。我们发现,这一领域的研究非常有限,大多数现有工作都集中在卷积神经网络的计算机视觉任务上,很少考虑大数据的影响。几种传统的类不平衡方法,例如数据采样和成本敏感学习,被证明适用于深度学习,而利用神经网络特征学习能力的更先进的方法显示出有希望的结果。调查的最后一个讨论强调了从类不平衡数据中进行深度学习的各种差距,以指导未来的研究。
The purpose of this study is to examine existing deep learning techniques for addressing class imbalanced data. Effective classification with imbalanced data is an important area of research, as high class imbalance is naturally inherent in many real-world applications, e.g., fraud detection and cancer detection. Moreover, highly imbalanced data poses added difficulty, as most learners will exhibit bias towards the majority class, and in extreme cases, may ignore the minority class altogether. Class imbalance has been studied thoroughly over the last two decades using traditional machine learning models, i.e. non-deep learning. Despite recent advances in deep learning, along with its increasing popularity, very little empirical work in the area of deep learning with class imbalance exists. Having achieved record-breaking performance results in several complex domains, investigating the use of deep neural networks for problems containing high levels of class imbalance is of great interest. Available studies regarding class imbalance and deep learning are surveyed in order to better understand the efficacy of deep learning when applied to class imbalanced data. This survey discusses the implementation details and experimental results for each study, and offers additional insight into their strengths and weaknesses. Several areas of focus include: data complexity, architectures tested, performance interpretation, ease of use, big data application, and generalization to other domains. We have found that research in this area is very limited, that most existing work focuses on computer vision tasks with convolutional neural networks, and that the effects of big data are rarely considered. Several traditional methods for class imbalance, e.g. data sampling and cost-sensitive learning, prove to be applicable in deep learning, while more advanced methods that exploit neural network feature learning abilities show promising results. The survey concludes with a discussion that highlights various gaps in deep learning from class imbalanced data for the purpose of guiding future research.