Effective data generation for imbalanced learning using conditional generative adversarial networks

Effective data generation for imbalanced learning using conditional generative adversarial networks
复制标题

DOI:
10.1016/j.eswa.2017.09.030
复制
发表时间:
2018-01-01
影响因子:
8.5
通讯作者:
Bacao, Fernando
Bacao, Fernando
中科院分区:
计算机科学1区
文献类型:
--
作者:
Douzas, Georgios;Bacao, Fernando

文献摘要

被引文献

相似文献

对于标准分类算法来说,从不平衡数据集中学习是一项常见但具有挑战性的任务。尽管有不同的策略来解决这个问题,但与算法修改相比,为少数类生成人工数据的方法构成了更通用的方法。标准过采样方法是 SMOTE 算法的变体,它沿着连接少数类样本的线段生成合成样本。因此,这些方法是基于局部信息,而不是整体少数群体分布。与这些算法相反,本文使用生成对抗网络(cGAN)的条件版本来近似真实的数据分布,并为各种不平衡数据集的少数类生成数据。 cGAN 的性能与多种标准过采样算法进行了比较。我们提出的实证结果表明,当使用 cGAN 作为过采样算法时,生成数据的质量显着提高。 (C) 2017 Elsevier Ltd. 保留所有权利。
Learning from imbalanced datasets is a frequent but challenging task for standard classification algorithms. Although there are different strategies to address this problem, methods that generate artificial data for the minority class constitute a more general approach compared to algorithmic modifications. Standard oversampling methods are variations of the SMOTE algorithm, which generates synthetic samples along the line segment that joins minority class samples. Therefore, these approaches are based on local information, rather on the overall minority class distribution. Contrary to these algorithms, in this paper the conditional version of Generative Adversarial Networks (cGAN) is used to approximate the true data distribution and generate data for the minority class of various imbalanced datasets. The performance of cGAN is compared against multiple standard oversampling algorithms. We present empirical results that show a significant improvement in the quality of the generated data when cGAN is used as an oversampling algorithm. (C) 2017 Elsevier Ltd. All rights reserved.