Dual autoencoders features for imbalance classification problem

Dual autoencoders features for imbalance classification problem
复制标题

DOI:
10.1016/j.patcog.2016.06.013
复制
发表时间:
2016-12
期刊:
Pattern Recognit.
影响因子:
--
通讯作者:
Wing W. Y. Ng;Guangjun Zeng;Jianjun Zhang;D. Yeung;W. Pedrycz
Wing W. Y. Ng;Guangjun Zeng;Jianjun Zhang;D. Yeung;W. Pedrycz
中科院分区:
其他
文献类型:
--
作者:
Wing W. Y. Ng;Guangjun Zeng;Jianjun Zhang;D. Yeung;W. Pedrycz

文献摘要

被引文献

相似文献

在实际应用中遇到的许多分类问题表现出不平衡数据的轮廓。目前的方法依赖于数据恢复。事实上,如果特征集提供了清晰的决策边界,则可能不需要重新排序来解决不平衡分类问题。因此,本文提出了一种基于自动编码器的特征学习方法,通过学习一组具有更好分类能力的少数类和多数类特征,来解决分类不平衡的问题。两组特征由具有不同激活函数的两个堆叠的自动编码器学习,以捕获数据的不同特征,并且它们被组合以形成双重自动编码特征。然后在以这种方式学习的新特征空间而不是原始输入空间中对样本进行分类。实验结果表明,该方法优于现有的重采样方法的统计意义上的不平衡模式分类问题。
Many classification problems encountered in real-world applications exhibit a profile of imbalanced data. Current methods depend on data resampling. In fact, if the feature set provides a clear decision boundary, resampling may not be needed to solve the imbalanced classification problem. Therefore, this work proposes a feature learning method based on the autoencoder to learn a set of features with better classification capabilities of the minority and the majority classes to address the imbalanced classification problems. Two sets of features are learned by two stacked autoencoders with different activation functions to capture different characteristics of the data and they are combined to form the Dual Autoencoding Features. Samples are then classified in the new feature space learned in this manner instead of the original input space. Experimental results show that the proposed method outperforms current resampling-based methods with statistical significance for imbalanced pattern classification problems.