SMOTE for high-dimensional class-imbalanced data.

SMOTE for high-dimensional class-imbalanced data.
复制标题

DOI:
10.1186/1471-2105-14-106
复制
发表时间:
2013-03-22
期刊:
影响因子:
3
通讯作者:
Lusa L
Lusa L
中科院分区:
生物学4区
文献类型:
--
作者:
Blagus R;Lusa L

文献摘要

参考文献

被引文献

相似文献

使用类别不平衡数据进行的分类偏向于多数类别。对于高维数据,偏差甚至更大,其中变量的数量大大超过样本的数量。这个问题可以通过欠采样或过采样来缓解,这会产生类平衡的数据。一般来说,欠采样是有帮助的,而随机过采样则没有。合成少数过采样技术(SMOTE)是一种非常流行的过采样方法,被提出来改进随机过采样,但其在高维数据上的行为尚未得到彻底研究。在本文中,我们使用模拟和真实的高维数据,从理论和实证的角度研究了 SMOTE 的特性。虽然在大多数情况下,SMOTE 似乎对低维数据有益,但当数据是高维时,它不会减弱大多数分类器在多数类中分类的偏差,并且它的效果不如随机欠采样。如果执行某种类型的变量选择时变量数量减少,SMOTE 对于高维数据的 k-NN 分类器是有益的;我们解释为什么 k-NN 分类偏向于少数类。此外,我们表明,在高维数据上,SMOTE 不会改变特定于类的平均值,同时会降低数据变异性并引入样本之间的相关性。我们解释了我们的发现如何影响高维数据的类别预测。实际上,在高维设置中,只有基于欧几里德距离的 k-NN 分类器似乎能从 SMOTE 的使用中受益匪浅,前提是在使用 SMOTE 之前执行变量选择;如果使用更多的邻居,好处就更大。不应该使用没有变量选择的 k-NN 的 SMOTE,因为它使分类强烈偏向于少数类。
Classification using class-imbalanced data is biased in favor of the majority class. The bias is even larger for high-dimensional data, where the number of variables greatly exceeds the number of samples. The problem can be attenuated by undersampling or oversampling, which produce class-balanced data. Generally undersampling is helpful, while random oversampling is not. Synthetic Minority Oversampling TEchnique (SMOTE) is a very popular oversampling method that was proposed to improve random oversampling but its behavior on high-dimensional data has not been thoroughly investigated. In this paper we investigate the properties of SMOTE from a theoretical and empirical point of view, using simulated and real high-dimensional data. While in most cases SMOTE seems beneficial with low-dimensional data, it does not attenuate the bias towards the classification in the majority class for most classifiers when data are high-dimensional, and it is less effective than random undersampling. SMOTE is beneficial for k-NN classifiers for high-dimensional data if the number of variables is reduced performing some type of variable selection; we explain why, otherwise, the k-NN classification is biased towards the minority class. Furthermore, we show that on high-dimensional data SMOTE does not change the class-specific mean values while it decreases the data variability and it introduces correlation between samples. We explain how our findings impact the class-prediction for high-dimensional data. In practice, in the high-dimensional setting only k-NN classifiers based on the Euclidean distance seem to benefit substantially from the use of SMOTE, provided that variable selection is performed before using SMOTE; the benefit is larger if more neighbors are used. SMOTE for k-NN without variable selection should not be used, because it strongly biases the classification towards the minority class.
DOI: 10.1016/j.csl.2005.06.002
发表时间: 2006-10-01
影响因子: 4.3
作者:
Liu, Yang;Chawla, Nitesh V.;Stolcke, Andreas
通讯作者: Stolcke, Andreas
DOI: 10.1186/1471-2105-11-523
发表时间: 2010-10-20
期刊: BMC bioinformatics
影响因子: 3
作者:
Blagus R;Lusa L
通讯作者: Lusa L
DOI: 10.1093/biostatistics/kxj035
发表时间: 2007-01-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Guo, Yaqian;Hastie, Trevor;Tibshirani, Robert
通讯作者: Tibshirani, Robert
DOI: 10.1093/bioinformatics/bti815
发表时间: 2006-02-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
MacIsaac, KD;Gordon, DB;Fraenkel, E
通讯作者: Fraenkel, E
DOI: 10.1186/1471-2105-12-424
发表时间: 2011-10-28
期刊: BMC bioinformatics
影响因子: 3
作者:
Doyle S;Monaco J;Feldman M;Tomaszewski J;Madabhushi A
通讯作者: Madabhushi A