Class prediction for high-dimensional class-imbalanced data.

Class prediction for high-dimensional class-imbalanced data.
复制标题

DOI:
10.1186/1471-2105-11-523
复制
发表时间:
2010-10-20
期刊:
影响因子:
3
通讯作者:
Lusa L
Lusa L
中科院分区:
生物学4区
文献类型:
--
作者:
Blagus R;Lusa L

文献摘要

参考文献

被引文献

相似文献

类别预测研究的目标是开发规则来准确预测新样本的类别成员。这些规则是使用每个主题可用的变量值导出的:高维数据的主要特征是变量的数量大大超过样本的数量。分类器通常使用类别不平衡的数据开发,即,数据集,其中每个类中的样本数量不相等。标准的分类方法用于类不平衡的数据往往产生分类器,不能准确地预测少数类;预测偏向于多数类。在本文中,我们研究了在处理类不平衡预测时,高维性是否会带来额外的挑战。我们评估了六种类型的分类器的类不平衡数据的性能,使用模拟数据和公开的数据集,从乳腺癌基因表达微阵列研究。我们还调查了一些策略,可用于克服类不平衡的影响的有效性。我们的研究结果表明,评估的分类器是高度敏感的类不平衡和变量选择引入了一个额外的偏见对大多数类的分类。大多数新样本被分配到训练集中的多数类,除非类之间的差异非常大。因此,特定类别的预测精度差异很大。当类别不平衡不太严重时,缩小和非对称装袋嵌入变量选择效果较好,而过采样则不好。变量归一化会进一步恶化分类器的性能。我们的研究结果表明,在训练集和测试集中匹配类的流行并不能保证分类器的良好性能,并且在处理高维数据时,与类不平衡数据的分类相关的问题会加剧。使用类不平衡数据的研究人员应该小心评估分类器的预测准确性,除非类不平衡是轻微的,否则他们应该始终使用适当的方法来处理类不平衡问题。
The goal of class prediction studies is to develop rules to accurately predict the class membership of new samples. The rules are derived using the values of the variables available for each subject: the main characteristic of high-dimensional data is that the number of variables greatly exceeds the number of samples. Frequently the classifiers are developed using class-imbalanced data, i.e., data sets where the number of samples in each class is not equal. Standard classification methods used on class-imbalanced data often produce classifiers that do not accurately predict the minority class; the prediction is biased towards the majority class. In this paper we investigate if the high-dimensionality poses additional challenges when dealing with class-imbalanced prediction. We evaluate the performance of six types of classifiers on class-imbalanced data, using simulated data and a publicly available data set from a breast cancer gene-expression microarray study. We also investigate the effectiveness of some strategies that are available to overcome the effect of class imbalance. Our results show that the evaluated classifiers are highly sensitive to class imbalance and that variable selection introduces an additional bias towards classification into the majority class. Most new samples are assigned to the majority class from the training set, unless the difference between the classes is very large. As a consequence, the class-specific predictive accuracies differ considerably. When the class imbalance is not too severe, down-sizing and asymmetric bagging embedding variable selection work well, while over-sampling does not. Variable normalization can further worsen the performance of the classifiers. Our results show that matching the prevalence of the classes in training and test set does not guarantee good performance of classifiers and that the problems related to classification with class-imbalanced data are exacerbated when dealing with high-dimensional data. Researchers using class-imbalanced data should be careful in assessing the predictive accuracy of the classifiers and, unless the class imbalance is mild, they should always use an appropriate method for dealing with the class imbalance problem.
DOI: 10.1038/nature03702
发表时间: 2005-06-09
期刊: NATURE
影响因子: 64.8
作者:
Lu, J;Getz, G;Golub, TR
通讯作者: Golub, TR
DOI: 10.1038/ng1547
发表时间: 2005-05-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Gunderson, KL;Steemers, FJ;Chee, MS
通讯作者: Chee, MS
DOI: 10.1073/pnas.97.1.262
发表时间: 2000-01-04
影响因子: 11.1
作者:
Brown, MPS;Grundy, WN;Haussler, D
通讯作者: Haussler, D
DOI: 10.1016/s0140-6736(05)17866-0
发表时间: 2005-02-05
期刊: LANCET
影响因子: 168.9
作者:
Michiels, S;Koscielny, S;Hill, C
通讯作者: Hill, C
DOI: 10.1109/tsmcb.2008.2007853
发表时间: 2009-04-01
影响因子: --
作者:
Liu, Xu-Ying;Wu, Jianxin;Zhou, Zhi-Hua
通讯作者: Zhou, Zhi-Hua