Iterative ensemble feature selection for multiclass classification of imbalanced microarray data.

Iterative ensemble feature selection for multiclass classification of imbalanced microarray data.
复制标题

不平衡微阵列数据多类分类的迭代集成特征选择

DOI:
10.1186/s40709-016-0045-8
复制
发表时间:
2016-05
期刊:
Journal of biological research (Thessalonike, Greece)
影响因子:
--
通讯作者:
Ji Z
Ji Z
中科院分区:
其他
文献类型:
--
作者:
Yang J;Zhou J;Zhu Z;Ma X;Ji Z

文献摘要

相似文献

背景微阵列技术使生物学家能够监测各种肿瘤组织中数千个基因的表达水平。鉴定相关基因用于不同肿瘤类型的样本分类,有利于临床研究。多类分类数据最广泛使用的分类策略之一是One-Versus-All(OVA)模式,该模式将原始问题划分为一个类对其余类的多个二进制分类。然而,多类微阵列数据往往遭受不平衡的类分布之间的多数和少数类,这不可避免地恶化的OVA classification.ResultsIn这项研究的性能,我们提出了一种新的迭代集成特征选择(IEFS)框架不平衡的微阵列数据的多类分类。特别是,过滤器特征选择和平衡采样迭代地和交替地执行,以提高在OVA模式中的每个二进制分类的性能。该框架进行了测试,并与其他代表性的国家的最先进的过滤器的特征选择方法,使用六个基准多类微阵列数据集进行比较。实验结果表明,IEFS框架提供了上级或可比的性能,无论是分类精度和接收器工作特征曲线下的面积的其他方法。类的数据的数量越多,IEFS框架实现更好的性能。ConclusionsBalanced sampling和特征选择一起工作,以及在提高多类分类的性能不平衡的微阵列数据。IEFS框架很容易适用于面临相同问题的其他生物数据分析任务。
BackgroundMicroarray technology allows biologists to monitor expression levels of thousands of genes among various tumor tissues. Identifying relevant genes for sample classification of various tumor types is beneficial to clinical studies. One of the most widely used classification strategies for multiclass classification data is the One-Versus-All (OVA) schema that divides the original problem into multiple binary classification of one class against the rest. Nevertheless, multiclass microarray data tend to suffer from imbalanced class distribution between majority and minority classes, which inevitably deteriorates the performance of the OVA classification.ResultsIn this study, we propose a novel iterative ensemble feature selection (IEFS) framework for multiclass classification of imbalanced microarray data. In particular, filter feature selection and balanced sampling are performed iteratively and alternatively to boost the performance of each binary classification in the OVA schema. The proposed framework is tested and compared with other representative state-of-the-art filter feature selection methods using six benchmark multiclass microarray data sets. The experimental results show that IEFS framework provides superior or comparable performance to the other methods in terms of both classification accuracy and area under receiver operating characteristic curve. The more number of classes the data have, the better performance of IEFS framework achieves.ConclusionsBalanced sampling and feature selection together work well in improving the performance of multiclass classification of imbalanced microarray data. The IEFS framework is readily applicable to other biological data analysis tasks facing the same problem.