Feature selection and classification of MAQC-II breast cancer and multiple myeloma microarray gene expression data.

Feature selection and classification of MAQC-II breast cancer and multiple myeloma microarray gene expression data.
复制标题

DOI:
10.1371/journal.pone.0008250
复制
发表时间:
2009-12-11
期刊:
影响因子:
3.7
通讯作者:
Deng Y
Deng Y
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Liu Q;Sung AH;Chen Z;Liu J;Huang X;Deng Y

文献摘要

参考文献

被引文献

相似文献

微阵列数据具有高维度的变量,但可用的数据集通常只有少量的样本,从而使得对此类数据集的研究有趣且具有挑战性。在分析微阵列数据的任务中,例如,在预测基因与疾病的关联性时,特征选择非常重要,因为它提供了一种通过利用遗传标记之间的关联引起的信息冗余来处理高维度的方法。在微阵列数据分析中,明智的特征选择可以显著降低成本,同时保持或提高用于整理数据集的学习机的分类或预测精度。在本文中,我们提出了一种基因选择方法称为递归特征添加(RFA),它结合了监督学习和统计相似性度量。我们将我们的方法与以下基因选择方法进行比较:支持向量机递归特征消除(SVMRFE)留一法计算顺序前向选择(LOOCSFS)基于梯度的留一法基因选择(GLGS)为了评估这些基因选择方法的性能,我们在微阵列质量控制第二阶段的预测建模(MAQC-II)中使用了几种流行的学习分类器乳腺癌数据集和MAQC-II多发性骨髓瘤数据集。实验结果表明,基因选择与学习分类器严格配对。总的来说,我们的方法优于其他比较方法。基于MAQC-II乳腺癌数据集的生物学功能分析使我们确信将我们的方法应用于表型预测。此外,学习分类器在微阵列数据的分类中也起着重要的作用,我们的实验结果表明,最近均值尺度分类器(NMSC)是一个很好的选择,因为它的预测可靠性和稳定性跨越三个性能指标:测试准确度,MCC值和AUC误差。
Microarray data has a high dimension of variables but available datasets usually have only a small number of samples, thereby making the study of such datasets interesting and challenging. In the task of analyzing microarray data for the purpose of, e.g., predicting gene-disease association, feature selection is very important because it provides a way to handle the high dimensionality by exploiting information redundancy induced by associations among genetic markers. Judicious feature selection in microarray data analysis can result in significant reduction of cost while maintaining or improving the classification or prediction accuracy of learning machines that are employed to sort out the datasets. In this paper, we propose a gene selection method called Recursive Feature Addition (RFA), which combines supervised learning and statistical similarity measures. We compare our method with the following gene selection methods: Support Vector Machine Recursive Feature Elimination (SVMRFE) Leave-One-Out Calculation Sequential Forward Selection (LOOCSFS) Gradient based Leave-one-out Gene Selection (GLGS) To evaluate the performance of these gene selection methods, we employ several popular learning classifiers on the MicroArray Quality Control phase II on predictive modeling (MAQC-II) breast cancer dataset and the MAQC-II multiple myeloma dataset. Experimental results show that gene selection is strictly paired with learning classifier. Overall, our approach outperforms other compared methods. The biological functional analysis based on the MAQC-II breast cancer dataset convinced us to apply our method for phenotype prediction. Additionally, learning classifiers also play important roles in the classification of microarray data and our experimental results indicate that the Nearest Mean Scale Classifier (NMSC) is a good choice due to its prediction reliability and its stability across the three performance measurements: Testing accuracy, MCC values, and AUC errors.
DOI: 10.1504/ijbra.2005.008443
发表时间: 2005-01-01
影响因子: --
作者:
Liang, Yulan;Kelemen, Arpad
通讯作者: Kelemen, Arpad
DOI: 10.1074/jbc.m010192200
发表时间: 2001-06-08
影响因子: 4.8
作者:
Long, AD;Mangalam, HJ;Baldi, P
通讯作者: Baldi, P
DOI: 10.1089/106652701300099074
发表时间: 2001-01-01
影响因子: 1.7
作者:
Newton, MA;Kendziorski, CM;Tsui, KW
通讯作者: Tsui, KW
DOI: 10.1016/0005-2795(75)90109-9
发表时间: 1975-01-01
期刊: BIOCHIMICA ET BIOPHYSICA ACTA
影响因子: --
作者:
MATTHEWS, BW
通讯作者: MATTHEWS, BW
DOI: 10.1186/gb-2001-2-10-research0042
发表时间: 2001
期刊: Genome biology
影响因子: 12.3
作者:
Pavlidis P;Noble WS
通讯作者: Noble WS