A classification-based machine learning approach for the analysis of genome-wide expression data

A classification-based machine learning approach for the analysis of genome-wide expression data
复制标题

DOI:
10.1101/gr.104003
复制
发表时间:
2003-03-01
期刊:
影响因子:
7
通讯作者:
Bhattacharya, S
Bhattacharya, S
中科院分区:
生物学1区
文献类型:
--
作者:
Lyons-Weiler, J;Patel, S;Bhattacharya, S

文献摘要

被引文献

相似文献

用于全局基因表达分析的数据分析的三个重要领域是类别发现、类别预测和寻找clysregulated基因(生物标志物)。微阵列数据的临床应用将需要其表达模式被充分理解的标记基因,以允许对疾病亚类成员的准确预测。常用的分析方法包括分层聚类算法、t-、F-和Z-检验以及机器学习方法。我们描述了一种称为最大差异子集(MDSS)算法的方法,该算法结合了分类算法,经典统计和机器学习的元素,并提供了一个连贯的框架。通过整合预测精度,MDSS算法学习统计显著性的临界阈值(a或P值),消除了设置统计显著性阈值的任意性,并最大限度地减少了正态假设的影响。为了降低假阳性率和提高预测基因集的外部有效性,使用了折刀步骤。该步骤鉴定并去除初始MDSS中具有低组合预测效用的基因。整体MDSS提供了一个预测,是不太依赖于任意的研究设计(样本纳入或排除),因此应该有很高的外部效度。我们证明,这种方法,不同于其他已发表的方法,确定生物标志物能够预测蒽环类药物-阿糖胞苷化疗的急性髓细胞白血病的情况下的结果。通过结合两个标准-统计显著性和预测效用-该方法学习与给定数据集相关的显著性水平。MDSS方法可以与任何测试和分类器操作符对一起使用。
Three important areas of data analysis for global gene expression analysis are class discovery, class prediction, and finding clysregulated genes (biomarkers). The clinical application of microarray data will require marker genes whose expression patterns are sufficiently well understood to allow accurate predictions on disease subclass membership. Commonly used methods of analysis include hierarchical clusterin g algorithms, t-, F-, and Z-tests, and machine learning approaches. We describe an approach called the maximum difference subset (MDSS) algorithm that combines classification algorithms, classical statistics, and elements of machine learning and provides a coherent framework. By integrating prediction accuracy, the MDSS algorithm learns the critical threshold of statistical significance (the a or P-value), eliminating the arbitrariness of setting a threshold of statistical significance and minimizing the effect of the normality assumptions. To reduce the false positive rate and to increase external validity of the predictive gene set, a jackknife step is used. This step identifies and removes genes in the initial MDSS with low combined predictive utility. The overall MDSS provides a prediction that is less dependent on an arbitrary study design (sample inclusion or exclusion) and should thus have high external validity. We demonstrate that this approach, unlike other published methods, identifies biomarkers capable of predicting the outcome of anthracycline-cytarabine chemotherapy in cases of acute myeloid leukemia. By incorporating two criteria-statistical significance and predictive utility-the approach learns the significance level relevant for a given data set. The MDSS approach can be used with any test and classifier operator pair.