A comprehensive comparison of random forests and support vector machines for microarray-based cancer classification.

A comprehensive comparison of random forests and support vector machines for microarray-based cancer classification.
复制标题

DOI:
10.1186/1471-2105-9-319
复制
发表时间:
2008-07-22
期刊:
影响因子:
3
通讯作者:
Aliferis, Constantin F.
Aliferis, Constantin F.
中科院分区:
生物学4区
文献类型:
--
作者:
Statnikov, Alexander;Wang, Lily;Aliferis, Constantin F.

文献摘要

参考文献

被引文献

相似文献

癌症诊断和临床结果预测是基因表达微阵列技术的最重要的新兴应用之一,其中有几个分子标记正在走向临床部署。使用可用于微阵列基因表达数据的最准确的分类算法是一个关键因素,以便为患者护理开发最佳可能的分子特征。如迄今为止的大量文献所建议的,支持向量机可以被认为是用于此类数据分类的“最佳”算法。然而,最近的工作表明,随机森林分类器在这一领域可能优于支持向量机。在本文中,我们确定了方法的偏见,以前的工作比较随机森林和支持向量机,并进行了新的严格的评估,纠正这些局限性的两种算法。我们的实验使用了22个诊断和预测数据集,并表明支持向量机的性能优于随机森林,通常是大幅度的。我们的数据还强调了合理的研究设计在生物信息学算法的基准测试和比较中的重要性。我们发现,平均而言,在大多数微阵列数据集,随机森林优于支持向量机的设置时,没有进行基因选择,当几个流行的基因选择方法被使用。
Cancer diagnosis and clinical outcome prediction are among the most important emerging applications of gene expression microarray technology with several molecular signatures on their way toward clinical deployment. Use of the most accurate classification algorithms available for microarray gene expression data is a critical ingredient in order to develop the best possible molecular signatures for patient care. As suggested by a large body of literature to date, support vector machines can be considered "best of class" algorithms for classification of such data. Recent work, however, suggests that random forest classifiers may outperform support vector machines in this domain. In the present paper we identify methodological biases of prior work comparing random forests and support vector machines and conduct a new rigorous evaluation of the two algorithms that corrects these limitations. Our experiments use 22 diagnostic and prognostic datasets and show that support vector machines outperform random forests, often by a large margin. Our data also underlines the importance of sound research design in benchmarking and comparison of bioinformatics algorithms. We found that both on average and in the majority of microarray datasets, random forests are outperformed by support vector machines both in the settings when no gene selection is performed and when several popular gene selection methods are used.
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1137/s0036144502411986
发表时间: 2003-12-01
期刊: SIAM REVIEW
影响因子: 10.2
作者:
Rifkin, R;Mukherjee, S;Mesirov, JP
通讯作者: Mesirov, JP
DOI: 10.1093/bioinformatics/bti033
发表时间: 2005-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Statnikov, A;Aliferis, CF;Levy, S
通讯作者: Levy, S
DOI: 10.1186/1471-2105-7-3
发表时间: 2006-01-06
期刊: BMC bioinformatics
影响因子: 3
作者:
Díaz-Uriarte R;Alvarez de Andrés S
通讯作者: Alvarez de Andrés S
DOI: 10.1016/j.patrec.2006.03.013
发表时间: 2006-10-15
影响因子: 5.1
作者:
Chen, Xue-wen;Zeng, Xiangyan;van Alphen, Deborah
通讯作者: van Alphen, Deborah