Robustness of Random Forest-based gene selection methods.

Robustness of Random Forest-based gene selection methods.
复制标题

DOI:
10.1186/1471-2105-15-8
复制
发表时间:
2014-01-13
期刊:
影响因子:
3
通讯作者:
Kursa MB
Kursa MB
中科院分区:
生物学4区
文献类型:
--
作者:
Kursa MB

文献摘要

参考文献

被引文献

相似文献

基因选择是微阵列数据分析的重要组成部分,因为它提供的信息可以导致对所调查现象的更好的机制理解。同时,由于微阵列数据的噪声特性,基因选择非常困难。因此,基因选择通常是用机器学习方法进行的。随机森林方法特别适合于这个目的。在这项工作中,在基因选择的背景下,比较了四种最先进的基于随机森林的特征选择方法。分析的重点是选择的稳定性,因为虽然它是确定结果的重要性所必需的,但在类似的研究中往往被忽视。随机森林分类器验证的选择后准确性的比较表明,在这种情况下,所有调查的方法都是等效的。然而,在选择基因的数量和选择的稳定性方面,这些方法有很大的不同。在分析的方法中,Boruta算法预测了大多数可能重要的基因。选择后分类器错误率是一种常用的测量方法,但被发现是一种潜在的欺骗性的基因选择质量测量方法。当考虑一致选择的基因数量时,Boruta算法显然是最好的。虽然它也是计算最密集的方法,但通过使用Random Ferns(一个类似但简化的分类器)的类似度量来替换Random Forest的重要性,Boruta算法的计算需求可以降低到与其他算法相当的水平。尽管他们的设计假设,最小最优选择方法,被发现选择假阳性的高比例。
Gene selection is an important part of microarray data analysis because it provides information that can lead to a better mechanistic understanding of an investigated phenomenon. At the same time, gene selection is very difficult because of the noisy nature of microarray data. As a consequence, gene selection is often performed with machine learning methods. The Random Forest method is particularly well suited for this purpose. In this work, four state-of-the-art Random Forest-based feature selection methods were compared in a gene selection context. The analysis focused on the stability of selection because, although it is necessary for determining the significance of results, it is often ignored in similar studies. The comparison of post-selection accuracy of a validation of Random Forest classifiers revealed that all investigated methods were equivalent in this context. However, the methods substantially differed with respect to the number of selected genes and the stability of selection. Of the analysed methods, the Boruta algorithm predicted the most genes as potentially important. The post-selection classifier error rate, which is a frequently used measure, was found to be a potentially deceptive measure of gene selection quality. When the number of consistently selected genes was considered, the Boruta algorithm was clearly the best. Although it was also the most computationally intensive method, the Boruta algorithm’s computational demands could be reduced to levels comparable to those of other algorithms by replacing the Random Forest importance with a comparable measure from Random Ferns (a similar but simplified classifier). Despite their design assumptions, the minimal optimal selection methods, were found to select a high fraction of false positives.
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.18637/jss.v036.i11
发表时间: 2010-09-01
影响因子: 5.8
作者:
Kursa, Miron B.;Rudnicki, Witold R.
通讯作者: Rudnicki, Witold R.
DOI: 10.1016/j.patcog.2013.05.018
发表时间: 2013-12-01
影响因子: 8
作者:
Deng, Houtao;Runger, George
通讯作者: Runger, George
DOI: 10.1109/tpami.2009.23
发表时间: 2010-03-01
影响因子: 23.6
作者:
Oezuysal, Mustafa;Calonder, Michael;Fua, Pascal
通讯作者: Fua, Pascal
DOI: 10.1186/1471-2105-7-3
发表时间: 2006-01-06
期刊: BMC bioinformatics
影响因子: 3
作者:
Díaz-Uriarte R;Alvarez de Andrés S
通讯作者: Alvarez de Andrés S