Gene selection and classification of microarray data using random forest.

Gene selection and classification of microarray data using random forest.
复制标题

DOI:
10.1186/1471-2105-7-3
复制
发表时间:
2006-01-06
期刊:
影响因子:
3
通讯作者:
Alvarez de Andrés S
Alvarez de Andrés S
中科院分区:
生物学4区
文献类型:
--
作者:
Díaz-Uriarte R;Alvarez de Andrés S

文献摘要

参考文献

被引文献

相似文献

在大多数基因表达研究中,选择相关基因进行样本分类是一项常见的任务,研究人员试图识别仍然可以实现良好预测性能的最小可能基因集(例如,用于临床实践中的诊断目的)。许多基因选择方法使用基因相关性的单变量(逐个基因)排名和任意阈值来选择基因的数量,只能应用于两类问题,并且使用与分类算法无关的基因选择排名标准。相比之下,随机森林是一种非常适合微阵列数据的分类算法:即使在大多数预测变量都是噪声的情况下,它也表现出出色的性能,当变量的数量远远大于观测值的数量时,以及涉及两个以上类别的问题时,它可以使用,并返回变量重要性的度量。因此,重要的是要了解随机森林与微阵列数据的性能及其可能用于基因选择。研究了随机森林在微阵列数据分类(包括多类问题)中的应用,提出了一种基于随机森林的基因选择方法。使用模拟和9个微阵列数据集,我们表明,随机森林具有可比的性能,其他分类方法,包括DLDA,KNN和SVM,新的基因选择程序产生非常小的基因集(通常小于替代方法),同时保持预测精度。由于其性能和特征,随机森林和使用随机森林的基因选择可能应该成为使用微阵列数据进行类别预测和基因选择的方法的“标准工具箱”的一部分。
Selection of relevant genes for sample classification is a common task in most gene expression studies, where researchers try to identify the smallest possible set of genes that can still achieve good predictive performance (for instance, for future use with diagnostic purposes in clinical practice). Many gene selection approaches use univariate (gene-by-gene) rankings of gene relevance and arbitrary thresholds to select the number of genes, can only be applied to two-class problems, and use gene selection ranking criteria unrelated to the classification algorithm. In contrast, random forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of observations and in problems involving more than two classes, and returns measures of variable importance. Thus, it is important to understand the performance of random forest with microarray data and its possible use for gene selection. We investigate the use of random forest for classification of microarray data (including multi-class problems) and propose a new method of gene selection in classification problems based on random forest. Using simulated and nine microarray data sets we show that random forest has comparable performance to other classification methods, including DLDA, KNN, and SVM, and that the new gene selection procedure yields very small sets of genes (often smaller than alternative methods) while preserving predictive accuracy. Because of its performance and features, random forest and gene selection using random forest should probably become part of the "standard tool-box" of methods for class prediction and gene selection with microarray data.
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
使用随机森林绘制复杂性状。
DOI: 10.1186/1471-2156-4-s1-s64
发表时间: 2003-12-31
期刊: BMC GENETICS
影响因子: 2.9
作者:
Bureau, A;Dupuis, J;Hayward, B;Falls, K;Van Eerdewegh, P
通讯作者: Van Eerdewegh, P
DOI: 10.1038/35000501
发表时间: 2000-02-03
期刊: NATURE
影响因子: 64.8
作者:
Alizadeh, AA;Eisen, MB;Staudt, LM
通讯作者: Staudt, LM
DOI: 10.1023/a:1009715923555
发表时间: 1998-06-01
影响因子: 4.8
作者:
Burges, CJC
通讯作者: Burges, CJC
DOI: 10.1093/bioinformatics/bth469
发表时间: 2005-01-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Ein-Dor, L;Kela, I;Domany, E
通讯作者: Domany, E