Empirical characterization of random forest variable importance measures

Empirical characterization of random forest variable importance measures
复制标题

DOI:
10.1016/j.csda.2007.08.015
复制
发表时间:
2008-01-10
影响因子:
1.8
通讯作者:
Kirnes, Ryan V.
Kirnes, Ryan V.
中科院分区:
数学3区
文献类型:
--
作者:
Archer, Kelfie J.;Kirnes, Ryan V.

文献摘要

被引文献

相似文献

微阵列研究产生的数据集包括大量的候选预测因子(基因)对少量的观察(样品)。当兴趣在于使用基因表达数据预测表型类时,通常目标是产生准确的分类器并揭示问题的预测结构。大多数机器学习方法,如k-最近邻,支持向量机和神经网络,都适用于分类。然而,这些方法没有提供关于对预测结构最有贡献的协变量的见解。其他方法,如线性判别分析,需要预测空间被大大减少之前,推导出的分类器。最近开发的方法,随机森林(RF),不需要减少的预测空间之前的分类。此外,RF产生每个候选预测因子的变量重要性度量。本研究检验了RF变量重要性度量在大量候选预测因子中识别真实预测因子的有效性。一个广泛的模拟研究进行了使用20个水平的预测变量之间的相关性和7个水平的真实预测和二分反应之间的关联。我们的结论是,RF方法是有吸引力的分类问题时,研究的目标是产生一个准确的分类器,并提供有关个人预测变量的判别能力的见解。这样的目标在微阵列研究中是常见的,因此,在微阵列数据集上证明了用于获得变量重要性度量的RF方法的应用。(c)2007 Elsevier B. V.保留所有权利。
Microarray studies yield data sets consisting of a large number of candidate predictors (genes) on a small number of observations (samples). When interest lies in predicting phenotypic class using gene expression data, often the goals are both to produce an accurate classifier and to uncover the predictive structure of the problem. Most machine learning methods, such as k-nearest neighbors, support vector machines, and neural networks, are useful for classification. However, these methods provide no insight regarding the covariates that best contribute to the predictive structure. Other methods, such as linear discriminant analysis, require the predictor space be substantially reduced prior to deriving the classifier. A recently developed method, random forests (RF), does not require reduction of the predictor space prior to classification. Additionally, RF yield variable importance measures for each candidate predictor. This study examined the effectiveness of RF variable importance measures in identifying the true predictor among a large number of candidate predictors. An extensive simulation study was conducted using 20 levels of correlation among the predictor variables and 7 levels of association between the true predictor and the dichotomous response. We conclude that the RF methodology is attractive for use in classification problems when the goals of the study are to produce an accurate classifier and to provide insight regarding the discriminative ability of individual predictor variables. Such goals are common among microarray studies, and therefore application of the RF methodology for the purpose of obtaining variable importance measures is demonstrated on a microarray data set.. (c) 2007 Elsevier B.V. All rights reserved.