Bias in random forest variable importance measures: illustrations, sources and a solution.

Bias in random forest variable importance measures: illustrations, sources and a solution.
复制标题

随机森林中的偏见重要性措施:插图,来源和解决方案。

DOI:
10.1186/1471-2105-8-25
复制
发表时间:
2007-01-25
期刊:
影响因子:
3
通讯作者:
Hothorn, Torsten
Hothorn, Torsten
中科院分区:
生物学4区
文献类型:
--
作者:
Strobl, Carolin;Boulesteix, Anne-Laure;Zeileis, Achim;Hothorn, Torsten

文献摘要

被引文献

相似文献

随机森林的变量重要性度量作为生物信息学和相关科学领域中许多分类任务中的变量选择手段,例如选择与预测某种疾病相关的遗传标记的子集,已经受到越来越多的关注。我们发现,随机森林变量的重要性措施是一个明智的手段,在许多应用程序中的变量选择,但不可靠的情况下,潜在的预测变量的测量尺度或类别的数量不同。这在基因组学和计算生物学中特别重要,其中预测因子通常包括不同类型的变量,例如当预测因子包括序列数据和连续变量(例如折叠能量)时,或者当氨基酸序列数据显示不同数量的类别时。模拟研究表明,当随机森林变量的重要性措施与不同类型的数据,结果是误导,因为次优预测变量可能是人为的变量选择的首选。这一缺陷的两个机制是有偏见的变量选择的个人分类树用于建立随机森林,一方面,与替换的自助抽样另一方面引起的影响。我们建议采用随机森林的另一种实现方式,在个体分类树中提供无偏的变量选择。当这种方法使用子采样而不替换时,即使在潜在的预测变量在其测量尺度或类别数量方面有所不同的情况下,所得到的变量重要性度量也可以可靠地用于变量选择。在R系统中使用随机森林算法及其变量重要性度量进行统计计算,并在重新分析RNA编辑研究数据的应用程序中进行了详细说明和记录。因此,建议的方法可以直接应用于生物信息学研究的科学家。
Variable importance measures for random forests have been receiving increased attention as a means of variable selection in many classification tasks in bioinformatics and related scientific fields, for instance to select a subset of genetic markers relevant for the prediction of a certain disease. We show that random forest variable importance measures are a sensible means for variable selection in many applications, but are not reliable in situations where potential predictor variables vary in their scale of measurement or their number of categories. This is particularly important in genomics and computational biology, where predictors often include variables of different types, for example when predictors include both sequence data and continuous variables such as folding energy, or when amino acid sequence data show different numbers of categories. Simulation studies are presented illustrating that, when random forest variable importance measures are used with data of varying types, the results are misleading because suboptimal predictor variables may be artificially preferred in variable selection. The two mechanisms underlying this deficiency are biased variable selection in the individual classification trees used to build the random forest on one hand, and effects induced by bootstrap sampling with replacement on the other hand. We propose to employ an alternative implementation of random forests, that provides unbiased variable selection in the individual classification trees. When this method is applied using subsampling without replacement, the resulting variable importance measures can be used reliably for variable selection even in situations where the potential predictor variables vary in their scale of measurement or their number of categories. The usage of both random forest algorithms and their variable importance measures in the R system for statistical computing is illustrated and documented thoroughly in an application re-analyzing data from a study on RNA editing. Therefore the suggested method can be applied straightforwardly by scientists in bioinformatics research.