A computationally fast variable importance test for random forests for high-dimensional data

A computationally fast variable importance test for random forests for high-dimensional data
复制标题

DOI:
10.1007/s11634-016-0276-4
复制
发表时间:
2018-12-01
影响因子:
1.6
通讯作者:
Boulesteix, Anne-Laure
Boulesteix, Anne-Laure
中科院分区:
计算机科学3区
文献类型:
--
作者:
Janitza, Silke;Celik, Ender;Boulesteix, Anne-Laure

文献摘要

被引文献

相似文献

随机森林是一种常用的工具,用于分类和基于所谓的变量重要性度量对候选预测因子进行排名。这些措施将分数归因于反映其重要性的变量。变量重要性度量的一个缺点是没有自然截止值可以用来区分重要和不重要的变量。为解决这一问题,开发了几种方法,例如基于假设检验的方法。现有的测试方法需要对随机森林进行重复计算。虽然对于低维设置,这些方法可能在计算上易于处理,但对于通常包括数千个候选预测因子的高维设置,计算时间是巨大的。在这篇文章中,提出了一种计算速度快的启发式变量重要性测试,适用于高维数据,其中许多变量不携带任何信息。测试方法是基于修改后的版本的置换变量的重要性,这是交叉验证程序的启发。新的方法进行了测试和比较的方法Altmann和同事使用模拟研究,这是基于真实的数据从高维二进制分类设置。新方法控制I类错误,并在研究中以更小的计算时间至少具有可比的功效。因此,它可能被用作一个计算快速替代现有的程序高维数据设置,其中许多变量不携带任何信息。新方法在R包vita中实现。
Random forests are a commonly used tool for classification and for ranking candidate predictors based on the so-called variable importance measures. These measures attribute scores to the variables reflecting their importance. A drawback of variable importance measures is that there is no natural cutoff that can be used to discriminate between important and non-important variables. Several approaches, for example approaches based on hypothesis testing, were developed for addressing this problem. The existing testing approaches require the repeated computation of random forests. While for low-dimensional settings those approaches might be computationally tractable, for high-dimensional settings typically including thousands of candidate predictors, computing time is enormous. In this article a computationally fast heuristic variable importance test is proposed that is appropriate for high-dimensional data where many variables do not carry any information. The testing approach is based on a modified version of the permutation variable importance, which is inspired by cross-validation procedures. The new approach is tested and compared to the approach of Altmann and colleagues using simulation studies, which are based on real data from high-dimensional binary classification settings. The new approach controls the type I error and has at least comparable power at a substantially smaller computation time in the studies. Thus, it might be used as a computationally fast alternative to existing procedures for high-dimensional data settings where many variables do not carry any information. The new approach is implemented in the R package vita.