Correlation and variable importance in random forests

Correlation and variable importance in random forests
复制标题

DOI:
10.1007/s11222-016-9646-1
复制
发表时间:
2017-05-01
影响因子:
2.2
通讯作者:
Saint-Pierre, Philippe
Saint-Pierre, Philippe
中科院分区:
数学2区
文献类型:
--
作者:
Gregorutti, Baptiste;Michel, Bertrand;Saint-Pierre, Philippe

文献摘要

被引文献

相似文献

本文是关于存在相关预测因子的随机森林算法的变量选择问题。在高维回归或分类框架中,变量选择是一项困难的任务,当存在高度相关的预测因素时,这一任务变得更加具有挑战性。首先,我们对加性回归模型的排列重要度进行了理论研究。这使我们能够描述预测者之间的相关性如何影响排列重要性。我们的结果推动了递归特征消除(RFE)算法在此背景下用于变量选择。该算法以排列重要度作为排序准则,递归地剔除变量。接下来,通过各种仿真实验验证了RFE算法在选择少量变量时的有效性,同时具有较好的预测误差。最后,利用UCI机器学习库中的Landsat卫星数据对该选择算法进行了测试。
This paper is about variable selection with the random forests algorithm in presence of correlated predictors. In high-dimensional regression or classification frameworks, variable selection is a difficult task, that becomes even more challenging in the presence of highly correlated predictors. Firstly we provide a theoretical study of the permutation importance measure for an additive regression model. This allows us to describe how the correlation between predictors impacts the permutation importance. Our results motivate the use of the recursive feature elimination (RFE) algorithm for variable selection in this context. This algorithm recursively eliminates the variables using permutation importance measure as a ranking criterion. Next various simulation experiments illustrate the efficiency of the RFE algorithm for selecting a small number of variables together with a good prediction error. Finally, this selection algorithm is tested on the Landsat Satellite data from the UCI Machine Learning Repository.