Variable selection using random forests

Variable selection using random forests
复制标题

DOI:
10.1016/j.patrec.2010.03.014
复制
发表时间:
2010-10-15
影响因子:
5.1
通讯作者:
Tuleau-Malot, Christine
Tuleau-Malot, Christine
中科院分区:
计算机科学3区
文献类型:
--
作者:
Genuer, Robin;Poggi, Jean-Michel;Tuleau-Malot, Christine

文献摘要

被引文献

相似文献

本文以随机森林为研究对象,针对Leo Breiman(2001)提出的分类和回归问题中越来越常用的统计方法,探讨了变量选择的两个经典问题。第一个是寻找重要的变量进行解释,第二个是更严格的限制,并试图设计一个好的简洁的预测模型。主要贡献有两个方面:提供了一些基于随机森林的变量重要性指数行为的实验见解,并提出了一种使用随机森林重要性评分对解释变量进行排序的策略和逐步上升的变量引入策略。(C) 2010 Elsevier B.V.版权所有
This paper proposes, focusing on random forests, the increasingly used statistical method for classification and regression problems introduced by Leo Breiman in 2001, to investigate two classical issues of variable selection. The first one is to find important variables for interpretation and the second one is more restrictive and try to design a good parsimonious prediction model. The main contribution is twofold: to provide some experimental insights about the behavior of the variable importance index based on random forests and to propose a strategy involving a ranking of explanatory variables using the random forests score of importance and a stepwise ascending variable introduction strategy. (C) 2010 Elsevier B.V. All rights reserved.