Random Forests: some methodological insights

Random Forests: some methodological insights
复制标题

DOI:
--
复制
发表时间:
2008-11
期刊:
arXiv: Machine Learning
影响因子:
--
通讯作者:
R. Genuer;Jean-Michel Poggi;Christine Tuleau
R. Genuer;Jean-Michel Poggi;Christine Tuleau
中科院分区:
其他
文献类型:
--
作者:
R. Genuer;Jean-Michel Poggi;Christine Tuleau

文献摘要

被引文献

相似文献

本文从实验的角度研究了随机森林,这是Leo Breiman在2001年提出的用于分类和回归问题的越来越多的统计方法。它首先旨在确认,已知的,但稀疏的,建议使用随机森林,并提出一些补充意见的标准问题,以及高维的变量的数量大大超过样本量。但本文的主要贡献是双重的:提供一些见解的行为的变量的重要性指数的基础上随机森林,此外,提出调查变量选择的两个经典问题。第一个是找到重要的变量进行解释,第二个是更具限制性,并试图设计一个好的预测模型。该策略包括使用随机森林重要性得分和逐步上升的变量引入策略对解释变量进行排名。
This paper examines from an experimental perspective random forests, the increasingly used statistical method for classification and regression problems introduced by Leo Breiman in 2001. It first aims at confirming, known but sparse, advice for using random forests and at proposing some complementary remarks for both standard problems as well as high dimensional ones for which the number of variables hugely exceeds the sample size. But the main contribution of this paper is twofold: to provide some insights about the behavior of the variable importance index based on random forests and in addition, to propose to investigate two classical issues of variable selection. The first one is to find important variables for interpretation and the second one is more restrictive and try to design a good prediction model. The strategy involves a ranking of explanatory variables using the random forests score of importance and a stepwise ascending variable introduction strategy.