A random forest guided tour

A random forest guided tour
复制标题

DOI:
10.1007/s11749-016-0481-7
复制
发表时间:
2016-06-01
期刊:
影响因子:
1.3
通讯作者:
Scornet, Erwan
Scornet, Erwan
中科院分区:
数学2区
文献类型:
--
作者:
Biau, Gerard;Scornet, Erwan

文献摘要

被引文献

相似文献

随机森林算法是L. Breiman在2001年提出的方法,作为一种通用的分类和回归方法已经非常成功。该方法结合了几个随机决策树,并通过平均来聚合它们的预测,在变量数量远大于观测数量的情况下表现出出色的性能。此外,它是通用的,足以适用于大规模的问题,很容易适应各种特设的学习任务,并返回变量的重要性的措施。本文综述了随机森林理论和方法的最新进展。重点放在数学力量驱动的算法,特别注意到参数的选择,resception机制,变量的重要性措施。这篇评论旨在为非专家提供方便的主要思想。
The random forest algorithm, proposed by L. Breiman in 2001, has been extremely successful as a general-purpose classification and regression method. The approach, which combines several randomized decision trees and aggregates their predictions by averaging, has shown excellent performance in settings where the number of variables is much larger than the number of observations. Moreover, it is versatile enough to be applied to large-scale problems, is easily adapted to various ad hoc learning tasks, and returns measures of variable importance. The present article reviews the most recent theoretical and methodological developments for random forests. Emphasis is placed on the mathematical forces driving the algorithm, with special attention given to the selection of parameters, the resampling mechanism, and variable importance measures. This review is intended to provide non-experts easy access to the main ideas.