An introduction to recursive partitioning: rationale, application, and characteristics of classification and regression trees, bagging, and random forests.

An introduction to recursive partitioning: rationale, application, and characteristics of classification and regression trees, bagging, and random forests.
复制标题

DOI:
10.1037/a0016973
复制
发表时间:
2009-12
影响因子:
7
通讯作者:
Tutz, Gerhard
Tutz, Gerhard
中科院分区:
心理学1区
文献类型:
--
作者:
Strobl, Carolin;Malley, James;Tutz, Gerhard

文献摘要

参考文献

被引文献

相似文献

递归划分方法在许多科学领域已经成为非参数回归和分类的流行和广泛使用的工具。特别是随机森林,即使在存在复杂相互作用的情况下也能处理大量的预测变量,在过去的几年里已经在遗传学、临床医学和生物信息学中得到了成功的应用。高维问题不仅在遗传学中很常见,在心理学研究的一些领域也是如此,由于时间或成本的限制,只有少数几个对象可以测量,但每个对象都会产生大量数据。随机森林已被证明在这类应用中实现了很高的预测精度,并提供了描述性变量重要性衡量标准,反映了每个变量在主要影响和相互作用中的影响。这项工作的目的是介绍标准递归划分方法的原理以及最近的方法改进,以说明它们在低维和高维数据探索中的用途,但也指出了这些方法的局限性和实际应用中的潜在陷阱。使用R系统中用于统计计算的可自由获得的实现来说明这些方法的应用。
Recursive partitioning methods have become popular and widely used tools for non-parametric regression and classification in many scientific fields. Especially random forests, that can deal with large numbers of predictor variables even in the presence of complex interactions, have been applied successfully in genetics, clinical medicine and bioinformatics within the past few years. High dimensional problems are common not only in genetics, but also in some areas of psychological research, where only few subjects can be measured due to time or cost constraints, yet a large amount of data is generated for each subject. Random forests have been shown to achieve a high prediction accuracy in such applications, and provide descriptive variable importance measures reflecting the impact of each variable in both main effects and interactions. The aim of this work is to introduce the principles of the standard recursive partitioning methods as well as recent methodological improvements, to illustrate their usage for low and high dimensional data exploration, but also to point out limitations of the methods and potential pitfalls in their practical application. Application of the methods is illustrated using freely available implementations in the R system for statistical computing.
DOI: 10.1214/009053606000000092
发表时间: 2006-04-01
影响因子: 4.5
作者:
Buhlmann, Peter
通讯作者: Buhlmann, Peter
DOI: 10.1109/tnn.1997.641482
发表时间: 1997-01-01
影响因子: --
作者:
Cherkassky, V
通讯作者: Cherkassky, V
DOI: 10.1023/a:1007515423169
发表时间: 1999-07-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Bauer, E;Kohavi, R
通讯作者: Kohavi, R
DOI: 10.1214/ss/1009213726
发表时间: 2001-08-01
影响因子: 5.7
作者:
Breiman, L
通讯作者: Breiman, L
DOI: 10.1198/016214503000125
发表时间: 2003-06-01
影响因子: 3.7
作者:
Bühlmann, P;Yu, B
通讯作者: Yu, B