Random forest versus logistic regression: a large-scale benchmark experiment.

Random forest versus logistic regression: a large-scale benchmark experiment.
复制标题

DOI:
10.1186/s12859-018-2264-5
复制
发表时间:
2018-07-17
期刊:
影响因子:
3
通讯作者:
Boulesteix AL
Boulesteix AL
中科院分区:
生物学4区
文献类型:
--
作者:
Couronné R;Probst P;Boulesteix AL

文献摘要

参考文献

被引文献

相似文献

自2001年推出以来,用于回归和分类的随机森林(RF)算法已经相当流行。与此同时,它已经发展成为一种标准的分类方法,在许多创新友好的科学领域与逻辑回归竞争。在这种情况下,我们提出了一个大规模的基准测试实验的基础上,243个真实的数据集比较的预测性能的原始版本的RF与默认参数和LR作为二进制分类工具。最重要的是,我们的基准实验的设计灵感来自临床试验方法,从而避免了常见的陷阱和主要的偏倚来源。根据在大约69%的数据集中测量的所考虑的准确度,RF的表现优于LR。RF和LR之间的平均差异为0.029(95%-CI =[0.022,0.038])准确度为0.041曲线下面积(95%-CI =[0.031,0.053])和-0.027(95%-CI =[-0.034,-0.021]),因此所有测量结果均表明RF的性能显著更好。作为基准测试实验的一个附带结果,我们观察到结果明显依赖于用于选择示例数据集的包含标准,因此强调了关于此数据集选择过程的明确声明的重要性。我们还强调,与我们类似的中立研究,基于大量数据集并经过精心设计,将来将有必要评估随机森林的其他变体,实现或参数,与默认值的原始版本相比,这些变体可能会提高准确性。本文的在线版本(10.1186/s12859-018-2264-5)包含补充材料,可供授权用户使用。
The Random Forest (RF) algorithm for regression and classification has considerably gained popularity since its introduction in 2001. Meanwhile, it has grown to a standard classification approach competing with logistic regression in many innovation-friendly scientific fields. In this context, we present a large scale benchmarking experiment based on 243 real datasets comparing the prediction performance of the original version of RF with default parameters and LR as binary classification tools. Most importantly, the design of our benchmark experiment is inspired from clinical trial methodology, thus avoiding common pitfalls and major sources of biases. RF performed better than LR according to the considered accuracy measured in approximately 69% of the datasets. The mean difference between RF and LR was 0.029 (95%-CI =[0.022,0.038]) for the accuracy, 0.041 (95%-CI =[0.031,0.053]) for the Area Under the Curve, and − 0.027 (95%-CI =[−0.034,−0.021]) for the Brier score, all measures thus suggesting a significantly better performance of RF. As a side-result of our benchmarking experiment, we observed that the results were noticeably dependent on the inclusion criteria used to select the example datasets, thus emphasizing the importance of clear statements regarding this dataset selection process. We also stress that neutral studies similar to ours, based on a high number of datasets and carefully designed, will be necessary in the future to evaluate further variants, implementations or parameters of random forests which may yield improved accuracy compared to the original version with default values. The online version of this article (10.1186/s12859-018-2264-5) contains supplementary material, which is available to authorized users.
DOI: 10.1186/s12859-016-1228-x
发表时间: 2016-09-01
期刊: BMC bioinformatics
影响因子: 3
作者:
Huang BF;Boutros PC
通讯作者: Boutros PC
DOI: 10.1198/106186006x133933
发表时间: 2006-09-01
影响因子: 2.4
作者:
Hothorn, Torsten;Hornik, Kurt;Zeileis, Achim
通讯作者: Zeileis, Achim
DOI: 10.1007/s10994-006-6226-1
发表时间: 2006-04-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Geurts, P;Ernst, D;Wehenkel, L
通讯作者: Wehenkel, L
DOI: 10.1093/pan/mpv024
发表时间: 2016-12-01
期刊: POLITICAL ANALYSIS
影响因子: 5.4
作者:
Muchlinski, David;Siroky, David;Kocher, Matthew
通讯作者: Kocher, Matthew
DOI: 10.1214/ss/1009213726
发表时间: 2001-08-01
影响因子: 5.7
作者:
Breiman, L
通讯作者: Breiman, L