Random forests and stochastic gradient boosting for predicting tree canopy cover: comparing tuning processes and model performance

Random forests and stochastic gradient boosting for predicting tree canopy cover: comparing tuning processes and model performance
复制标题

DOI:
10.1139/cjfr-2014-0562
复制
发表时间:
2016-03-01
影响因子:
2.2
通讯作者:
Wilson, Barry T.
Wilson, Barry T.
中科院分区:
农林科学3区
文献类型:
--
作者:
Freeman, Elizabeth A.;Moisen, Gretchen G.;Wilson, Barry T.

文献摘要

被引文献

相似文献

作为2011年国家土地覆盖数据库(NLCD)树冠覆盖层开发的一部分,启动了一个试点项目,以测试使用高分辨率摄影加上广泛的辅助数据,以绘制美国接壤的四个研究区域的树冠覆盖分布图。两种随机建模技术,随机森林(RF)和随机梯度提升(SGB),进行了比较。本研究的目的是第一,探讨RF和SGB的灵敏度,选择在调整参数,第二,通过评估的重要性,并相互作用,预测变量,来自一个独立的测试集的全球准确性指标,以及视觉质量的树冠覆盖所得到的地图,比较两个最终模型的性能。RF和SGB的预测准确性在我们的所有四个试点地区都非常相似。在所有四个研究区域中,独立检验集均方误差(MSE)均为小数点后三位,其中堪萨斯差异最大,RF的MSE为0.0113,SGB的MSE为0.0117。与相关的预测变量,SGB有倾向于集中变量的重要性在较少的变量,而RF倾向于传播的重要性在更多的变量。RF比SGB更容易实现,因为RF需要调整的参数更少,并且对这些参数不太敏感。作为随机技术,RF和SGB都引入了新的不确定性成分:重复的模型运行可能会导致不同的最终预测。我们演示了如何RF允许生产的空间上明确的地图,这种随机的不确定性的最终模型。
As part of the development of the 2011 National Land Cover Database (NLCD) tree canopy cover layer, a pilot project was launched to test the use of high-resolution photography coupled with extensive ancillary data to map the distribution of tree canopy cover over four study regions in the conterminous US. Two stochastic modeling techniques, random forests (RF) and stochastic gradient boosting (SGB), are compared. The objectives of this study were first to explore the sensitivity of RF and SGB to choices in tuning parameters and, second, to compare the performance of the two final models by assessing the importance of, and interaction between, predictor variables, the global accuracy metrics derived from an independent test set, as well as the visual quality of the resultant maps of tree canopy cover. The predictive accuracy of RF and SGB was remarkably similar on all four of our pilot regions. In all four study regions, the independent test set mean squared error (MSE) was identical to three decimal places, with the largest difference in Kansas where RF gave an MSE of 0.0113 and SGB gave an MSE of 0.0117. With correlated predictor variables, SGB had a tendency to concentrate variable importance in fewer variables, whereas RF tended to spread importance among more variables. RF is simpler to implement than SGB, as RF has fewer parameters needing tuning and also was less sensitive to these parameters. As stochastic techniques, both RF and SGB introduce a new component of uncertainty: repeated model runs will potentially result in different final predictions. We demonstrate how RF allows the production of a spatially explicit map of this stochastic uncertainty of the final model.