Improved prediction of tree species richness and interpretability of environmental drivers using a machine learning approach

Improved prediction of tree species richness and interpretability of environmental drivers using a machine learning approach
复制标题

DOI:
10.1016/j.foreco.2023.120972
复制
发表时间:
2023-04-21
影响因子:
3.7
通讯作者:
Kedron,Peter
Kedron,Peter
中科院分区:
农林科学1区
文献类型:
--
作者:
Brugere,Lian;Kwon,Youngsang;Kedron,Peter

文献摘要

相似文献

全球生物多样性正在下降,如果要扭转目前的趋势,预测物种多样性至关重要。树种丰富度(TSR)长期以来一直是生物多样性的一个关键指标,但目前的模型存在相当大的不确定性,特别是考虑到经典的统计假设和机器学习结果的生态可解释性较差。在这里,我们测试了几种生态可解释的机器学习方法来预测TSR并解释美国大陆的驱动环境因素。我们开发了两个人工神经网络(ANN)和一个随机森林(RF)模型来预测TSR使用森林资源清查和分析数据和20个环境协变量,并将它们与经典的广义线性模型(GLM)进行比较。使用R2和平均绝对误差(MAE)和残差空间自相关分析在独立的、看不见的测试数据集上对模型进行评估。一个可解释的机器学习方法,SHapley加法解释(SHAP),被用来解释驱动TSR的主要环境因素。与基线GLM(R2=0.7; MAE = 4.7)相比,ANN和RF模型的R2大于0.9,MAE<3.1。此外,ANN和RF模型产生的空间聚类TSR残差比GLM少。SHAP分析表明,干旱指数、森林面积、海拔高度、最干季平均降水量和年平均气温对TSR的预测效果最好。SHAP进一步揭示了环境协变量与TSR的非线性关系以及GLM未揭示的复杂相互作用。这项研究强调,需要努力保护森林地区,减少降水对低森林但干旱地区树种造成的生理压力。这里使用的机器学习方法可用于其他生物的生物多样性研究或未来气候情景下的TSR预测。
Biodiversity is in decline globally and predicting species diversity is critically important if current trends are to be reversed. Tree species richness (TSR) has long been a key measure of biodiversity, but considerable uncertainties exist in current models, particularly given the classic statistical assumptions and poor ecological interpretability of machine learning outcomes. Here, we test several ecologically interpretable machine learning approaches to predict TSR and interpret the driving environmental factors in the continental United States. We develop two artificial neural networks (ANN) and one random forest (RF) model to predict TSR using Forest Inventory and Analysis data and 20 environmental covariates and compare them to a classic generalized linear model (GLM). Models were evaluated on an independent, unseen testing dataset usingR2and Mean Absolute Error (MAE) and residual spatial autocorrelation analysis. An Interpretable Machine Learning approach, SHapley Additive exPlanations (SHAP), was adopted to explain the major environmental factors driving TSR. Compared to a baseline GLM (R2=0.7; MAE = 4.7), the ANN and RF models achievedR2greater than 0.9 and MAE<3.1. Additionally, the ANN and RF models produced less spatially clustered TSR residuals than the GLM. SHAP analysis suggested that TSR is best predicted by Aridity Index, Forest Area, Altitude, Mean Precipitation of the Driest Quarter and Mean Annual Temperature. SHAP further revealed a non-linear relationship of environmental covariates with TSR and complex interactions that were not revealed by the GLM. The study highlights the need for conservation efforts of forest areas and reducing precipitation-related physiological stress on tree species in low forested but arid regions. The machine learning approach used here is transferrable for studies of biodiversity for other organisms or prediction of TSR under future climatic scenarios.