Explainable Boosting Machines for Slope Failure Spatial Predictive Modeling

Explainable Boosting Machines for Slope Failure Spatial Predictive Modeling
复制标题

DOI:
10.3390/rs13244991
复制
发表时间:
2021-12
期刊:
Remote. Sens.
影响因子:
--
通讯作者:
Aaron E. Maxwell;Maneesh Sharma;Kurt A. Donaldson
Aaron E. Maxwell;Maneesh Sharma;Kurt A. Donaldson
中科院分区:
其他
文献类型:
--
作者:
Aaron E. Maxwell;Maneesh Sharma;Kurt A. Donaldson

文献摘要

被引文献

相似文献

机器学习(ML)方法,如人工神经网络(ANN)、k近邻(kNN)、随机森林(RF)、支持向量机(SVM)和增强决策树(dt),对于特定的映射和建模任务,可能比更传统的参数方法(如线性回归、多元线性回归和逻辑回归(LR))提供更强的预测性能。然而,这种性能的提高通常伴随着模型复杂性的增加和可解释性的降低,导致对其“黑箱”性质的批评,这突出了对既能提供强大的预测性能又能提供可解释性的算法的需求。当全局模型和特定数据点的预测需要解释时,这一点尤其正确,以便模型能够使用。可解释增强机(EBM)是广义可加性模型(GAMs)的一种增强和改进,是一种既可解释结果又具有较强预测性能的经验建模方法。经过训练的模型可以图形化地概括为将每个预测变量与因变量相关的一组函数,以及表示选定的预测变量对之间相互作用的热图。在这项研究中,我们评估了基于数字地形特征的EBMs在美国西弗吉尼亚州四个独立的主要土地资源区(MLRAs)预测边坡破坏发生的可能性或概率,并将结果与LR、kNN、RF和SVM进行了比较。EBM提供的预测精度与RF和SVM相当,优于LR和kNN。为每个预测变量生成的函数和可视化,包括预测变量对之间的相互作用,基于平均绝对分数的变量重要性估计,并为新预测提供每个预测变量的分数,增加了可解释性,但需要额外的工作来量化这些输出如何受到变量相关性,交互项的包含和大特征空间的影响。EBM的进一步探索尤其值得用于地质灾害制图和建模,以及一般的空间预测制图和建模,特别是当结果预测的价值或使用将通过改善全球可解释性和预测解释的可用性而大大增强时,在绘制或建模范围内的每个细胞或聚集单元。
Machine learning (ML) methods, such as artificial neural networks (ANN), k-nearest neighbors (kNN), random forests (RF), support vector machines (SVM), and boosted decision trees (DTs), may offer stronger predictive performance than more traditional, parametric methods, such as linear regression, multiple linear regression, and logistic regression (LR), for specific mapping and modeling tasks. However, this increased performance is often accompanied by increased model complexity and decreased interpretability, resulting in critiques of their “black box” nature, which highlights the need for algorithms that can offer both strong predictive performance and interpretability. This is especially true when the global model and predictions for specific data points need to be explainable in order for the model to be of use. Explainable boosting machines (EBM), an augmentation and refinement of generalize additive models (GAMs), has been proposed as an empirical modeling method that offers both interpretable results and strong predictive performance. The trained model can be graphically summarized as a set of functions relating each predictor variable to the dependent variable along with heat maps representing interactions between selected pairs of predictor variables. In this study, we assess EBMs for predicting the likelihood or probability of slope failure occurrence based on digital terrain characteristics in four separate Major Land Resource Areas (MLRAs) in the state of West Virginia, USA and compare the results to those obtained with LR, kNN, RF, and SVM. EBM provided predictive accuracies comparable to RF and SVM and better than LR and kNN. The generated functions and visualizations for each predictor variable and included interactions between pairs of predictor variables, estimation of variable importance based on average mean absolute scores, and provided scores for each predictor variable for new predictions add interpretability, but additional work is needed to quantify how these outputs may be impacted by variable correlation, inclusion of interaction terms, and large feature spaces. Further exploration of EBM is merited for geohazard mapping and modeling in particular and spatial predictive mapping and modeling in general, especially when the value or use of the resulting predictions would be greatly enhanced by improved interpretability globally and availability of prediction explanations at each cell or aggregating unit within the mapped or modeled extent.