Landslide susceptibility estimation by random forests technique: sensitivity and scaling issues

Landslide susceptibility estimation by random forests technique: sensitivity and scaling issues
复制标题

DOI:
10.5194/nhess-13-2815-2013
复制
发表时间:
2013-01-01
影响因子:
4.6
通讯作者:
Tofani, V.
Tofani, V.
中科院分区:
地球科学3区
文献类型:
--
作者:
Catani, F.;Lagomarsino, D.;Tofani, V.

文献摘要

被引文献

相似文献

尽管最近在滑坡敏感性制图(LSM)方面取得了大量进展和发展,但仍然缺乏针对LSM模型敏感性特定方面的研究。例如,滑坡条件变量(LCVs)的调查规模,映射单元(MUR)的分辨率和LCVs的最佳数量和排名等因素的影响从来没有被分析研究,特别是在大数据集,在本文中,我们尝试这种实验集中在模型调整选择对最终结果的影响,而不是方法的比较。为此,我们采用了一个简单的实现随机森林(RF),机器学习技术,产生一组不同的模型设置,输入数据类型和规模的滑坡易感性地图的合奏。随机森林是贝叶斯树的组合,它将一组预测因子与实际滑坡发生联系起来。作为一个非参数模型,它可以包含一系列数值或分类数据层,并且不需要选择单峰训练数据,例如在线性判别分析中。许多公认的滑坡诱发因素主要与岩性、土地利用、地貌、构造和人为因素有关。此外,对于每个因素,我们还包括在预测集的标准偏差(数值变量)或品种(分类的)在地图unit.As在其他系统中,使用RF使人们能够估计的相对重要性的单一输入参数,并选择最佳配置的分类模型。该模型最初使用完整的输入变量集应用,然后实施迭代过程,并考虑参数空间的逐渐较小的子集。的规模和输入变量的准确性,以及对敏感性结果的RF模型的随机组件的效果的影响,也检查。该模型在阿尔诺河流域(意大利中部)进行了测试。我们发现,参数空间的维度,映射单元(规模)和训练过程强烈影响分类精度和预测过程,这反过来又意味着,在产生所有级别和规模的最终磁化率图之前,应始终使用传统和新的工具进行仔细的敏感性分析。
Despite the large number of recent advances and developments in landslide susceptibility mapping (LSM) there is still a lack of studies focusing on specific aspects of LSM model sensitivity. For example, the influence of factors such as the survey scale of the landslide conditioning variables (LCVs), the resolution of the mapping unit (MUR) and the optimal number and ranking of LCVs have never been investigated analytically, especially on large data sets.In this paper we attempt this experimentation concentrating on the impact of model tuning choice on the final result, rather than on the comparison of methodologies. To this end, we adopt a simple implementation of the random forest (RF), a machine learning technique, to produce an ensemble of landslide susceptibility maps for a set of different model settings, input data types and scales. Random forest is a combination of Bayesian trees that relates a set of predictors to the actual landslide occurrence. Being it a nonparametric model, it is possible to incorporate a range of numerical or categorical data layers and there is no need to select unimodal training data as for example in linear discriminant analysis. Many widely acknowledged landslide predisposing factors are taken into account as mainly related to the lithology, the land use, the geomorphology, the structural and anthropogenic constraints. In addition, for each factor we also include in the predictors set a measure of the standard deviation (for numerical variables) or the variety (for categorical ones) over the map unit.As in other systems, the use of RF enables one to estimate the relative importance of the single input parameters and to select the optimal configuration of the classification model. The model is initially applied using the complete set of input variables, then an iterative process is implemented and progressively smaller subsets of the parameter space are considered. The impact of scale and accuracy of input variables, as well as the effect of the random component of the RF model on the susceptibility results, are also examined. The model is tested in the Arno River basin (central Italy). We find that the dimension of parameter space, the mapping unit (scale) and the training process strongly influence the classification accuracy and the prediction process.This, in turn, implies that a careful sensitivity analysis making use of traditional and new tools should always be performed before producing final susceptibility maps at all levels and scales.