Species-specific tuning increases robustness to sampling bias in models of species distributions: An implementation with Maxent

Species-specific tuning increases robustness to sampling bias in models of species distributions: An implementation with Maxent
复制标题

DOI:
10.1016/j.ecolmodel.2011.04.011
复制
发表时间:
2011-08-10
影响因子:
3.1
通讯作者:
Gonzalez, Israel, Jr.
Gonzalez, Israel, Jr.
中科院分区:
环境科学与生态学3区
文献类型:
--
作者:
Anderson, Robert P.;Gonzalez, Israel, Jr.

文献摘要

被引文献

相似文献

利用记录物种存在的研究区域和发生地点的环境数据(通常来自博物馆和草本植物),存在各种方法来模拟物种生态位和地理分布。在仅存在的模型中,地理抽样偏差和小样本量是许多物种面临的挑战。对这类数据集的偏差和/或噪声特性的过度拟合可能会严重损害模型的通用性和可转移性,这对许多当前的应用--包括入侵物种的研究、气候变化的影响和生态位演变--至关重要。即使在可转移性不是必要的情况下,许多领域的应用,包括保护生物学、宏观生态学和人畜共患疾病,也需要不太适合的模型。我们使用最大熵方法(Maxent)对委内瑞拉Cordillera de Merida特有的隐翅虫(Cryptotis Merdensis)进行了评估。为了模拟强抽样偏差,我们将地点分为两个数据集:来自物种范围中采样努力较高的部分的数据集(用于模型校准)和来自物种范围的其他区域的数据集,其中发生的采样较少(用于模型评估)。在建模之前,我们评估了两个数据集中各地的气候值,以确定是否有任何环境偏差伴随着地理偏差。然后,为了确定模型复杂性的最佳水平(并最大限度地减少过度匹配),我们建立了模型并调整了模型设置,将性能与使用默认设置实现的性能进行了比较。我们随机选择用于模型校准的地点(5、10、15和20个地点的集合),并改变所考虑的模型复杂性水平(线性特征与线性特征和二次特征)以及防止过度拟合的两个方面(正则化)。环境偏差实际上对应于数据集之间的地理偏差,某些变量的中位数和观测范围(最小值和/或最大值)存在差异。模型的性能因正规化程度的不同而有很大差异。中等正则化始终导致最佳模型,在低正则化时性能下降,而在高正则化时通常性能下降。正则化的最优水平在样本量依赖和样本量无关的方法中有所不同,但两者都达到了相似的最大性能水平。在某些情况下,最佳正则化值不同于(通常高于)默认正则化值。同时使用线性和二次特征进行校准的模型表现优于仅使用线性特征的模型。在所检查的样本大小中,结果非常一致。当采用适当的正则化方法并确定最优模型复杂性时,用较少和有偏的位置建立的模型具有较高的预测能力。与使用默认设置相比,特定于物种的模型设置调整有很大的好处。(C)2011爱思唯尔B.V.保留所有权利。
Various methods exist to model a species niche and geographic distribution using environmental data for the study region and occurrence localities documenting the species' presence (typically from museums and herbaria). In presence-only modelling, geographic sampling bias and small sample sizes represent challenges for many species. Overfitting to the bias and/or noise characteristic of such datasets can seriously compromise model generality and transferability, which are critical to many current applications - including studies of invasive species, the effects of climatic change, and niche evolution. Even when transferability is not necessary, applications to many areas, including conservation biology, macroecology, and zoonotic diseases, require models that are not overfit. We evaluated these issues using a maximum entropy approach (Maxent) for the shrew Cryptotis meridensis, which is endemic to the Cordillera de Merida in Venezuela. To simulate strong sampling bias, we divided localities into two datasets: those from a portion of the species' range that has seen high sampling effort (for model calibration) and those from other areas of the species' range, where less sampling has occurred (for model evaluation). Before modelling, we assessed the climatic values of localities in the two datasets to determine whether any environmental bias accompanies the geographic bias. Then, to identify optimal levels of model complexity (and minimize overfitting), we made models and tuned model settings, comparing performance with that achieved using default settings. We randomly selected localities for model calibration (sets of 5, 10, 15, and 20 localities) and varied the level of model complexity considered (linear versus both linear and quadratic features) and two aspects of the strength of protection against overfitting (regularization). Environmental bias indeed corresponded to the geographic bias between datasets, with differences in median and observed range (minima and/or maxima) for some variables. Model performance varied greatly according to the level of regularization. Intermediate regularization consistently led to the best models, with decreased performance at low and generally at high regularization. Optimal levels of regularization differed between sample-size-dependent and sample-size-independent approaches, but both reached similar levels of maximal performance. In several cases, the optimal regularization value was different from (usually higher than) the default one. Models calibrated with both linear and quadratic features outperformed those made with just linear features. Results were remarkably consistent across the examined sample sizes. Models made with few and biased localities achieved high predictive ability when appropriate regularization was employed and optimal model complexity was identified. Species-specific tuning of model settings can have great benefits over the use of default settings. (C) 2011 Elsevier B.V. All rights reserved.