Selecting pseudo-absences for species distribution models: how, where and how many?

Selecting pseudo-absences for species distribution models: how, where and how many?
复制标题

DOI:
10.1111/j.2041-210x.2011.00172.x
复制
发表时间:
2012-04-01
影响因子:
6.6
通讯作者:
Thuiller, Wilfried
Thuiller, Wilfried
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Barbet-Massin, Morgane;Jiguet, Frederic;Thuiller, Wilfried

文献摘要

被引文献

相似文献

1.物种分布模型越来越多地用于解决保护生物学,生态学和进化中的问题。最有效的物种分布模型需要有关该地区物种存在和现有环境条件的数据(称为背景或伪缺失数据)。然而,对于如何以及在何处对这些假缺席进行采样以及采样数量,仍然没有达成共识。2.在这项研究中,我们进行了一个全面的比较分析的基础上,简单的模拟物种分布,提出指导方针,如何,在哪里和多少伪缺席应产生建立可靠的物种分布模型。根据初始存在数据的数量和质量(无偏与气候或空间偏差),我们评估了选择伪缺席(随机与环境或空间分层)的方法及其数量对七种常见建模技术(回归,分类和机器学习技术)的预测准确性的相对影响。3.当使用回归技术时,用于选择伪缺席的方法对模型的预测准确性影响最大。随机选择的假缺席产生了最可靠的分布模型。模型拟合了大量的伪缺席,但同样加权的存在(即。e.存在的加权和等于伪不存在的加权和)产生最准确的预测分布。对于分类和机器学习技术,伪缺失的数量对模型准确性的影响最大,并且对具有比回归技术更少的伪缺失的几次运行进行平均,产生了最具预测性的模型。4.总的来说,我们建议使用一个大的数字(e。G. 10 000)的伪缺席与平等的权重存在和缺席时,使用回归技术(e。G.广义线性模型和广义加性模型);对几次运行进行平均(例如。G. 10)具有较少的伪缺席(e. G. 100)使用多重自适应回归样条和判别分析对存在和不存在进行相等的加权;以及使用与可用存在相同数量的伪不存在(如果很少伪不存在,则对几次运行进行平均)用于分类技术,诸如增强回归树、分类树和随机森林。此外,我们建议在使用回归技术时随机选择伪缺席,在使用分类和机器学习技术时随机选择地理和环境分层的伪缺席。
1. Species distribution models are increasingly used to address questions in conservation biology, ecology and evolution. The most effective species distribution models require data on both species presence and the available environmental conditions (known as background or pseudo-absence data) in the area. However, there is still no consensus on how and where to sample these pseudo-absences and how many. 2. In this study, we conducted a comprehensive comparative analysis based on simple simulated species distributions to propose guidelines on how, where and how many pseudo-absences should be generated to build reliable species distribution models. Depending on the quantity and quality of the initial presence data (unbiased vs. climatically or spatially biased), we assessed the relative effect of the method for selecting pseudo-absences (random vs. environmentally or spatially stratified) and their number on the predictive accuracy of seven common modelling techniques (regression, classification and machine-learning techniques). 3. When using regression techniques, the method used to select pseudo-absences had the greatest impact on the model's predictive accuracy. Randomly selected pseudo-absences yielded the most reliable distribution models. Models fitted with a large number of pseudo-absences but equally weighted to the presences (i. e. the weighted sum of presence equals the weighted sum of pseudoabsence) produced the most accurate predicted distributions. For classification and machine-learning techniques, the number of pseudo-absences had the greatest impact on model accuracy, and averaging several runs with fewer pseudo-absences than for regression techniques yielded the most predictive models. 4. Overall, we recommend the use of a large number (e. g. 10 000) of pseudo-absences with equal weighting for presences and absences when using regression techniques (e. g. generalised linear model and generalised additive model); averaging several runs (e. g. 10) with fewer pseudo-absences (e. g. 100) with equal weighting for presences and absences with multiple adaptive regression splines and discriminant analyses; and using the same number of pseudo-absences as available presences (averaging several runs if few pseudo-absences) for classification techniques such as boosted regression trees, classification trees and random forest. In addition, we recommend the random selection of pseudo-absences when using regression techniques and the random selection of geographically and environmentally stratified pseudo-absences when using classification and machine-learning techniques.