Sample selection bias and presence-only distribution models: implications for background and pseudo-absence data

Sample selection bias and presence-only distribution models: implications for background and pseudo-absence data
复制标题

DOI:
10.1890/07-2153.1
复制
发表时间:
2009-01-01
影响因子:
5
通讯作者:
Ferrier, Simon
Ferrier, Simon
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Phillips, Steven J.;Dudik, Miroslav;Ferrier, Simon

文献摘要

被引文献

相似文献

大多数根据发生记录对物种分布进行建模的方法都需要表示建模区域环境条件范围的额外数据。这些数据称为背景或伪缺失数据,通常是从整个区域随机抽取的,而事件收集通常在空间上偏向于易于访问的区域。由于空间偏差通常会导致环境偏差,因此事件收集和背景采样之间的差异可能会导致模型不准确。为了纠正估计,我们建议选择与出现数据具有相同偏差的背景数据。我们研究了这种方法的理论和实践意义。通常缺乏有关空间偏差的准确信息,因此可能无法对背景站点进行显式偏差采样。然而,通过类似方法观察到的整个目标物种组很可能会出现类似的偏差。因此,我们探索使用目标组内的所有事件作为有偏见的背景数据。我们使用目标群体背景和随机抽样背景对来自世界不同地区的 226 个物种的全面数据集合的模型性能进行比较。我们发现目标群体背景提高了我们考虑的所有建模方法的平均性能,背景数据的选择对预测性能的影响与建模方法的选择一样大。当目标群体存在记录存在强烈偏差时,目标群体背景带来的绩效改进最为显着。我们的方法适用于基于回归的建模方法,这些方法适用于发生数据,例如广义线性或加性模型和增强回归树,以及 Maxent(一种概率密度估计方法)。我们认为,提高对调查中空间偏差影响的认识以及可能的建模补救措施,将大大改善对物种分布的预测。
Most methods for modeling species distributions from occurrence records require additional data representing the range of environmental conditions in the modeled region. These data, called background or pseudo-absence data, are usually drawn at random from the entire region, whereas occurrence collection is often spatially biased toward easily accessed areas. Since the spatial bias generally results in environmental bias, the difference between occurrence collection and background sampling may lead to inaccurate models. To correct the estimation, we propose choosing background data with the same bias as occurrence data. We investigate theoretical and practical implications of this approach. Accurate information about spatial bias is usually lacking, so explicit biased sampling of background sites may not be possible. However, it is likely that an entire target group of species observed by similar methods will share similar bias. We therefore explore the use of all occurrences within a target group as biased background data. We compare model performance using target-group background and randomly sampled background on a comprehensive collection of data for 226 species from diverse regions of the world. We find that target-group background improves average performance for all the modeling methods we consider, with the choice of background data having as large an effect on predictive performance as the choice of modeling method. The performance improvement due to target-group background is greatest when there is strong bias in the target-group presence records. Our approach applies to regression-based modeling methods that have been adapted for use with occurrence data, such as generalized linear or additive models and boosted regression trees, and to Maxent, a probability density estimation method. We argue that increased awareness of the implications of spatial bias in surveys, and possible modeling remedies, will substantially improve predictions of species distributions.