Effects of sample size on accuracy of species distribution models

Effects of sample size on accuracy of species distribution models
复制标题

DOI:
10.1016/s0304-3800(01)00388-x
复制
发表时间:
2002-02-01
影响因子:
3.1
通讯作者:
Peterson, AT
Peterson, AT
中科院分区:
环境科学与生态学3区
文献类型:
--
作者:
Stockwell, DRB;Peterson, AT

文献摘要

被引文献

相似文献

随着对大量生物多样性信息的获取,一种强大的能力是模拟生态位和预测地理分布。由于采样物种的分布是昂贵的,我们通过对采样良好的物种的数据进行重新采样,探索了三种预测建模方法的准确建模所需的样本量,并绘制了随着样本量增加的模型改进曲线。一般来说,在粗略的替代模型和机器学习方法下,预测某一地点物种出现的平均成功率或准确度在10个样本点内为最大值的90%,在50个数据点处接近最大值。然而,随着样本量的增加,精细的替代模型和逻辑回归模型的准确性增加率显著降低,在100个数据点处达到相似的最大准确性。环境变量的选择也对逻辑回归方法在样本量范围内的准确性产生了不可预测的影响,而机器学习方法在整个过程中具有稳健的性能。检查跨物种的模型性能的相关性,地理分布的范围是唯一显着的生态因素。(C)2002 Elsevier Science B.V.保留所有权利。
Given increasing access to large amounts of biodiversity information, a powerful capability is that of modeling ecological niches and predicting geographic distributions. Because, sampling species' distributions is costly, we explored sample size needs for accurate modeling for three predictive modeling methods via re-sampling of data for well-sampled species, and developed curves of model improvement with increasing sample size. In general, under a coarse surrogate model, and machine-learning methods, average success rate at predicting occurrence of a species at a location, or accuracy, was 90% of maximum within ten sample points, and was near maximal at 50 data points. However, a fine surrogate model and logistic regression model had significantly lower rates of increase in accuracy with increasing sample size, reaching similar maximum accuracy at 100 data points. The choice of environmental variables also produced unpredictable effects on accuracy over the range of sample sizes on the logistic regression method, while the machine-learning method had robust performance throughout. Examining correlates of model performance across species, extent of geographic distribution was the only significant ecological factor. (C) 2002 Elsevier Science B.V. All rights reserved.