Is more data always better? A simulation study of benefits and limitations of integrated distribution models

Is more data always better? A simulation study of benefits and limitations of integrated distribution models
复制标题

DOI:
10.1111/ecog.05146
复制
发表时间:
2020-07-14
期刊:
影响因子:
5.9
通讯作者:
O'Hara, Robert B.
O'Hara, Robert B.
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Simmonds, Emily G.;Jarvis, Susan G.;O'Hara, Robert B.

文献摘要

被引文献

相似文献

物种分布模型是一种流行且应用广泛的生态学工具。最近数据可用性的增加为物种分布建模带来了机遇和挑战。每个数据源都有不同的质量,这取决于它的收集方式。由于几个数据源可以提供单一物种的信息,生态学家通常只分析其中一个数据源,但由于一些数据源被丢弃,这丢失了信息。开发了集成分布模型(idm),以便在考虑不同数据收集协议的同时,在单个模型中包含多个数据集。这是有利的,因为它允许有效地使用所有可用的数据,可以改进估计和解释数据收集中的偏差。目前尚不清楚的是,何时集成不同的数据源不会带来优势。在这里,我们首次使用模拟研究来探索idm的潜在局限性,该研究将空间偏差、机会主义、仅存在的数据集与结构化、存在-不存在的数据集相结合。我们基于现实生态问题探讨了四种情景;样本量小,检测概率低,协变量之间的相关性以及缺乏对数据收集中偏倚驱动因素的了解。对于每种情况,我们都会问;与单独对任一数据源建模相比,我们是否看到了IDM在参数估计或空间模式预测准确性方面的改进?我们发现,仅靠整合无法纠正仅存在数据中的空间偏差。包括一个协变量来解释偏差或添加一个灵活的空间项可以提高IDM的性能,超越单一数据集模型,模型包括一个灵活的空间项,产生最准确和稳健的估计。增加存在-缺失数据的样本量和没有相关协变量也可以改善估计。这些结果表明,在哪些条件下,集成模型比单一数据源建模更有优势。
Species distribution models are popular and widely applied ecological tools. Recent increases in data availability have led to opportunities and challenges for species distribution modelling. Each data source has different qualities, determined by how it was collected. As several data sources can inform on a single species, ecologists have often analysed just one of the data sources, but this loses information, as some data sources are discarded. Integrated distribution models (IDMs) were developed to enable inclusion of multiple datasets in a single model, whilst accounting for different data collection protocols. This is advantageous because it allows efficient use of all data available, can improve estimation and account for biases in data collection. What is not yet known is when integrating different data sources does not bring advantages. Here, for the first time, we explore the potential limits of IDMs using a simulation study integrating a spatially biased, opportunistic, presence-only dataset with a structured, presence-absence dataset. We explore four scenarios based on real ecological problems; small sample sizes, low levels of detection probability, correlations between covariates and a lack of knowledge of the drivers of bias in data collection. For each scenario we ask; do we see improvements in parameter estimation or the accuracy of spatial pattern prediction in the IDM versus modelling either data source alone? We found integration alone was unable to correct for spatial bias in presence-only data. Including a covariate to explain bias or adding a flexible spatial term improved IDM performance beyond single dataset models, with the models including a flexible spatial term producing the most accurate and robust estimates. Increasing the sample size of presence-absence data and having no correlated covariates also improved estimation. These results demonstrate under which conditions integrated models provide benefits over modelling single data sources.