Presence-only data and the em algorithm.

Presence-only data and the em algorithm.
复制标题

DOI:
10.1111/j.1541-0420.2008.01116.x
复制
发表时间:
2009-06
期刊:
影响因子:
1.9
通讯作者:
Leathwick JR
Leathwick JR
中科院分区:
数学3区
文献类型:
--
作者:
Ward G;Hastie T;Barry S;Elith J;Leathwick JR

文献摘要

被引文献

相似文献

在物种栖息地的生态建模中,确定物种缺失的成本可能会高得令人望而却步。仅存在的数据由观察到存在的位置样本和从整个景观中采样的未知存在的单独一组位置组成。我们提出了一种期望最大化算法来估计仅存在数据的潜在存在-不存在逻辑模型。该算法可用于任何现成的逻辑模型。对于具有逐步拟合过程的模型,例如增强树,可以通过在过程中交叉期望步骤来加速拟合过程。基于对新西兰河流中鱼类存在-缺失记录取样的初步分析表明,这种新程序可以减少在实践中经常使用的朴素模型中出现的偏差和边际效应估计的收缩。最后,研究表明,只有当逻辑模型的结构存在一些不切实际的约束时,物种的种群流行率才能被识别。实际上,强烈建议提供人口流行率的估计数。
In ecological modeling of the habitat of a species, it can be prohibitively expensive to determine species absence. Presence-only data consist of a sample of locations with observed presences and a separate group of locations sampled from the full landscape, with unknown presences. We propose an expectation–maximization algorithm to estimate the underlying presence–absence logistic model for presence-only data. This algorithm can be used with any off-the-shelf logistic model. For models with stepwise fitting procedures, such as boosted trees, the fitting process can be accelerated by interleaving expectation steps within the procedure. Preliminary analyses based on sampling from presence–absence records of fish in New Zealand rivers illustrate that this new procedure can reduce both deviance and the shrinkage of marginal effect estimates that occur in the naive model often used in practice. Finally, it is shown that the population prevalence of a species is only identifiable when there is some unrealistic constraint on the structure of the logistic model. In practice, it is strongly recommended that an estimate of population prevalence be provided.