Modeling of species distributions with Maxent:: new extensions and a comprehensive evaluation

Modeling of species distributions with Maxent:: new extensions and a comprehensive evaluation
复制标题

DOI:
10.1111/j.0906-7590.2008.5203.x
复制
发表时间:
2008-04-01
期刊:
影响因子:
5.9
通讯作者:
Dudik, Miroslav
Dudik, Miroslav
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Phillips, Steven J.;Dudik, Miroslav

文献摘要

被引文献

相似文献

物种地理分布的精确建模对于生态学和保护学的各种应用至关重要。性能最好的技术通常需要一些参数调整,这可能是非常耗时的,为每个物种单独做,或不可靠的小或有偏见的数据集。此外,即使有大量高质量的数据,对物种模型应用感兴趣的用户也不需要具备详细调优所需的统计知识。在这种情况下,最好使用“默认设置”,在不同的数据集上进行调整和验证。Maxent是最近推出的建模技术,实现了高预测精度,并享有几个额外的有吸引力的属性。Maxent的性能受到中等数量的参数的影响。本文的第一个贡献是这些参数的经验调整。由于许多数据集缺乏有关物种缺席的信息,我们提出了一种使用仅存在数据的调优方法。我们对独立收集的高质量存在-不存在数据评估我们的方法。除了调优之外,我们还介绍了几个概念,这些概念可以提高Maxent的预测准确性和运行时间。我们引入了“铰链功能”,模型更复杂的关系,在训练数据中,我们描述了一个新的逻辑输出格式,给出了估计的存在概率,最后,我们探讨了“背景抽样”的策略,科普样本选择偏差,减少模型的建立时间。我们的评估,基于来自6个地区的226个物种的不同数据集,显示:1)默认设置调整的存在唯一的数据实现的性能几乎一样好,如果他们已经调整了评估数据本身; 2)铰链功能大大提高模型的性能; 3)逻辑输出改善了模型校准,使得输出值的大差异更好地对应于适合性的大差异; 4)“目标组”背景采样比随机背景采样具有更好的预测性能; 5)随机背景采样在不降低模型性能的情况下显著减少了运行时间。
Accurate modeling of geographic distributions of species is crucial to various applications in ecology and conservation. The best performing techniques often require some parameter tuning, which may be prohibitively time-consuming to do separately for each species, or unreliable for small or biased datasets. Additionally, even with the abundance of good quality data, users interested in the application of species models need not have the statistical knowledge required for detailed tuning. In such cases, it is desirable to use "default settings", tuned and validated on diverse datasets. Maxent is a recently introduced modeling technique, achieving high predictive accuracy and enjoying several additional attractive properties. The performance of Maxent is influenced by a moderate number of parameters. The first contribution of this paper is the empirical tuning of these parameters. Since many datasets lack information about species absence, we present a tuning method that uses presence-only data. We evaluate our method on independently collected high-quality presence-absence data. In addition to tuning, we introduce several concepts that improve the predictive accuracy and running time of Maxent. We introduce "hinge features" that model more complex relationships in the training data; we describe a new logistic output format that gives an estimate of probability of presence; finally we explore "background sampling" strategies that cope with sample selection bias and decrease model-building time. Our evaluation, based on a diverse dataset of 226 species from 6 regions, shows: 1) default settings tuned on presence-only data achieve performance which is almost as good as if they had been tuned on the evaluation data itself; 2) hinge features substantially improve model performance; 3) logistic output improves model calibration, so that large differences in output values correspond better to large differences in suitability; 4) "target-group" background sampling can give much better predictive performance than random background sampling; 5) random background sampling results in a dramatic decrease in running time, with no decrease in model performance.