A statistical explanation of MaxEnt for ecologists

A statistical explanation of MaxEnt for ecologists
复制标题

DOI:
10.1111/j.1472-4642.2010.00725.x
复制
发表时间:
2011-01-01
影响因子:
4.6
通讯作者:
Yates, Colin J.
Yates, Colin J.
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Elith, Jane;Phillips, Steven J.;Yates, Colin J.

文献摘要

被引文献

相似文献

MaxEnt是一个用于从仅存在的物种记录中模拟物种分布的程序。这篇论文是为生态学家写的,从统计学的角度描述了MaxEnt模型,在模型的结构、产生模拟分布所需的决策以及有关物种和可能影响这些决策的数据的知识之间建立了明确的联系。开始,我们讨论的特点存在的数据,突出的影响建模分布。我们特别关注样本偏差和缺乏物种流行率信息的问题。本文的重点是MaxEnt的一个新的统计解释,它表明该模型最小化定义在协变量空间中的两个概率密度(一个从存在数据估计,一个从景观)之间的相对熵。对于许多用户来说,这种观点可能是一种比以前依赖机器学习概念的模型更容易理解的方式。然后,我们详细解释了MaxEnt描述的关键组件(例如协变量和特征,以及景观范围的定义),模型拟合的机制(例如特征选择,约束和正则化)和输出。使用的案例研究的Banksia物种原产于澳大利亚西南部和河流鱼类,我们适合模型和解释它们,探索为什么某些选择会影响结果,这意味着什么。鱼的例子说明了使用线性河流段的矢量数据而不是栅格(网格化)数据的模型。适当的处理调查偏差,未预测的数据,当地限制的物种,并预测到训练数据范围之外的环境进行了演示,并讨论了新的功能。在线附录包括模型的其他细节以及以前的解释与本解释之间的数学联系,示例代码和数据,以及案例研究的进一步信息。
MaxEnt is a program for modelling species distributions from presence-only species records. This paper is written for ecologists and describes the MaxEnt model from a statistical perspective, making explicit links between the structure of the model, decisions required in producing a modelled distribution, and knowledge about the species and the data that might affect those decisions. To begin we discuss the characteristics of presence-only data, highlighting implications for modelling distributions. We particularly focus on the problems of sample bias and lack of information on species prevalence. The keystone of the paper is a new statistical explanation of MaxEnt which shows that the model minimizes the relative entropy between two probability densities (one estimated from the presence data and one, from the landscape) defined in covariate space. For many users, this viewpoint is likely to be a more accessible way to understand the model than previous ones that rely on machine learning concepts. We then step through a detailed explanation of MaxEnt describing key components (e.g. covariates and features, and definition of the landscape extent), the mechanics of model fitting (e.g. feature selection, constraints and regularization) and outputs. Using case studies for a Banksia species native to south-west Australia and a riverine fish, we fit models and interpret them, exploring why certain choices affect the result and what this means. The fish example illustrates use of the model with vector data for linear river segments rather than raster (gridded) data. Appropriate treatments for survey bias, unprojected data, locally restricted species, and predicting to environments outside the range of the training data are demonstrated, and new capabilities discussed. Online appendices include additional details of the model and the mathematical links between previous explanations and this one, example code and data, and further information on the case studies.