Evaluating the predictive performance of habitat models developed using logistic regression

Evaluating the predictive performance of habitat models developed using logistic regression
复制标题

DOI:
10.1016/s0304-3800(00)00322-7
复制
发表时间:
2000-09-03
影响因子:
3.1
通讯作者:
Ferrier, S
Ferrier, S
中科院分区:
环境科学与生态学3区
文献类型:
--
作者:
Pearce, J;Ferrier, S

文献摘要

被引文献

相似文献

使用统计模型来预测物种可能出现或分布正在成为保护规划和野生动物管理中越来越重要的工具。使用独立数据评估模型的预测性能是模型开发的重要一步。这种评估有助于确定模型对特定应用的适用性,有助于对竞争模型和建模技术进行比较评估,并确定模型最需要改进的方面。使用逻辑回归开发的栖息地模型的预测性能需要根据两个组成部分进行评估:可靠性或校准(预测发生概率与观察到的占用地点比例之间的一致性)和区分能力(模型正确区分占用和未占用地点的能力)。可靠性的缺乏可归因于两个系统来源:校准偏差和扩散。描述了用于评估这两种误差源的技术。逻辑回归模型的辨别能力通常通过在二乘二的表中对观察值和预测进行交叉分类并计算分类性能指数来衡量。然而,这种方法基本上依赖于阈值概率的任意选择来确定站点是否被预测被占用。描述了一种替代方法,该方法根据相对操作特征(ROC)曲线下的面积来测量辨别能力,该曲线与宽且连续的阈值水平范围内正确和错误分类的预测的相对比例相关。本文推广的技术的更广泛应用可以极大地提高对为保护规划和野生动物管理而开发的栖息地模型的有用性和潜在局限性的理解。 (C) 2000 Elsevier Science B.V. 保留所有权利。
The use of statistical models to predict the likely occurrence or distribution of species is becoming an increasingly important tool in conservation planning and wildlife management. Evaluating the predictive performance of models using independent data is a vital step in model development. Such evaluation assists in determining the suitability of a model for specific applications, facilitates comparative assessment of competing models and modelling techniques, and identities aspects of a model most in need of improvement. The predictive performance of habitat models developed using logistic regression needs to be evaluated in terms of two components: reliability or calibration (the agreement between predicted probabilities of occurrence and observed proportions of sites occupied), and discrimination capacity (the ability of a model to correctly distinguish between occupied and unoccupied sites). Lack of reliability can be attributed to two systematic sources, calibration bias and spread. Techniques are described for evaluating both of these sources of error. The discrimination capacity of logistic regression models is often measured by cross-classifying observations and predictions in a two-by-two table, and calculating indices of classification performance. However, this approach relies on the essentially arbitrary choice of a threshold probability to determine whether or not a site is predicted to be occupied. An alternative approach is described which measures discrimination capacity in terms of the area under a relative operating characteristic (ROC) curve relating relative proportions of correctly and incorrectly classified predictions over a wide and continuous range of threshold levels. Wider application of the techniques promoted in this paper could greatly improve understanding of the usefulness, and potential limitations, of habitat models developed for use in conservation planning and wildlife management. (C) 2000 Elsevier Science B.V. All rights reserved.