Predictive models of fish species distributions: A note on proper validation and chance predictions

Predictive models of fish species distributions: A note on proper validation and chance predictions
复制标题

DOI:
10.1577/1548-8659(2002)131
复制
发表时间:
2002-03-01
影响因子:
1.4
通讯作者:
Peres-Neto, PR
Peres-Neto, PR
中科院分区:
农林科学3区
文献类型:
--
作者:
Olden, JD;Jackson, DA;Peres-Neto, PR

文献摘要

被引文献

相似文献

物种分布预测是渔业资源研究、保护和管理的主要目标。在这方面,将物种存在或不存在的模式与多尺度生境变量联系起来的统计模型发挥着重要作用。然而,研究人员很少注意到不适当的模型验证和偶然预测如何导致对这些模型的性能和实用性的毫无根据的信心。使用模拟和经验数据的40个湖泊和溪流鱼类,我们证明了常用的再替换方法模型验证(其中相同的数据用于模型的构建和预测)产生高度偏倚的正确分类率的估计,因此真正的模型性能的不准确的看法。相比之下,折刀验证方法导致模型性能的相对无偏估计。模型正确分类的估计率也显示出受到物种流行率的显著影响(即,物种存在的地点的比例),并且通常导致表现不佳的模型被视为强大。我们使用模拟数据来展示如何从模型的机会预测的预期频率是物种流行率和样本大小的函数。最后,我们使用经验数据来说明一种随机化方法,用于评估鱼类栖息地模型的性能是否在统计上大于基于机会预测的预期。总之,我们敦促研究人员采用适当的和可辩护的方法进行模型验证和预测评估;不这样做只会增加渔业文献中已发表的物种栖息地模型的数量,这些模型的使用和可靠性有限。
The prediction of species distributions is a primary goal in the study, conservation, and management of fisheries resources. Statistical models relating patterns of species presence or absence to multiscale habitat variables play an important role in this regard. Researchers, however, have paid little attention to how improper model validation and chance predictions can result in unfounded confidence in the performance and utility of such models. Using simulated and empirical data for 40 lake and stream fish species, we demonstrate that the commonly employed resubstitution approach to model validation (in which the same data are used for both model construction and prediction) produces highly biased estimates of correct classification rates and consequently an inaccurate perception of true model performance. In contrast, a jackknife approach to validation resulted in relatively unbiased estimates of model performance. The estimated rates of model correct classification are also shown to be substantially influenced by species prevalence (i.e., the proportion of sites at which a species is present), and often result in poorly performing models being viewed as powerful. We use simulated data to show how the expected frequency of chance predictions from models is a function of species prevalence and sample size. Finally, we use empirical data to illustrate a randomization approach for assessing whether the performances of the fish habitat models are statistically greater than expectations based on chance predictions. In summary, we urge researchers to employ proper and defensible methodologies for model validation and prediction assessment; failing to do so will only add to the accumulating number of published species habitat models in the fisheries literature that are of limited use and reliability.