AIC MODEL SELECTION IN OVERDISPERSED CAPTURE-RECAPTURE DATA

AIC MODEL SELECTION IN OVERDISPERSED CAPTURE-RECAPTURE DATA
复制标题

DOI:
10.2307/1939637
复制
发表时间:
1994-09-01
期刊:
影响因子:
4.8
通讯作者:
WHITE, GC
WHITE, GC
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
ANDERSON, DR;BURNHAM, KP;WHITE, GC

文献摘要

被引文献

相似文献

选择适当的模型作为从捕获-再捕获数据进行统计推断的基础是至关重要的。在分析多个相互关联的数据集(例如,男性和女性,2-3个年龄段,3-5个地区和10-15岁)。为这些数据集考虑的最一般的模型可能包含1000个存活和再捕获参数。本文给出了三种信息理论方法的数值结果,当数据过分散时(即,缺乏独立性,从而发生额外的二项式变化)。Akaike的信息准则(AIC),二阶调整AIC的偏差(AIC(c)),和维度一致性准则(CAIC)进行了修改,使用平均过度分散的经验估计,准似然理论的基础上。基于标准化和theta(参数theta为向量值)之间的欧几里得距离评价模型选择的质量;该量(一种残差平方和,因此表示为RSS)是平方偏差和方差的组合。五个结果似乎是普遍感兴趣的这些产品多项式模型。首先,当存在过度分散时,方差膨胀因子的最直接估计值有正偏差,并且相对偏差随着过度分散的量而增加。其次,AIC和AIC(c),未使用准似然理论调整过度分散,在选择具有小RSS值的模型时表现不佳,当数据过度分散时(即,当与具有最小RRS值的模型相比时,选择过拟合模型)。第三,信息理论的标准,调整过度分散,表现良好,选择parismonious模型,并有一个良好的平衡下和过拟合的数据。第四,一般来说,尺寸一致的标准选择的模型比其他标准的参数少,有较小的RSS值,但显然是错误的欠拟合时,与最小的RSS值的模型相比。第五,即使真实的模型结构(但不是模型中的实际参数值)是已知的,当真实模型包括几个(更不用说许多)与O无显著差异的估计参数时,真实模型在拟合到数据(通过参数估计)时对于统计推断是相对较差的基础。
Selection of a proper model as a basis for statistical inference from capture-recapture data is critical. This is especially so when using open models in the analysis of multiple, interrelated data sets (e.g., males and females, with 2-3 age classes, over 3-5 areas and 10-15 yr). The most general model considered for such data sets might contain 1000 survival and recapture parameters. This paper presents numerical results on three information-theoretic methods for model selection when the data are overdispersed (i.e., a lack of independence so that extra-binomial variation occurs). Akaike's information criterion (AIC), a second-order adjustment to AIC for bias (AIC(c)), and a dimension-consistent criterion (CAIC) were modified using an empirical estimate of the average overdispersion, based on quasi-likelihood theory. Quality of model selection was evaluated based on the Euclidian distance between standardized and theta (parameter theta is vector valued); this quantity (a type of residual sum of squares, hence denoted as RSS) is a combination of squared bias and variance. Five results seem to be of general interest for these product-multinomial models. First, when there was overdispersion the most direct estimator of the variance inflation factor was positively biased and the relative bias increased with the amount of overdispersion. Second, AIC and AIC(c), unadjusted for overdispersion using quasi-likelihood theory, performed poorly in selecting a model with a small RSS value when the data were overdispersed (i.e., overfitted models were selected when compared to the model with the minimum RRS value). Third, the information-theoretic criteria, adjusted for overdispersion, performed well, selected parismonious models, and had a good balance between under- and overfitting the data. Fourth, generally, the dimension-consistent criterion selected models with fewer parameters than the other criteria, had smaller RSS values, but clearly was in error by underfitting when compared with the model with the minimum RSS value. Fifth, even if the true model structure (but not the actual parameter values in the model) is known, that true model, when fitted to the data (by parameter estimation) is a relatively poor basis for statistical inference when that true model includes several, let alone many, estimated parameters that are not significantly different from O.