Simple Formula for Calculating Bias-Corrected AIC in Generalized Linear Models

Simple Formula for Calculating Bias-Corrected AIC in Generalized Linear Models
复制标题

计算广义线性模型中偏差校正 AIC 的简单公式

DOI:
10.1111/sjos.12049
复制
发表时间:
2014
影响因子:
1
通讯作者:
H.
H.
中科院分区:
数学4区
文献类型:
--
作者:
Imori;S.;Yanagihara;H. & Wa kaki;H.

文献摘要

相似文献

在实际数据分析中,从一组候选模型中确定最佳模型是一个重要问题。有很多文献从不同的角度考虑了此类模型选择问题。例如,通常考虑回归模型中解释变量的子集选择以预测未来数据。模型选择方法通常通过基于预期 Kullback-Leibler (KL) 信息的风险函数来衡量模型对未来数据的拟合优度 (Kullback & Leibler, 1951)。对于实际使用,我们必须估计风险函数,该函数取决于未知参数。最著名的风险函数估计是 Akaike (1973, 1974) 提出的 Akaike 信息准则 (AIC)。由于AIC可以简单地定义为− 2ד最大对数似然”+ 2ד参数数量”,因此AIC广泛应用于化学计量学、工程学、计量经济学、心理计量学和许多其他领域,用于使用一组解释变量选择合适的模型(有关统计模型选择的详细信息,请参见例如,Konishi,1999;Burnham & Anderson,2002;Konishi &北川,2008)。候选模型中具有最小AIC的模型被认为是最佳模型。另外,AIC对风险函数的偏差阶数为O(n−1),这隐含地表明,当样本量n不太大时,AIC有时对风险函数具有不可忽略的偏差。 AIC 往往会低估风险函数,并且 AIC 的偏差容易随着模型中参数数量的增加而增加。潜在地,AIC 倾向于选择比真实模型具有更多参数的模型作为最佳模型 Shibata (1980)。结合这些特征,偏差将导致一个缺点,即
In real data analysis, deciding the best model among a set of candidate models is an important problem. There have been a lot of literature to consider such model selection problems from the various standpoints. For example, a subset selection of explanatory variables in regression models in order to predict the future data is often considered. It is common for a model selection method to measure the goodness of fit of the model for the future data by the risk function based on the expected Kullback-Leibler (KL) information (Kullback & Leibler, 1951). For actual use, we must estimate the risk function, which depends on unknown parameters. The most famous estimator of the risk function is Akaike’s information criterion (AIC) proposed by Akaike (1973, 1974). Since the AIC can be simply defined as− 2דthe maximum log-likelihood”+ 2דthe number of parameters”, the AIC is widely applied in chemometrics, engineering, econometrics, psychometrics, and many other fields for selecting appropriate models using a set of explanatory variables (for details of statistical model selection, see eg., Konishi, 1999; Burnham & Anderson, 2002; Konishi & Kitagawa, 2008). The model having the smallest AIC among the candidate models is regarded as the best model. In addition, the order of the bias of the AIC to the risk function is O (n− 1), which indicates implicitly that the AIC sometimes has a nonnegligible bias to the risk function when the sample size n is not so large. The AIC tends to underestimate the risk function and the bias of AIC is apt to increase with the number of parameters in the model. Potentially, the AIC has a tendency to choose the model that has more parameters than the true model as the best model Shibata (1980). Combined with these characteristics, the bias will cause a disadvantage whereby the