STRONG LIMIT THEOREMS ON MODEL SELECTION IN GENERALIZED LINEAR REGRESSION WITH BINOMIAL RESPONSES
STRONG LIMIT THEOREMS ON MODEL SELECTION IN GENERALIZED LINEAR REGRESSION WITH BINOMIAL RESPONSES
复制标题
二项式响应广义线性回归模型选择的强极限定理
作者:
G. Qian;Yuehua Wu
We prove a law of iterated logarithm for the maximum likelihood es- timator of the parameters in a generalized linear regression model with binomial response. This result is then used to derive an asymptotic bound for the dierence between the maximum log-likelihood function and the true log-likelihood. It is further used to establish the strong consistency of some penalized likelihood based model selection criteria. We have shown that, under some general conditions, a model selection criterion will select the simplest correct model almost surely if the penalty term is an increasing function of the model dimension and has an order between O(log log n) and O(n). Cases involving the commonly used link functions are discussed for illustration of the results. An important task in linear regression is to identify an optimal subset of available explanatory variables to form a model for best predicting the response variable. We refer to George (2002) and Rao and Wu (2001) for a detailed survey in this area of research. Among the many model selection methods, the classical ones like AIC and BIC are still widely used in practice. It is therefore of interest to investigate the asymptotic properties of model selection criteria which have not yet been established for many problems. In this paper, we focus on variable selection in generalized linear models with binomial responses. We consider a set of model selection criteria, such as AIC, BIC, Cp and the stochastic complexity criterion, that follow the form of a penalized log-likelihood. We assume that all the explanatory variables af- fecting the response variable are available in observations, so that selecting the simplest correct model is possible. We establish a strong representation for the maximum log-likelihood function relative to the true log-likelihood under some general conditions. Based on this representation we show that, when the sample size n is sucien tly large, the simplest correct model is selected almost surely if