Frequentist Model Averaging with missing observations

Frequentist Model Averaging with missing observations
复制标题

DOI:
10.1016/j.csda.2009.07.023
复制
发表时间:
2010-12
期刊:
Comput. Stat. Data Anal.
影响因子:
--
通讯作者:
M. Schomaker;Alan T. K. Wan;C. Heumann
M. Schomaker;Alan T. K. Wan;C. Heumann
中科院分区:
其他
文献类型:
--
作者:
M. Schomaker;Alan T. K. Wan;C. Heumann

文献摘要

被引文献

相似文献

模型平均或组合通常被认为是模型选择的另一种选择。对频率模型平均(FMA)进行了广泛的研究,提出了基于两种不同方法的FMA方法在存在缺失数据情况下的应用策略。第一种方法结合了来自一组适当模型的估计,这些模型通过在最近的模型选择文献中开发的缺失数据调整标准的分数来加权。第二种方法基于传统的模型选择标准,对一组模型的估计值进行平均,但在估计模型之前,用输入值来替换缺失的数据。为此目的,考虑了在当前可用的统计软件中已编程的三种易于使用的推算方法,并且进一步采用简单的递归算法来实现广义回归推算,使得缺失值被连续地预测。后一种算法被发现在给定行观测中同时面临两个或多个缺失值时非常有用。以二元Logistic回归模型为研究对象,通过蒙特卡罗方法研究了这些策略所产生的FMA估计量的性质。结果表明,在许多情况下,补偿后的平均比使用针对缺失数据进行调整的权重进行平均更好,并且模型平均估计器通常比任何单一模型的估计值提供更好的估计。作为说明,所提出的方法被应用于Duchenne肌营养不良检测研究的数据集。
Model averaging or combining is often considered as an alternative to model selection. Frequentist Model Averaging (FMA) is considered extensively and strategies for the application of FMA methods in the presence of missing data based on two distinct approaches are presented. The first approach combines estimates from a set of appropriate models which are weighted by scores of a missing data adjusted criterion developed in the recent literature of model selection. The second approach averages over the estimates of a set of models with weights based on conventional model selection criteria but with the missing data replaced by imputed values prior to estimating the models. For this purpose three easy-to-use imputation methods that have been programmed in currently available statistical software are considered, and a simple recursive algorithm is further adapted to implement a generalized regression imputation in a way such that the missing values are predicted successively. The latter algorithm is found to be quite useful when one is confronted with two or more missing values simultaneously in a given row of observations. Focusing on a binary logistic regression model, the properties of the FMA estimators resulting from these strategies are explored by means of a Monte Carlo study. The results show that in many situations, averaging after imputation is preferred to averaging using weights that adjust for the missing data, and model average estimators often provide better estimates than those resulting from any single model. As an illustration, the proposed methods are applied to a dataset from a study of Duchenne muscular dystrophy detection.