DOUBLY ROBUST NONPARAMETRIC MULTIPLE IMPUTATION FOR IGNORABLE MISSING DATA

DOUBLY ROBUST NONPARAMETRIC MULTIPLE IMPUTATION FOR IGNORABLE MISSING DATA
复制标题

DOI:
10.5705/ss.2010.069
复制
发表时间:
2012-01-01
期刊:
影响因子:
1.4
通讯作者:
Li, Yisheng
Li, Yisheng
中科院分区:
数学3区
文献类型:
--
作者:
Long, Qi;Hsu, Chiu-Hsieh;Li, Yisheng

文献摘要

被引文献

相似文献

缺失数据在医学和社会科学研究中很常见,在数据分析中往往构成严重的挑战。多重补偿方法是处理缺失数据的流行而自然的工具,用一组代表潜在值不确定性的似是而非的值来替换每个缺失值。我们考虑随机缺失(MAR)的情形,并研究当有一组完全观察到的协变量时,在存在缺失值的情况下结果变量的边际均值的估计。我们提出了一种新的非参数多重填充(MI)方法,它使用两个工作模型来实现降维,并定义了缺失观测值的填充集合。与现有的非参数插补方法相比,我们的方法可以更好地处理高维协变量,并且在两个工作模型中的任何一个被正确指定的情况下,所得到的估计量保持一致,因此具有双重稳健性。与现有的双稳健方法相比,我们的非参数MI方法对两个工作模型的错误指定具有更强的鲁棒性;它还避免了使用逆加权,因此对接近1的丢失概率不那么敏感。我们提出了一种灵敏度分析来评估工作模型的有效性,允许调查人员选择最优权重,从而得到的估计器完全或更多地依赖于可能被正确指定的工作模型,从而获得更高的效率。我们研究了所提出的估计量的渐近性质,并进行了仿真研究,结果表明,在有限样本下,所提出的方法优于现有的一些方法。使用结直肠腺瘤研究的数据进一步说明了所提出的方法。
Missing data are common in medical and social science studies and often pose a serious challenge in data analysis. Multiple imputation methods are popular and natural tools for handling missing data, replacing each missing value with a set of plausible values that represent the uncertainty about the underlying values. We consider a case of missing at random (MAR) and investigate the estimation of the marginal mean of an outcome variable in the presence of missing values when a set of fully observed covariates is available. We propose a new nonparametric multiple imputation (MI) approach that uses two working models to achieve dimension reduction and define the imputing sets for the missing observations. Compared with existing nonparametric imputation procedures, our approach can better handle covariates of high dimension, and is doubly robust in the sense that the resulting estimator remains consistent if either of the working models is correctly specified. Compared with existing doubly robust methods, our nonparametric MI approach is more robust to the misspecification of both working models; it also avoids the use of inverse-weighting and hence is less sensitive to missing probabilities that are close to 1. We propose a sensitivity analysis for evaluating the validity of the working models, allowing investigators to choose the optimal weights so that the resulting estimator relies either completely or more heavily on the working model that is likely to be correctly specified and achieves improved efficiency. We investigate the asymptotic properties of the proposed estimator, and perform simulation studies to show that the proposed method compares favorably with some existing methods in finite samples. The proposed method is further illustrated using data from a colorectal adenoma study.