Partially parametric techniques for multiple imputation

Partially parametric techniques for multiple imputation
复制标题

DOI:
10.1016/0167-9473(95)00057-7
复制
发表时间:
1996-08-10
影响因子:
1.8
通讯作者:
Taylor, JMG
Taylor, JMG
中科院分区:
数学3区
文献类型:
--
作者:
Schenker, N;Taylor, JMG

文献摘要

被引文献

相似文献

多重填补是一种处理含有缺失值数据集的技术。该方法多次填补缺失值,创建多个完整的数据集用于分析。每个数据集使用为完整数据设计的技术分别进行分析,然后将结果以一种可以纳入填补导致的变异性的方式进行合并。填补缺失值的方法可以从完全参数化到非参数化有所不同。在本文中,我们比较了基于部分参数回归和完全参数回归的多重填补方法。我们所考虑的完全参数方法是通过从回归模型下的预测分布中抽取来填补缺失的回归结果,而部分参数方法是基于使用从完整案例中抽取的值来填补不完全案例的结果或残差。对于部分参数方法,我们提出了一种选择完整案例以从中抽取值的新方法。在回归设定下的蒙特卡洛研究中,我们调查了多重填补方案对数据潜在模型误设的稳健性。所考虑的模型误设来源包括均值结构建模错误以及关于尾部厚重性和异方差性的误差分布设定错误。这些方法在点估计的偏差和效率以及结果的边际均值和分布函数的置信区间覆盖率方面进行了比较。我们发现,当均值结构正确设定时,所有方法都表现良好,即使误差分布被误设。然而,完全参数方法对结果的边际分布函数产生的估计比部分参数方法略为更有效。当均值结构被误设时,所有方法在估计边际均值方面仍然表现良好,尽管完全参数方法在偏差和方差上略有增加。然而,在估计边际分布函数时,完全参数方法在几种情况下失效,而部分参数方法保持其良好性能。在一个类似于但比蒙特卡洛研究稍复杂的艾滋病研究应用中,我们研究了对于具有右删失数据的受试者,从感染艾滋病毒到艾滋病发病的时间分布的估计如何随着用于填补到艾滋病的剩余时间的方法而变化。完全参数和部分参数技术产生了相似的结果,这表明用于完全参数填补的模型选择是合适的。我们的应用提供了一个多重填补如何用于结合信息的例子:从两个队列中结合信息以估计不能从单独的任何一个队列中直接估计的量。
Multiple imputation is a technique for handling data sets with missing values. The method fills in the missing values several times, creating several completed data sets for analysis. Each data set is analyzed separately using techniques designed for complete data, and the results are then combined in such a way that the variability due to imputation may be incorporated. Methods of imputing the missing values can vary from fully parametric to nonparametric. In this paper, we compare partially parametric and fully parametric regression-based multiple-imputation methods. The fully parametric method that we consider imputes missing regression outcomes by drawing them from their predictive distribution under the regression model, whereas the partially parametric methods are based on imputing outcomes or residuals for incomplete cases using values drawn from the complete cases. For the partially parametric methods, we suggest a new approach to choosing complete cases from which to draw values. In a Monte Carlo study in the regression setting, we investigate the robustness of the multiple-imputation schemes to misspecification of the underlying model for the data. Sources of model misspecification considered include incorrect modeling of the mean structure as well as incorrect specification of the error distribution with regard to heaviness of the tails and heteroscedasticity. The methods are compared with respect to the bias and efficiency of point estimates and the coverage rates of confidence intervals for the marginal mean and distribution function of the outcome. We find that when the mean structure is specified correctly, all of the methods perform well, even if the error distribution is misspecified. The fully parametric approach, however, produces slightly more efficient estimates of the marginal distribution function of the outcome than do the partially parametric approaches. When the mean structure is misspecified, all of the methods still perform well for estimating the marginal mean, although the fully parametric method shows slight increases in bias and variance. For estimating the marginal distribution function, however, the fully parametric method breaks down in several situations, whereas the partially parametric methods maintain their good performance. In an application to AIDS research in a setting that is similar to although slightly more complicated than that of the Monte Carlo study, we examine how estimates for the distribution of the time from infection with HIV to the onset of AIDS vary with the method used to impute the residual time to AIDS for subjects with right-censored data. The fully parametric and partially parametric techniques produce similar results, suggesting that the model selection used for fully parametric imputation was adequate. Our application provides an example of how multiple imputation can be used to combine information:from two cohorts to estimate quantities that cannot be estimated directly from either one of the cohorts separately.