Multiple imputation using multivariate gh transformations

Multiple imputation using multivariate gh transformations
复制标题

DOI:
10.1080/02664763.2012.702268
复制
发表时间:
2012-01-01
影响因子:
1.5
通讯作者:
Raghunathan, Trivellore E.
Raghunathan, Trivellore E.
中科院分区:
数学4区
文献类型:
--
作者:
He, Yulei;Raghunathan, Trivellore E.

文献摘要

被引文献

相似文献

多重插补已成为处理缺失值数据集的一种流行方法。对于不完全连续变量,通常使用多变量正态模型进行插补。然而,这种方法对于具有强非正态形状的变量可能存在问题,因为它会产生与实际分布不一致的估算,从而导致不正确的推断。对于非正态数据,我们考虑Tukey gh分布/变换的多变量扩展[38],以适应偏度和/或峰度并捕获变量之间的相关性。我们提出了一个算法来适应不完整的数据与模型,并产生插补。我们将该方法应用于一个国家的数据集,医院的性能在几个标准的质量措施,这是高度向左倾斜,并大大相互关联。我们使用Monte Carlo研究来评估所提出的方法的性能。我们讨论了可能的推广,并就如何处理非正态不完整数据给从业者一些建议。
Multiple imputation has emerged as a popular approach to handling data sets with missing values. For incomplete continuous variables, imputations are usually produced using multivariate normal models. However, this approach might be problematic for variables with a strong non-normal shape, as it would generate imputations incoherent with actual distributions and thus lead to incorrect inferences. For non-normal data, we consider a multivariate extension of Tukey's gh distribution/transformation [38] to accommodate skewness and/or kurtosis and capture the correlation among the variables. We propose an algorithm to fit the incomplete data with the model and generate imputations. We apply the method to a national data set for hospital performance on several standard quality measures, which are highly skewed to the left and substantially correlated with each other. We use Monte Carlo studies to assess the performance of the proposed approach. We discuss possible generalizations and give some advices to practitioners on how to handle non-normal incomplete data.