Multiple Imputation for Multilevel Data with Continuous and Binary Variables

Multiple Imputation for Multilevel Data with Continuous and Binary Variables
复制标题

DOI:
10.1214/18-sts646
复制
发表时间:
2018-05-01
影响因子:
5.7
通讯作者:
Resche-Rigon, Matthieu
Resche-Rigon, Matthieu
中科院分区:
数学2区
文献类型:
--
作者:
Audigier, Vincent;White, Ian R.;Resche-Rigon, Matthieu

文献摘要

被引文献

相似文献

我们提出并比较了多水平连续数据和二元数据的多种插值方法,其中变量系统地和零星地缺失。这些方法从理论角度进行了比较,并通过由多个研究组成的真实数据集驱动的广泛模拟研究。比较结果表明,这些多重插值方法最适合处理多级设置中的缺失值,以及为什么它们的相对性能会根据缺失数据模式、多级结构和缺失变量类型而变化。该研究表明,只有当数据集包含大量聚类时,才能获得有效的推断。此外,它还强调了异方差多重插值方法比均方差方法提供了更准确的推断,而均方差方法应保留在每个聚类中个体较少的数据中。最后给出了根据数据结构选择最合适的多重插值方法的指导原则。
We present and compare multiple imputation methods for multilevel continuous and binary data where variables are systematically and sporadically missing. The methods are compared from a theoretical point of view and through an extensive simulation study motivated by a real dataset comprising multiple studies. The comparisons show that these multiple imputation methods are the most appropriate to handle missing values in a multilevel setting and why their relative performances can vary according to the missing data pattern, the multilevel structure and the type of missing variables. This study shows that valid inferences can only be obtained if the dataset includes a large number of clusters. In addition, it highlights that heteroscedastic multiple imputation methods provide more accurate inferences than homoscedastic methods, which should be reserved for data with few individuals per cluster. Finally, guidelines are given to choose the most suitable multiple imputation method according to the structure of the data.