Nonparametric Bayesian Multiple Imputation for Incomplete Categorical Variables in Large-Scale Assessment Surveys

Nonparametric Bayesian Multiple Imputation for Incomplete Categorical Variables in Large-Scale Assessment Surveys
复制标题

DOI:
10.3102/1076998613480394
复制
发表时间:
2013-10-01
影响因子:
2.4
通讯作者:
Reiter, Jerome P.
Reiter, Jerome P.
中科院分区:
心理学4区
文献类型:
--
作者:
Si, Yajuan;Reiter, Jerome P.

文献摘要

被引文献

相似文献

在许多调查中,数据包括大量的分类变量,这些变量受到项目无应答的影响。多重插补的标准方法,如对数线性模型或序贯回归插补,可能无法捕获复杂的依赖关系,并且难以在高维中有效实施。我们提出了一个完全贝叶斯,联合建模方法的多项分布的Dirichlet过程混合物的基础上的分类数据的多重插补。该方法自动建模复杂的依赖关系,同时计算方便。Dirichlet过程先验分布使分析人员能够避免将混合组分的数量固定在任意数量。我们说明了重复采样的方法使用模拟数据的属性。我们应用该方法对2007年国际数学和科学趋势研究中缺失的背景数据进行了估算。
In many surveys, the data comprise a large number of categorical variables that suffer from item nonresponse. Standard methods for multiple imputation, like log-linear models or sequential regression imputation, can fail to capture complex dependencies and can be difficult to implement effectively in high dimensions. We present a fully Bayesian, joint modeling approach to multiple imputation for categorical data based on Dirichlet process mixtures of multinomial distributions. The approach automatically models complex dependencies while being computationally expedient. The Dirichlet process prior distributions enable analysts to avoid fixing the number of mixture components at an arbitrary number. We illustrate repeated sampling properties of the approach using simulated data. We apply the methodology to impute missing background data in the 2007 Trends in International Mathematics and Science Study.