Bayesian Simultaneous Edit and Imputation for Multivariate Categorical Data

Bayesian Simultaneous Edit and Imputation for Multivariate Categorical Data
复制标题

DOI:
10.1080/01621459.2016.1231612
复制
发表时间:
2017-01-01
影响因子:
3.7
通讯作者:
Reiter, Jerome P.
Reiter, Jerome P.
中科院分区:
数学1区
文献类型:
--
作者:
Manrique-Vallier, Daniel;Reiter, Jerome P.

文献摘要

被引文献

相似文献

在分类数据中,典型的情况是某些变量的组合在理论上是不可能的,例如已婚的3岁孩子或怀孕的男人。然而,在实践中,由于例如答复错误或数据处理错误,报告的值通常包括这样的结构零。为了清除这种错误,许多统计组织使用一种称为编辑-归罪的过程。其基本思想是首先根据某个启发式或损失函数选择要更改的报告值,然后用合理的推算来替换这些值。在确定错误的位置时,这两个阶段的过程通常不会充分利用数据中的信息,也不会适当地反映编辑和推算所产生的不确定性。我们提出了另一种方法来编辑和补充具有结构零的分类微数据,以解决这些缺点。具体地说,我们使用贝叶斯分层模型,该模型将测量误差过程的随机模型与基础无误差值的多项式分布的狄利克雷过程混合。后一种模型被限制为仅在理论上可能的组合集合上具有支持。我们使用2000年美国人口普查数据的模拟研究来说明这种综合的编辑和填充方法,并将其与两阶段编辑-填充例程进行比较。补充材料可以在网上找到。
In categorical data, it is typically the case that some combinations of variables are theoretically impossible, such as a 3-year-old child who is married or a man who is pregnant. In practice, however, reported values often include such structural zeros due to, for example, respondent mistakes or data processing errors. To purge data of such errors, many statistical organizations use a process known as edit-imputation. The basic idea is first to select reported values to change according to some heuristic or loss function, and second to replace those values with plausible imputations. This two-stage process typically does not fully use information in the data when determining locations of errors, nor does it appropriately reflect uncertainty resulting from the edits and imputations. We present an alternative approach to editing and imputation for categorical microdata with structural zeros that addresses these shortcomings. Specifically, we use a Bayesian hierarchical model that couples a stochastic model for the measurement error process with a Dirichlet process mixture of multinomial distributions for the underlying, error-free values. The latter model is restricted to have support only on the set of theoretically possible combinations. We illustrate this integrated approach to editing and imputation using simulation studies with data from the 2000 U.S. census, and compare it to a two-stage edit-imputation routine. Supplementary material is available online.