Bayesian Latent Class Models for the Multiple Imputation of Categorical Data

Bayesian Latent Class Models for the Multiple Imputation of Categorical Data
复制标题

DOI:
10.1027/1614-2241/a000146
复制
发表时间:
2018-04-01
影响因子:
3.1
通讯作者:
Van Deun, Katrijn
Van Deun, Katrijn
中科院分区:
心理学4区
文献类型:
--
作者:
Vidotto, Davide;Vermunt, Jeroen K.;Van Deun, Katrijn

文献摘要

被引文献

相似文献

潜类分析最近提出了用于缺失分类数据的多重插补(MI),使用标准频率论方法或称为Dirichlet过程混合多项分布(DPMM)的非参数贝叶斯模型。使用潜在类模型进行多重插补的主要优点是,它非常灵活,因为它可以捕获数据中的复杂关系,前提是潜在类的数量足够大。然而,现有的两种方法也具有某些缺点。频率论方法的计算要求很高,因为它需要估计许多LC模型:首先应估计具有不同类别数量的模型以确定所需的类别数量,随后重新估计多个自助样本的选定模型,以考虑插补阶段的参数不确定性。而贝叶斯。Dirichlet过程模型自动进行模型选择和参数不确定性的处理,其缺点是在Gibbs抽样过程中使用的聚类数太少,导致模型欠拟合,产生无效插补。在本文中,我们提出了一种替代方法,结合了现有的两种方法的优势,即我们使用贝叶斯标准的潜在类模型作为插补模型。我们展示了如何在使用单次运行的Gibbs采样器进行插补步骤之前进行模型选择,此外,还展示了如何通过使用混合权重的超参数的大值来防止拟合不足。两项模拟研究和一项真实数据研究的结果表明,通过适当设置先验分布,贝叶斯潜在类模型产生有效的插补,并优于竞争方法。
Latent class analysis has beer recently proposed for the multiple imputation (MI) of missing categorical data, using either a standard frequentist approach or a nonparametric Bayesian model called Dirichlet process mixture of multinomial distributions (DPMM). The main advantage of using a latent class model for multiple imputation is that it is very flexible in the sense that it car capture complex relationships in the data given that the number of latent classes is large enough. However, the two existing approaches also have certain disadvantages. The frequentist approach is computationally demanding because it requires estimating many LC models: first models with different number of classes should be estimated to determine the required number of classes and subsequently the selected model is reestimated for multiple bootstrap samples to take into account parameter uncertainty during the imputation stage. Whereas the Bayesian. Dirichlet process models perform the model selection and the handling of the parameter uncertainty automatically, the disadvantage of this method is that it tends to use a too small number of clusters during the Gibbs sampling, leading to an underfitting model yielding invalid imputations. In this paper, we propose an alternative approach which combined the strengths of the two existing approaches; that is, we use the Bayesian standard latent class model as an imputation model. We show how model selection can be performed prior to the imputation step using a single run of the Gibbs sampler and, moreover, show how underfitting is prevented by using large values for the hyperparameters of the mixture weights. The results of two simulation studies and one real-data study indicate that with a proper setting of the prior distributions, the Bayesian latent class model yields valid imputations and outperforms competing methods.