A Tensor-EM Method for Large-Scale Latent Class Analysis with Binary Responses

A Tensor-EM Method for Large-Scale Latent Class Analysis with Binary Responses
复制标题

DOI:
10.1007/s11336-022-09887-1
复制
发表时间:
2021-03
期刊:
影响因子:
3
通讯作者:
Zhenghao Zeng;Yuqi Gu;Gongjun Xu
Zhenghao Zeng;Yuqi Gu;Gongjun Xu
中科院分区:
心理学4区
文献类型:
--
作者:
Zhenghao Zeng;Yuqi Gu;Gongjun Xu

文献摘要

相似文献

潜在类模型是广泛应用于心理学、行为学和社会科学的强大统计建模工具。在现代数据科学时代,研究人员经常可以访问从大规模调查或评估中收集的响应数据,这些数据包含许多项目(大J)和许多主题(大N)。这与传统的固定J和大N的制度相反。为了分析如此大规模的数据,重要的是开发计算效率和理论有效的方法。在计算方面,传统的EM算法的潜在类模型往往有一个缓慢的算法收敛速度为大规模的数据,并可能收敛到一些局部最优值,而不是最大似然估计(MLE)。出于这一动机,我们引入张量分解的角度到潜在的类分析与二进制响应。在方法上,我们建议在第一步中使用基于矩的张量幂方法,然后在第二步中使用所获得的估计作为EM算法的初始化。从理论上讲,我们建立了聚类一致性的最大似然估计分配到潜在类时,N和J都走向无穷大。仿真研究表明,所提出的张量EM管道具有良好的精度和计算效率的大规模数据与二进制响应。我们还将所提出的方法应用于教育评估数据集作为例证。
Latent class models are powerful statistical modeling tools widely used in psychological, behavioral, and social sciences. In the modern era of data science, researchers often have access to response data collected from large-scale surveys or assessments, featuring many items (large J) and many subjects (large N). This is in contrary to the traditional regime with fixed J and large N. To analyze such large-scale data, it is important to develop methods that are both computationally efficient and theoretically valid. In terms of computation, the conventional EM algorithm for latent class models tends to have a slow algorithmic convergence rate for large-scale data and may converge to some local optima instead of the maximum likelihood estimator (MLE). Motivated by this, we introduce the tensor decomposition perspective into latent class analysis with binary responses. Methodologically, we propose to use a moment-based tensor power method in the first step and then use the obtained estimates as initialization for the EM algorithm in the second step. Theoretically, we establish the clustering consistency of the MLE in assigning subjects into latent classes when N and J both go to infinity. Simulation studies suggest that the proposed tensor-EM pipeline enjoys both good accuracy and computational efficiency for large-scale data with binary responses. We also apply the proposed method to an educational assessment dataset as an illustration.