A Spectral Method for Identifiable Grade of Membership Analysis with Binary Responses

A Spectral Method for Identifiable Grade of Membership Analysis with Binary Responses
复制标题

DOI:
10.1007/s11336-024-09951-y
复制
发表时间:
2024-02-15
期刊:
影响因子:
3
通讯作者:
Gu,Yuqi
Gu,Yuqi
中科院分区:
心理学4区
文献类型:
--
作者:
Chen,Ling;Gu,Yuqi

文献摘要

相似文献

隶属度模型(Grade of Membership,GoM)是一种流行的多变量分类数据的个体水平混合模型。GoM允许每个主题在多个极端潜在配置文件中具有混合成员资格。因此,GoM模型比限制每个主题属于单个配置文件的潜在类模型具有更丰富的建模能力。GoM的灵活性是以更具挑战性的可识别性和估计问题为代价的。在这项工作中,我们提出了一个奇异值分解(SVD)为基础的频谱方法GoM分析与多元二进制响应。我们的方法取决于观察到的期望的数据矩阵有一个低秩分解下的GoM模型。对于可识别性,我们发展了期望可识别性概念的充分和几乎必要条件。预测,我们只提取几个领先的奇异向量的观测数据矩阵,并利用这些向量的单纯形几何估计的混合隶属度得分和其他参数。我们还建立了我们的估计在双渐近制度的主题和项目的数量都增长到无穷大的一致性。我们的谱方法比贝叶斯或基于似然的方法具有巨大的计算优势,并且可扩展到大规模和高维数据。大量的仿真研究表明,我们的方法的上级效率和准确性。我们还通过将其应用于个性测试数据集来说明我们的方法。
Grade of membership (GoM) models are popular individual-level mixture models for multivariate categorical data. GoM allows each subject to have mixed memberships in multiple extreme latent profiles. Therefore, GoM models have a richer modeling capacity than latent class models that restrict each subject to belong to a single profile. The flexibility of GoM comes at the cost of more challenging identifiability and estimation problems. In this work, we propose a singular value decomposition (SVD)-based spectral approach to GoM analysis with multivariate binary responses. Our approach hinges on the observation that the expectation of the data matrix has a low-rank decomposition under a GoM model. Foridentifiability, we develop sufficient and almost necessary conditions for a notion of expectation identifiability. Forestimation, we extract only a few leading singular vectors of the observed data matrix and exploit the simplex geometry of these vectors to estimate the mixed membership scores and other parameters. We also establish the consistency of our estimator in the double-asymptotic regime where both the number of subjects and the number of items grow to infinity. Our spectral method has a huge computational advantage over Bayesian or likelihood-based methods and is scalable to large-scale and high-dimensional data. Extensive simulation studies demonstrate the superior efficiency and accuracy of our method. We also illustrate our method by applying it to a personality test dataset.