Dimension-Grouped Mixed Membership Models for Multivariate Categorical Data

Dimension-Grouped Mixed Membership Models for Multivariate Categorical Data
复制标题

DOI:
--
复制
发表时间:
2021-09
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Yuqi Gu;E. Erosheva;Gongjun Xu;D. Dunson
Yuqi Gu;E. Erosheva;Gongjun Xu;D. Dunson
中科院分区:
其他
文献类型:
--
作者:
Yuqi Gu;E. Erosheva;Gongjun Xu;D. Dunson

文献摘要

被引文献

相似文献

混合隶属度模型(MMMs)是一种流行的复杂多元数据潜在结构模型。mm没有强迫每个主题属于单个集群,而是结合了一个特定主题的权重向量,该权重表示跨集群的部分隶属关系。有了这种灵活性,在唯一地识别、估计和解释参数方面就出现了挑战。在本文中,我们提出了一种新的多维分类数据的维度分组hmm (Gro-M$^3$s),它提高了简化性和可解释性。在Gro-M$^3$s中,观察到的变量被划分成组,使得组内变量的潜在隶属度是恒定的,但组间可能不同。传统的潜在类模型是在所有变量都在一组时得到的,而传统的hmm是在每个变量都在自己的组时得到的。新模型对应于一种新的概率张量分解。理论上,我们导出了在一般情况下未知分组结构和模型参数的透明可辨识性条件。在方法上,我们提出了Dirichlet grom $^3$s的贝叶斯方法来推断变量分组结构和估计模型参数。仿真结果显示了良好的计算性能,并从经验上验证了可辨识性结果。我们通过对功能性残疾调查数据集和个性测试数据集的应用来说明新方法。
Mixed Membership Models (MMMs) are a popular family of latent structure models for complex multivariate data. Instead of forcing each subject to belong to a single cluster, MMMs incorporate a vector of subject-specific weights characterizing partial membership across clusters. With this flexibility come challenges in uniquely identifying, estimating, and interpreting the parameters. In this article, we propose a new class of Dimension-Grouped MMMs (Gro-M$^3$s) for multivariate categorical data, which improve parsimony and interpretability. In Gro-M$^3$s, observed variables are partitioned into groups such that the latent membership is constant for variables within a group but can differ across groups. Traditional latent class models are obtained when all variables are in one group, while traditional MMMs are obtained when each variable is in its own group. The new model corresponds to a novel decomposition of probability tensors. Theoretically, we derive transparent identifiability conditions for both the unknown grouping structure and model parameters in general settings. Methodologically, we propose a Bayesian approach for Dirichlet Gro-M$^3$s to inferring the variable grouping structure and estimating model parameters. Simulation results demonstrate good computational performance and empirically confirm the identifiability results. We illustrate the new methodology through applications to a functional disability survey dataset and a personality test dataset.