Mixture of latent trait analyzers for model-based clustering of categorical data

Mixture of latent trait analyzers for model-based clustering of categorical data
复制标题

DOI:
10.1007/s11222-013-9389-1
复制
发表时间:
2014-07-01
影响因子:
2.2
通讯作者:
Murphy, Thomas Brendan
Murphy, Thomas Brendan
中科院分区:
数学2区
文献类型:
--
作者:
Gollini, Isabella;Murphy, Thomas Brendan

文献摘要

被引文献

相似文献

连续数据的基于模型的聚类方法已经得到了很好的建立,并且在广泛的应用中得到了广泛的应用。然而,分类数据的基于模型的聚类方法不太标准。潜在类分析是一种常用的基于模型的二进制数据和/或分类数据的聚类方法,但由于假设的局部独立结构,估计的潜在类和感兴趣的群体中的组之间可能没有对应关系。潜在特质分析器模型的混合扩展潜在类分析,假设一个模型的分类响应变量,取决于一个分类的潜在类和一个连续的潜在特质变量;离散的潜在类容纳组结构和连续的潜在特质容纳这些组内的依赖。拟合潜在特质分析器模型的混合物是潜在困难的,因为似然函数涉及不能通过分析评估的积分。我们开发了一个变分的方法来拟合潜在特质模型的混合物,这提供了一个有效的模型拟合策略。潜在特质分析者的混合模型的分析数据来自国家长期护理调查(NLTCS)和美国国会的投票。该模型产生直观的聚类结果,它给出了一个更好的适合无论是潜在的类分析或潜在性状分析。
Model-based clustering methods for continuous data are well established and commonly used in a wide range of applications. However, model-based clustering methods for categorical data are less standard. Latent class analysis is a commonly used method for model-based clustering of binary data and/or categorical data, but due to an assumed local independence structure there may not be a correspondence between the estimated latent classes and groups in the population of interest. The mixture of latent trait analyzers model extends latent class analysis by assuming a model for the categorical response variables that depends on both a categorical latent class and a continuous latent trait variable; the discrete latent class accommodates group structure and the continuous latent trait accommodates dependence within these groups. Fitting the mixture of latent trait analyzers model is potentially difficult because the likelihood function involves an integral that cannot be evaluated analytically. We develop a variational approach for fitting the mixture of latent trait models and this provides an efficient model fitting strategy. The mixture of latent trait analyzers model is demonstrated on the analysis of data from the National Long Term Care Survey (NLTCS) and voting in the U.S. Congress. The model is shown to yield intuitive clustering results and it gives a much better fit than either latent class analysis or latent trait analysis alone.