Extending mixtures of multivariate t-factor analyzers

Extending mixtures of multivariate t-factor analyzers
复制标题

DOI:
10.1007/s11222-010-9175-2
复制
发表时间:
2011-07-01
影响因子:
2.2
通讯作者:
McNicholas, Paul D.
McNicholas, Paul D.
中科院分区:
数学2区
文献类型:
--
作者:
Andrews, Jeffrey L.;McNicholas, Paul D.

文献摘要

被引文献

相似文献

基于模型的聚类通常涉及开发一系列混合模型并将这些模型强加于数据。然后,使用某种标准来选择该家族中的最佳成员,并且相关的参数估计导致预测的组成员关系,或集群。本文描述了多变量混合t因子分析模型的扩展,包括对自由度、因子载荷和误差方差矩阵的约束。结果是一个由六个混合模型组成的家族,其中包括节俭模型。这一族模型的参数估计是使用交替期望-条件最大化算法得到的,并基于Aitken的加速来确定收敛。模型选择使用贝叶斯信息准则(BIC)和积分完全似然(ICL)。然后将这种新的混合模型系列应用于模拟数据和真实数据,在这些数据中,聚类性能达到或超过已建立的基于模型的聚类方法。仿真研究包括对BIC和ICL作为这一新模型家族的模型选择技术的比较。并探讨了其在高维模拟数据中的应用。
Model-based clustering typically involves the development of a family of mixture models and the imposition of these models upon data. The best member of the family is then chosen using some criterion and the associated parameter estimates lead to predicted group memberships, or clusterings. This paper describes the extension of the mixtures of multivariate t-factor analyzers model to include constraints on the degrees of freedom, the factor loadings, and the error variance matrices. The result is a family of six mixture models, including parsimonious models. Parameter estimates for this family of models are derived using an alternating expectation-conditional maximization algorithm and convergence is determined based on Aitken's acceleration. Model selection is carried out using the Bayesian information criterion (BIC) and the integrated completed likelihood (ICL). This novel family of mixture models is then applied to simulated and real data where clustering performance meets or exceeds that of established model-based clustering methods. The simulation studies include a comparison of the BIC and the ICL as model selection techniques for this novel family of models. Application to simulated data with larger dimensionality is also explored.