Mixtures of probabilistic principal component analyzers

Mixtures of probabilistic principal component analyzers
复制标题

DOI:
10.1162/089976699300016728
复制
发表时间:
1999-02-15
期刊:
影响因子:
2.9
通讯作者:
Bishop, CM
Bishop, CM
中科院分区:
计算机科学4区
文献类型:
--
作者:
Tipping, ME;Bishop, CM

文献摘要

被引文献

相似文献

主成分分析(PCA)是处理、压缩和可视化数据的最流行的技术之一,尽管其有效性受到其全局线性的限制。虽然已经提出了PCA的非线性变体,但另一种范式是通过局部线性PCA投影的组合来捕获数据复杂性。然而,传统的PCA不对应于概率密度,因此没有唯一的方式来联合收割机PCA模型。因此,以前的尝试制定混合物模型的PCA已经在一定程度上特设。在本文中,基于高斯潜变量模型的特定形式,在最大似然框架内制定了PCA。这导致了一个定义良好的混合模型的概率主成分分析,其参数可以使用期望最大化算法来确定。我们在聚类、密度建模和局部降维的背景下讨论了该模型的优势,并展示了它在图像压缩和手写数字识别中的应用。
Principal component analysis (PCA) is one of the most popular techniques for processing, compressing, and visualizing data, although its effectiveness is limited by its global linearity. While nonlinear variants of PCA have been proposed, an alternative paradigm is to capture data complexity by a combination of local linear PCA projections. However, conventional PCA does not correspond to a probability density, and so there is no unique way to combine PCA models. Therefore, previous attempts to formulate mixture models for PCA have been ad hoc to some extent. In this article, PCA is formulated within a maximum likelihood framework, based on a specific form of gaussian latent variable model. This leads to a well-defined mixture model for probabilistic principal component analyzers, whose parameters can be determined using an expectation-maximization algorithm. We discuss the advantages of this model in the context of clustering, density modeling, and local dimensionality reduction, and we demonstrate its application to image compression and handwritten digit recognition.