On the equivalence between Non-negative Matrix Factorization and Probabilistic Latent Semantic Indexing

On the equivalence between Non-negative Matrix Factorization and Probabilistic Latent Semantic Indexing
复制标题

DOI:
10.1016/j.csda.2008.01.011
复制
发表时间:
2008-04-15
影响因子:
1.8
通讯作者:
Peng, Wei
Peng, Wei
中科院分区:
数学3区
文献类型:
--
作者:
Ding, Chris;Li, Tao;Peng, Wei

文献摘要

被引文献

相似文献

最近,非负矩阵分解(NMF)和概率潜在语义索引(PLSI)被成功地应用于文档聚类。在本文中,我们证明了PLSI和NMF(具有I-散度目标函数)优化相同的目标函数,尽管经实验验证,PLSI和NMF是不同的算法。这为一种新的混合方法提供了理论依据,这种方法交替运行PLSI和NMF,每一种方法都先后跳出另一种方法的局部极小点,从而获得更好的最终解。在五个真实数据集上的大量实验表明了NMF和PLSI之间的关系,并表明该混合方法比仅NMF或仅PLSI的方法有显著的改进。我们还证明了在一阶近似下,NMF与X-2统计量是相同的。(C)2008年,爱思唯尔出版。
Non-negative Matrix Factorization (NMF) and Probabilistic Latent Semantic Indexing (PLSI) have been successfully applied to document clustering recently. In this paper, we show that PLSI and NMF (with the I-divergence objective function) optimize the same objective function, although PLSI and NMF are different algorithms as verified by experiments. This provides a theoretical basis for a new hybrid method that runs PLSI and NMF alternatively, each jumping out of the local minima of the other method successively, thus achieving a better final solution. Extensive experiments on five real-life datasets show relations between NMF and PLSI, and indicate that the hybrid method leads to significant improvements over NMF-only or PLSI-only methods. We also show that at first-order approximation, NMF is identical to the X-2-statistic. (c) 2008 Published by Elsevier B.V.