Modeling hidden topics on document manifold

Modeling hidden topics on document manifold
复制标题

DOI:
10.1145/1458082.1458202
复制
发表时间:
2008-10
期刊:
--
影响因子:
--
通讯作者:
Deng Cai;Qiaozhu Mei;Jiawei Han;ChengXiang Zhai
Deng Cai;Qiaozhu Mei;Jiawei Han;ChengXiang Zhai
中科院分区:
其他
文献类型:
--
作者:
Deng Cai;Qiaozhu Mei;Jiawei Han;ChengXiang Zhai

文献摘要

被引文献

相似文献

主题建模一直是文档分析的关键问题。用于主题建模的规范方法之一是概率潜在语义索引,其最大化语料库中的文档和术语的联合概率。PLSI的主要缺点是它独立地估计每个文档在隐藏主题上的概率分布,并且模型中的参数数量随着语料库的大小线性增长,这导致严重的过拟合问题。潜在的Dirichlet分配(LDA)被提出来克服这个问题,通过处理的概率分布的每个文档的主题作为一个隐藏的随机变量。这两种方法都能发现欧氏空间中隐藏的主题。然而,没有令人信服的证据表明,文件空间是欧几里得,或平坦的。因此,假设文档空间是线性或非线性的流形是更自然和合理的。在本文中,我们考虑的问题,主题建模的内在文档流形。具体来说,我们提出了一种新的算法称为拉普拉斯概率潜在语义索引(LapPLSI)的主题建模。LapPLSI将文档空间建模为嵌入在环境空间中的子流形,并直接在该文档流形上执行主题建模。我们比较了三个文本数据集的建议LapPLSI方法与PLSI和LDA。实验结果表明,LapPLSI在语义结构上具有更好的表示效果.
Topic modeling has been a key problem for document analysis. One of the canonical approaches for topic modeling is Probabilistic Latent Semantic Indexing, which maximizes the joint probability of documents and terms in the corpus. The major disadvantage of PLSI is that it estimates the probability distribution of each document on the hidden topics independently and the number of parameters in the model grows linearly with the size of the corpus, which leads to serious problems with overfitting. Latent Dirichlet Allocation (LDA) is proposed to overcome this problem by treating the probability distribution of each document over topics as a hidden random variable. Both of these two methods discover the hidden topics in the Euclidean space. However, there is no convincing evidence that the document space is Euclidean, or flat. Therefore, it is more natural and reasonable to assume that the document space is a manifold, either linear or nonlinear. In this paper, we consider the problem of topic modeling on intrinsic document manifold. Specifically, we propose a novel algorithm called Laplacian Probabilistic Latent Semantic Indexing (LapPLSI) for topic modeling. LapPLSI models the document space as a submanifold embedded in the ambient space and directly performs the topic modeling on this document manifold in question. We compare the proposed LapPLSI approach with PLSI and LDA on three text data sets. Experimental results show that LapPLSI provides better representation in the sense of semantic structure.