Probabilistic Topic Modeling for Comparative Analysis of Document Collections

Probabilistic Topic Modeling for Comparative Analysis of Document Collections
复制标题

DOI:
10.1145/3369873
复制
发表时间:
2020-03-01
影响因子:
3.6
通讯作者:
Reddy, Chandan K.
Reddy, Chandan K.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hua, Ting;Lu, Chang-Tien;Reddy, Chandan K.

文献摘要

被引文献

相似文献

概率主题模型可以发现文档中隐藏的模式,已经得到了广泛的研究。然而,而不是从一个单一的文档集合中学习,许多现实世界的应用程序需要全面了解各种文档集之间的关系。为了满足这些需求,本文提出了一种新的模型,可以识别多个数据集的共同和区别方面。具体来说,我们提出的方法是一种贝叶斯方法,它将每个文档表示为共同主题(在所有文档集上共享)和独特主题(在特定数据集上专有的单词上的分布)的组合。通过大量的实验,我们证明了我们的方法相比,国家的最先进的模型的有效性。所提出的模型可用于现实世界中的文档集合的“比较思维”分析。
Probabilistic topic models, which can discover hidden patterns in documents, have been extensively studied. However, rather than learning from a single document collection, numerous real-world applications demand a comprehensive understanding of the relationships among various document sets. To address such needs, this article proposes a new model that can identify the common and discriminative aspects of multiple datasets. Specifically, our proposed method is a Bayesian approach that represents each document as a combination of common topics (shared across all document sets) and distinctive topics (distributions over words that are exclusive to a particular dataset). Through extensive experiments, we demonstrate the effectiveness of our method compared with state-of-the-art models. The proposedmodel can be useful for "comparative thinking" analysis in real-world document collections.