Probabilistic Topic Modeling for Comparative Analysis of Document Collections
Probabilistic Topic Modeling for Comparative Analysis of Document Collections
复制标题
DOI:
10.1145/3369873
复制
发表时间:
2020-03-01
影响因子:
3.6
通讯作者:
Reddy, Chandan K.
中科院分区:
文献类型:
--
作者:
Hua, Ting;Lu, Chang-Tien;Reddy, Chandan K.
Probabilistic topic models, which can discover hidden patterns in documents, have been extensively studied. However, rather than learning from a single document collection, numerous real-world applications demand a comprehensive understanding of the relationships among various document sets. To address such needs, this article proposes a new model that can identify the common and discriminative aspects of multiple datasets. Specifically, our proposed method is a Bayesian approach that represents each document as a combination of common topics (shared across all document sets) and distinctive topics (distributions over words that are exclusive to a particular dataset). Through extensive experiments, we demonstrate the effectiveness of our method compared with state-of-the-art models. The proposedmodel can be useful for "comparative thinking" analysis in real-world document collections.