Self-Supervised Neural Topic Modeling

Self-Supervised Neural Topic Modeling
复制标题

DOI:
10.18653/v1/2021.findings-emnlp.284
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Seyed Ali Bahrainian;Martin Jaggi;Carsten Eickhoff
Seyed Ali Bahrainian;Martin Jaggi;Carsten Eickhoff
中科院分区:
其他
文献类型:
--
作者:
Seyed Ali Bahrainian;Martin Jaggi;Carsten Eickhoff

文献摘要

相似文献

主题模型是分析和解释大型文本语料库的主要潜在主题的有用工具。大多数主题模型依赖于词共现来计算主题,即,一组加权的词,它们共同代表一个高级语义概念。在本文中,我们提出了一个新的轻量级的自监督神经主题模型(SNTM),学习丰富的上下文,通过学习主题表示联合从三个共现的单词和一个文件,三重起源。我们的实验结果表明,我们提出的神经主题模型SNTM在一致性指标和文档聚类准确性方面优于之前现有的主题模型。此外,除了主题一致性和聚类性能外,所提出的神经主题模型具有许多优点,即计算效率高,易于训练。
Topic models are useful tools for analyzing and interpreting the main underlying themes of large corpora of text. Most topic models rely on word co-occurrence for computing a topic, i.e., a weighted set of words that together represent a high-level semantic concept. In this paper, we propose a new light-weight Self-Supervised Neural Topic Model (SNTM) that learns a rich context by learning a topic representation jointly from three co-occurring words and a document that the triplet originates from. Our experimental results indicate that our proposed neural topic model, SNTM, outperforms previously existing topic models in coherence metrics as well as document clustering accuracy. Moreover, apart from the topic coherence and clustering performance, the proposed neural topic model has a number of advantages, namely, being computationally efficient and easy to train.