GraphBTM: Graph Enhanced Autoencoded Variational Inference for Biterm Topic Model

GraphBTM: Graph Enhanced Autoencoded Variational Inference for Biterm Topic Model
复制标题

DOI:
10.18653/v1/d18-1495
复制
发表时间:
2018
期刊:
--
影响因子:
--
通讯作者:
Qile Zhu;Zheng Feng;Xiaolin Li
Qile Zhu;Zheng Feng;Xiaolin Li
中科院分区:
其他
文献类型:
--
作者:
Qile Zhu;Zheng Feng;Xiaolin Li

文献摘要

被引文献

相似文献

发现文本中的潜在主题一直是许多应用的基本任务。然而,传统的主题模型在不同的设置中遭受不同的问题。潜在狄利克雷分配(LDA)可能无法很好地工作,由于短文本的数据稀疏性(即稀疏的词同现模式在短文档)。双术语主题模型(BTM)通过对整个语料库中的双术语词对进行建模来学习主题。当文档很长,主题信息丰富,并且没有表现出双项的传递性时,这个假设非常强。在本文中,我们提出了一种新的方式称为GraphBTM表示为图的双项和设计一个图卷积网络(GCN)与剩余连接提取传递功能的双项。为了克服LDA的数据稀疏性和BTM的强假设,我们抽取固定数量的文档形成一个小型语料库作为样本。我们还提出了一个数据集,称为所有新闻提取15个新闻出版商,其中的文件远远超过20个新闻组。我们提出了一个摊销变分推理方法GraphBTM。与以前的方法相比,我们的方法生成更连贯的主题。实验结果表明,该采样策略在很大程度上提高了性能.
Discovering the latent topics within texts has been a fundamental task for many applications. However, conventional topic models suffer different problems in different settings. The Latent Dirichlet Allocation (LDA) may not work well for short texts due to the data sparsity (i.e. the sparse word co-occurrence patterns in short documents). The Biterm Topic Model (BTM) learns topics by modeling the word-pairs named biterms in the whole corpus. This assumption is very strong when documents are long with rich topic information and do not exhibit the transitivity of biterms. In this paper, we propose a novel way called GraphBTM to represent biterms as graphs and design a Graph Convolutional Networks (GCNs) with residual connections to extract transitive features from biterms. To overcome the data sparsity of LDA and the strong assumption of BTM, we sample a fixed number of documents to form a mini-corpus as a sample. We also propose a dataset called All News extracted from 15 news publishers, in which documents are much longer than 20 Newsgroups. We present an amortized variational inference method for GraphBTM. Our method generates more coherent topics compared with previous approaches. Experiments show that the sampling strategy improves performance by a large margin.