A Correlated Topic Model Using Word Embeddings

A Correlated Topic Model Using Word Embeddings
复制标题

DOI:
10.24963/ijcai.2017/588
复制
发表时间:
2017-08
期刊:
--
影响因子:
--
通讯作者:
Guangxu Xun;Yaliang Li;Wayne Xin Zhao;Jing Gao;Aidong Zhang
Guangxu Xun;Yaliang Li;Wayne Xin Zhao;Jing Gao;Aidong Zhang
中科院分区:
其他
文献类型:
--
作者:
Guangxu Xun;Yaliang Li;Wayne Xin Zhao;Jing Gao;Aidong Zhang

文献摘要

被引文献

相似文献

传统的相关主题模型通过用logistic正态分布代替狄利克雷先验来捕捉潜在主题之间的关联结构。词嵌入已经被证明能够捕捉语言中的语义规律。因此,可以直接在词嵌入空间中计算词之间的语义相关性和相关性,例如通过余弦值。本文提出了一种基于词嵌入的相关主题模型。该模型使我们能够利用词嵌入中额外的词级相关信息,并在连续词嵌入空间中直接建模主题相关性。该模型将文档中的单词替换为有意义的词嵌入,将主题建模为词嵌入上的多元高斯分布,并在连续高斯主题之间学习主题相关性。给出了一种带有数据增广的Gibbs抽样解来进行推理。我们在20个新闻组数据集和路透社-21578数据集上定性和定量地评估了我们的模型。实验结果表明了该模型的有效性。
Conventional correlated topic models are able to capture correlation structure among latent topics by replacing the Dirichlet prior with the logistic normal distribution. Word embeddings have been proven to be able to capture semantic regularities in language. Therefore, the semantic relatedness and correlations between words can be directly calculated in the word embedding space, for example, via cosine values. In this paper, we propose a novel correlated topic model using word embed-dings. The proposed model enables us to exploit the additional word-level correlation information in word embeddings and directly model topic correlation in the continuous word embedding space. In the model, words in documents are replaced with meaningful word embeddings, topics are modeled as multivariate Gaussian distributions over the word embeddings and topic correlations are learned among the continuous Gaussian topics. A Gibbs sampling solution with data augmentation is given to perform inference. We evaluate our model on the 20 Newsgroups dataset and the Reuters-21578 dataset qualitatively and quantitatively. The experimental results show the effectiveness of our proposed model.