Learning Topic Models by Neighborhood Aggregation

Learning Topic Models by Neighborhood Aggregation
复制标题

DOI:
10.24963/ijcai.2019/347
复制
发表时间:
2018-02
期刊:
--
影响因子:
--
通讯作者:
Ryohei Hisano
Ryohei Hisano
中科院分区:
其他
文献类型:
--
作者:
Ryohei Hisano

文献摘要

相似文献

主题模型由于其高度的可解释性和模块化的结构,经常被用于机器学习。然而,扩展主题模型以包括监督信号、合并预训练的词嵌入向量并包括非线性输出函数并不是一件容易的任务,因为必须诉诸于高度复杂的近似推理过程。本文表明,使用预先训练的词嵌入向量进行主题建模可以被视为实现邻域聚合算法,其中消息通过定义在词上的网络传递。从主题模型的网络视图来看,节点对应于文档中的单词,边对应于描述文档中共现单词的关系或描述语料库中相同单词的关系。网络视图允许我们扩展模型以包括监督信号,包含预训练的单词嵌入向量,并以简单的方式包括非线性输出函数。在实验中,我们表明,我们的方法优于国家的最先进的监督潜在狄利克雷分配实施方面举行了文件分类任务。
Topic models are frequently used in machine learning owing to their high interpretability and modular structure. However, extending a topic model to include a supervisory signal, to incorporate pre-trained word embedding vectors and to include a nonlinear output function is not an easy task because one has to resort to a highly intricate approximate inference procedure. The present paper shows that topic modeling with pre-trained word embedding vectors can be viewed as implementing a neighborhood aggregation algorithm where messages are passed through a network defined over words. From the network view of topic models, nodes correspond to words in a document and edges correspond to either a relationship describing co-occurring words in a document or a relationship describing the same word in the corpus. The network view allows us to extend the model to include supervisory signals, incorporate pre-trained word embedding vectors and include a nonlinear output function in a simple manner. In experiments, we show that our approach outperforms the state-of-the-art supervised Latent Dirichlet Allocation implementation in terms of held-out document classification tasks.