Neural Embedding Allocation: Distributed Representations of Topic Models

Neural Embedding Allocation: Distributed Representations of Topic Models
复制标题

DOI:
10.1162/coli_a_00457
复制
发表时间:
2019-09
影响因子:
9.3
通讯作者:
Kamrun Keya;Yannis Papanikolaou;James R. Foulds
Kamrun Keya;Yannis Papanikolaou;James R. Foulds
中科院分区:
计算机科学3区
文献类型:
--
作者:
Kamrun Keya;Yannis Papanikolaou;James R. Foulds

文献摘要

相似文献

摘要提出了一种利用神经嵌入来提高任意给定lda风格主题模型性能的方法。我们的方法称为神经嵌入分配(NEA),通过学习神经嵌入来模仿主题模型,将主题模型(LDA或其他)解构为词、主题、文档、作者等的可解释向量空间嵌入。我们证明,当主题数量较大时,NEA通过平滑噪声主题来提高原始主题模型的相干性分数。此外,我们展示了NEA在解构和平滑LDA、作者-主题模型和最近的混合成员跳过-克主题模型方面的有效性和通用性,并且与几个最先进的模型相比,通过嵌入获得了更好的性能。
Abstract We propose a method that uses neural embeddings to improve the performance of any given LDA-style topic model. Our method, called neural embedding allocation (NEA), deconstructs topic models (LDA or otherwise) into interpretable vector-space embeddings of words, topics, documents, authors, and so on, by learning neural embeddings to mimic the topic model. We demonstrate that NEA improves coherence scores of the original topic model by smoothing out the noisy topics when the number of topics is large. Furthermore, we show NEA’s effectiveness and generality in deconstructing and smoothing LDA, author-topic models, and the recent mixed membership skip-gram topic model and achieve better performance with the embeddings compared to several state-of-the-art models.