LTSG: Latent Topical Skip-Gram for Mutually Learning Topic Model and Vector Representations

LTSG: Latent Topical Skip-Gram for Mutually Learning Topic Model and Vector Representations
复制标题

DOI:
--
复制
发表时间:
2017-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Jarvan Law;Hankui Zhuo;Junhua He;Erhu Rong
Jarvan Law;Hankui Zhuo;Junhua He;Erhu Rong
中科院分区:
其他
文献类型:
--
作者:
Jarvan Law;Hankui Zhuo;Junhua He;Erhu Rong

文献摘要

被引文献

相似文献

在文本挖掘中,主题模型被广泛应用于发现跨文档共享的潜在主题。向量表示、词嵌入和主题嵌入将词和主题映射到低维密集的实值向量空间,在自然语言处理任务中取得了较高的性能。然而,现有的大多数模型都假设其中一个模型的训练结果是完全正确的,并将其作为先验知识来改进另一个模型。其他一些模型使用从外部大型语料库训练的信息来帮助改进较小的语料库。在本文中,我们的目标是建立这样一个算法框架,使主题模型和向量表示在同一语料库内相互改进。采用EM风格的算法框架迭代优化主题模型和向量表示。实验结果表明,在不同的自然语言处理任务上,我们的模型比现有的方法具有更好的性能。
Topic models have been widely used in discovering latent topics which are shared across documents in text mining. Vector representations, word embeddings and topic embeddings, map words and topics into a low-dimensional and dense real-value vector space, which have obtained high performance in NLP tasks. However, most of the existing models assume the result trained by one of them are perfect correct and used as prior knowledge for improving the other model. Some other models use the information trained from external large corpus to help improving smaller corpus. In this paper, we aim to build such an algorithm framework that makes topic models and vector representations mutually improve each other within the same corpus. An EM-style algorithm framework is employed to iteratively optimize both topic model and vector representations. Experimental results show that our model outperforms state-of-art methods on various NLP tasks.