A Bidirectional Hierarchical Skip-Gram model for text topic embedding

A Bidirectional Hierarchical Skip-Gram model for text topic embedding
复制标题

DOI:
10.1109/ijcnn.2016.7727289
复制
发表时间:
2016-07
期刊:
2016 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Suncong Zheng;Hongyun Bao;Jiaming Xu;Yuexing Hao;Zhenyu Qi;Hongwei Hao
Suncong Zheng;Hongyun Bao;Jiaming Xu;Yuexing Hao;Zhenyu Qi;Hongwei Hao
中科院分区:
其他
文献类型:
--
作者:
Suncong Zheng;Hongyun Bao;Jiaming Xu;Yuexing Hao;Zhenyu Qi;Hongwei Hao

文献摘要

被引文献

相似文献

利用网络上大规模的语料库有效、高效地挖掘文本中的主题是大数据时代的一个重要问题。我们专注于以无监督的方式学习文本主题嵌入的问题,它具有效率和可扩展性的特性。文本主题嵌入表示语义主题空间中的单词和文档,其中具有相似主题的单词和文档将彼此靠近地嵌入。与隐式捕获文档级单词共现模式的传统主题模型相比,文本主题嵌入缓解了数据稀疏问题并捕获了不同单词和文档之间的语义相关性。为了对文本主题嵌入进行建模,我们提出了一种基于skip-gram模型的双向分层Skip-Gram模型(BHSG)。 BHSG 包括两个组件:语义生成模块,用于学习文本之间的语义相关性;主题增强模块,用于基于在前一个模块中学习的文本嵌入来生成文本主题嵌入。我们在两种与主题相关的任务上评估了我们的方法:文本分类和信息检索。在四个公共数据集和我们提供的一个数据集上的实验结果都表明我们提出的方法可以获得更好的性能。
Taking advantage of the large scale corpus on the web to effectively and efficiently mine the topics within texts is an essential problem in the era of big data. We focus on the problem of learning text topic embedding in an unsupervised manner, which enjoys the properties of efficiency and scalability. Text topic embedding represents words and documents in a semantic topic space, in which the words and documents with similar topic will be embedded close to each other. When compared with conventional topic models, which implicitly capture the document-level word co-occurrence patterns, text topic embedding alleviates the data sparsity problem and captures the semantic relevance between different words and documents. To model text topic embedding, we propose a Bidirectional Hierarchical Skip-Gram model (BHSG) based on skip-gram model. BHSG includes two components: semantic generation module to learn semantic relevance between texts and topic enhance module to produce the text topic embedding based on text embedding learned in the former module. We evaluated our method on two kinds of topic-related tasks: text classification and information retrieval. The experimental results on four public datasets and one dataset we provide all demonstrate that our proposed method can achieve a better performance.