A novel contextual topic model for multi-document summarization

A novel contextual topic model for multi-document summarization
复制标题

DOI:
10.1016/j.eswa.2014.09.015
复制
发表时间:
2015-02-15
影响因子:
8.5
通讯作者:
Sutinen, Erkki
Sutinen, Erkki
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yang, Guangbing;Wen, Dunwei;Sutinen, Erkki

文献摘要

被引文献

相似文献

信息过载成为数字时代的一个严重问题。它会对有用信息的理解产生负面影响。如何缓解这一问题是自然语言处理,特别是多文档自动摘要研究的主要关注点。为了寻求一种新的方法来帮助证明相似句子在多文档摘要中的重要性,本研究提出了一种新的方法,基于最近的分层贝叶斯主题模型。该模型将n-gram的概念融入到分层的潜在主题中,以捕获出现在单词的局部上下文中的单词依赖关系。定量和定性的评价结果表明,该模型在文档建模方面优于hLDA和LDA。此外,在实践中的实验结果表明,我们的摘要系统实现该模型可以显着提高性能,使其媲美国家的最先进的摘要系统。(C)2014爱思唯尔有限公司版权所有。
Information overload becomes a serious problem in the digital age. It negatively impacts understanding of useful information. How to alleviate this problem is the main concern of research on natural language processing, especially multi-document summarization. With the aim of seeking a new method to help justify the importance of similar sentences in multi-document summarizations, this study proposes a novel approach based on recent hierarchical Bayesian topic models. The proposed model incorporates the concepts of n-grams into hierarchically latent topics to capture the word dependencies that appear in the local context of a word. The quantitative and qualitative evaluation results show that this model has outperformed both hLDA and LDA in document modeling. In addition, the experimental results in practice demonstrate that our summarization system implementing this model can significantly improve the performance and make it comparable to the state-of-the-art summarization systems. (C) 2014 Elsevier Ltd. All rights reserved.