An Information Distillation Framework for Extractive Summarization

An Information Distillation Framework for Extractive Summarization
复制标题

DOI:
10.1109/taslp.2017.2764545
复制
发表时间:
2018
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Kuan-Yu Chen;Shih-Hung Liu;Berlin Chen;H. Wang
Kuan-Yu Chen;Shih-Hung Liu;Berlin Chen;H. Wang
中科院分区:
其他
文献类型:
--
作者:
Kuan-Yu Chen;Shih-Hung Liu;Berlin Chen;H. Wang

文献摘要

被引文献

相似文献

在自然语言处理的背景下,表征学习由于其在许多应用中的优异性能而成为一个新兴的研究课题。学习表征的话是一个开拓性的研究,在这所学校的研究。然而,段落(或句子和文档)嵌入学习更适合/合理的一些现实任务,如文档摘要。然而,经典的段落嵌入方法通过考虑段落中出现的所有单词来推断给定段落的表示。因此,那些频繁出现的停用词或功能词可能会误导嵌入学习过程,从而产生模糊的段落表示。受这些观察的启发,我们在本文中的主要贡献有三个方面。首先,我们提出了一种新的无监督段落嵌入方法,命名为本质向量(EV)模型,其目的是不仅从段落中提取最具代表性的信息,但也排除了一般的背景信息,以产生一个更丰富的低维向量表示的感兴趣的段落。其次,鉴于语音内容处理的重要性日益增加,EV模型的扩展,命名为去噪本质向量(D-EV)模型,提出。D-EV模型不仅继承了EV模型的优点,而且可以针对不完美的语音识别推断出针对给定口语段落的更鲁棒的表示。第三,提出了一种新的文摘框架,它可以同时考虑相关性和冗余信息。我们评估所提出的嵌入方法(即,EV和D-EV)和基于两个标准摘要语料库的摘要框架。实验结果表明,所提出的框架的有效性和适用性的几个实践和国家的最先进的摘要方法。
In the context of natural language processing, representation learning has emerged as a newly active research subject because of its excellent performance in many applications. Learning representations of words is a pioneering study in this school of research. However, paragraph (or sentence and document) embedding learning is more suitable/reasonable for some realistic tasks such as document summarization. Nevertheless, classic paragraph embedding methods infer the representation of a given paragraph by considering all of the words occurring in the paragraph. Consequently, those stop or function words that occur frequently may mislead the embedding learning process to produce a misty paragraph representation. Motivated by these observations, our major contributions in this paper are threefold. First, we propose a novel unsupervised paragraph embedding method, named the essence vector (EV) model, which aims at not only distilling the most representative information from a paragraph but also excluding the general background information to produce a more informative low-dimensional vector representation for the paragraph of interest. Second, in view of the increasing importance of spoken content processing, an extension of the EV model, named the denoising essence vector (D-EV) model, is proposed. The D-EV model not only inherits the advantages of the EV model but also can infer a more robust representation for a given spoken paragraph against imperfect speech recognition. Third, a new summarization framework, which can take both relevance and redundancy information into account simultaneously, is also introduced. We evaluate the proposed embedding methods (i.e., EV and D-EV) and the summarization framework on two benchmark summarization corpora. The experimental results demonstrate the effectiveness and applicability of the proposed framework in relation to several well-practiced and state-of-the-art summarization methods.