A framework for multi-document abstractive summarization based on semantic role labelling

A framework for multi-document abstractive summarization based on semantic role labelling
复制标题

DOI:
10.1016/j.asoc.2015.01.070
复制
发表时间:
2015-05-01
影响因子:
8.7
通讯作者:
Kumar, Yogan Jaya
Kumar, Yogan Jaya
中科院分区:
计算机科学2区
文献类型:
--
作者:
Khan, Atif;Salim, Naomie;Kumar, Yogan Jaya

文献摘要

被引文献

相似文献

本文提出了一个多文档摘要框架,其目的是从源文档的语义表达中选择摘要的内容,而不是从源文档的句子中选择摘要的内容。在该框架中,源文档的内容通过语义角色标注由谓词论元结构表示。摘要的内容选择是通过基于优化的特征对谓词论元结构进行排序,并使用语言生成来从谓词论元结构生成句子。我们提出的框架不同于其他抽象摘要方法在几个方面。首先,它采用语义角色标注的文本的语义表示。其次,利用语义相似性度量对源文本进行语义分析,对文本中语义相似的谓词论元结构进行聚类;最后,利用遗传算法对谓词论元结构进行特征加权排序。本文的实验是在文本摘要标准语料库DUC-2002上进行的。实验结果表明,该方法的性能优于其他摘要系统。(C)2015爱思唯尔B. V.保留所有权利。
We propose a framework for abstractive summarization of multi-documents, which aims to select contents of summary not from the source document sentences but from the semantic representation of the source documents. In this framework, contents of the source documents are represented by predicate argument structures by employing semantic role labeling. Content selection for summary is made by ranking the predicate argument structures based on optimized features, and using language generation for generating sentences from predicate argument structures. Our proposed framework differs from other abstractive summarization approaches in a few aspects. First, it employs semantic role labeling for semantic representation of text. Secondly, it analyzes the source text semantically by utilizing semantic similarity measure in order to cluster semantically similar predicate argument structures across the text; and finally it ranks the predicate argument structures based on features weighted by genetic algorithm (GA). Experiment of this study is carried out using DUC-2002, a standard corpus for text summarization. Results indicate that the proposed approach performs better than other summarization systems. (C) 2015 Elsevier B.V. All rights reserved.