Extractive Summarization: Limits, Compression, Generalized Model and Heuristics

Extractive Summarization: Limits, Compression, Generalized Model and Heuristics
复制标题

提取概括:限制、压缩、广义模型和启发式

DOI:
10.13053/cys-21-4-2885
复制
发表时间:
2017
期刊:
Computación y Sistemas
影响因子:
--
通讯作者:
Daniel Lee
Daniel Lee
中科院分区:
--
文献类型:
--
作者:
Rakesh M. Verma;Daniel Lee

文献摘要

被引文献

相似文献

摘要技术因其在缓解信息过载方面的优势而引起了研究者的广泛关注。然而,这仍然是一个严峻的挑战。在这里,我们首先证明了在单文档和多文档摘要任务的ROUGE评估下DUC数据集上提取摘要器的召回率(和F1分数)的经验限制。接下来,我们定义了文档的可压缩性的概念,并提出了一种新的摘要模型,该模型概括了文献中现有的模型,并集成了摘要的几个维度,即,抽象与提取、单文档与多文档以及句法与语义。最后,我们研究了一些新的和现有的单文档摘要算法在一个单一的框架,并比较与国家的最先进的摘要DUC数据。
Due to its promise to alleviate information overload, text summarization has attracted the attention of many researchers. However, it has remained a serious challenge. Here, we first prove empirical limits on the recall (and F1-scores) of extractive summarizers on the DUC datasets under ROUGE evaluation for both the single-document and multi-document summarization tasks. Next we define the concept of compressibility of a document and present a new model of summarization, which generalizes existing models in the literature and integrates several dimensions of the summarization, viz., abstractive versus extractive, single versus multi-document, and syntactic versus semantic. Finally, we examine some new and existing single-document summarization algorithms in a single framework and compare with state of the art summarizers on DUC data.