Summarization beyond sentence extraction: A probabilistic approach to sentence compression

Summarization beyond sentence extraction: A probabilistic approach to sentence compression
复制标题

DOI:
10.1016/s0004-3702(02)00222-9
复制
发表时间:
2002-07
期刊:
Artif. Intell.
影响因子:
--
通讯作者:
Kevin Knight;D. Marcu
Kevin Knight;D. Marcu
中科院分区:
其他
文献类型:
--
作者:
Kevin Knight;D. Marcu

文献摘要

被引文献

相似文献

当人类制作文档摘要时,他们不会简单地提取句子并将它们连接起来。相反,它们创建的新句子符合语法,彼此连贯,并捕获了原始文档中最重要的信息。鉴于大量的文本/摘要对可以在网上获得,现在可以设想经过训练的算法来模拟这一过程。在这篇文章中,我们关注句子压缩,这是这一更大挑战的一个更简单的版本。我们的目标是同时实现两个目标:我们的压缩应该是符合语法的,并且它们应该保留最重要的信息。这两个目标可能会发生冲突。我们设计了噪声通道和决策树方法来解决该问题,并针对手动压缩和简单的基线对结果进行了评估。
When humans produce summaries of documents, they do not simply extract sentences and concatenate them. Rather, they create new sentences that are grammatical, that cohere with one another, and that capture the most salient pieces of information in the original document. Given that large collections of text/abstract pairs are available online, it is now possible to envision algorithms that are trained to mimic this process. In this paper, we focus on sentence compression, a simpler version of this larger challenge. We aim to achieve two goals simultaneously: our compressions should be grammatical, and they should retain the most important pieces of information. These two goals can conflict. We devise both a noisy-channel and a decision-tree approach to the problem, and we evaluate results against manual compressions and a simple baseline.