Information fusion for multidocument summarization: paraphrasing and generation

Information fusion for multidocument summarization: paraphrasing and generation
复制标题

DOI:
--
复制
发表时间:
2003
期刊:
--
影响因子:
--
通讯作者:
K. McKeown;R. Barzilay
K. McKeown;R. Barzilay
中科院分区:
其他
文献类型:
--
作者:
K. McKeown;R. Barzilay

文献摘要

被引文献

相似文献

在线新闻来源的数量和多样性使得人们很难追踪到哪怕是关于一个事件的新闻。冗余导致这样的跟踪非常耗时:同一事件的多个新闻提要往往包含相似的信息。此类新闻提要的摘要可以在一个简短的文本中呈现重要信息,从而大大减少阅读时间。这篇论文的重点是信息融合,这是一种在给定多个文档的情况下识别冗余信息并合成连贯摘要的技术。这种技术体现在MultiGen中,这是我在攻读博士期间设计、实现和评估的一个系统。与该领域以前的工作不同,MultiGen是一个独立于领域的系统:它生成关于任何领域各种主题的新闻摘要。对现有技术的另一个贡献是,系统通过重复使用和更改输入文章中的短语来生成摘要,从而创建更流畅和连贯的文本。这与其他现有系统形成了鲜明对比,其他现有系统只是从输入文章中提取句子并将它们连接在一起,从而导致流畅性问题。目前,MultiGen是哥伦比亚大学新闻播报系统的一部分。每天,Newsblaster从各种来源下载所有新闻文章,按主题对文章进行分类,并为每个文档分类生成有凝聚力的、可读的自动摘要。多文档摘要中的一个关键挑战是消除生成的摘要中的冗余信息。关于同一事件的文章经常使用不同的措辞描述同一事实。为了解决这个问题,我们需要一种方法来识别释义--表达相似意思的文本片段,即使它们在措辞上不相同。释义的自动识别在以前的研究中没有涉及,尽管它对于许多应用是必要的,包括问题回答、信息提取和自然语言生成。本文提出了一种基于多个平行文本语料库的无监督学习技术来识别释义。这种语料库提供了许多释义的例子,因为这些文本保留了原始来源的意思,但可能使用不同的词来传达意思。数据和方法都与过去基于语料库的技术方法不同。我们的评测实验表明,该算法提取复述的准确率很高,显著优于为机器翻译相关任务开发的最新算法。
The number and variety of online news sources makes it difficult for people to track the news concerning even a single event. Redundancy causes such tracking to be extremely time-consuming: multiple news feeds on the same event tend to contain similar information. A summary of such news feeds can present important information in one short text, dramatically reducing reading time. The focus of this thesis is information fusion, a technique which, given multiple documents, identifies redundant information and synthesizes a coherent summary. This technique is embodied in MultiGen, a system that I have designed, implemented and evaluated over the course of my Ph.D. Unlike previous work in the area, MultiGen is a domain-independent system: it generates news summaries on a variety of topics in any domain. Another contribution to the state of the art is that the system generates the summary by reusing and altering phrases from the input articles, creating a more fluent and cohesive text. This is in contrast with other existing systems, which simply extract sentences from input articles and concatenate them together, leading to fluency problems. Currently MultiGen operates as part of Columbia's Newsblaster system. Everyday, Newsblaster downloads all news articles from a variety of sources, clusters articles by topic, and generates a cohesive, readable automatic summary of each document cluster. One key challenge in multidocument summarization is eliminating redundant information in the produced summaries. Articles about the same event often contain descriptions of the same fact using different wording. To address this issue, we need a method to identify paraphrases—fragments of text that convey similar meaning even if they are not identical in wording. Automatic identification of paraphrases was not addressed in previous research, although it is necessary for many applications, including question-answering, information extraction and natural language generation. This thesis presents unsupervised learning techniques to identify paraphrases given a corpus of multiple parallel texts. This type of corpus provides many instances of paraphrasing, because these texts preserve the meaning of the original source, but may use different words to convey the meaning. Both the data and the method are departures from past approaches to corpus based techniques. Our evaluation experiments show that the algorithm extracts paraphrases with high accuracy and significantly outperforms a state of the art algorithm developed for related tasks in machine translation.