Toward a Unified Framework for Standard and Update Multi-Document Summarization

Toward a Unified Framework for Standard and Update Multi-Document Summarization
复制标题

DOI:
10.1145/2184436.2184438
复制
发表时间:
2012-06
期刊:
ACM Trans. Asian Lang. Inf. Process.
影响因子:
--
通讯作者:
Hongling Wang;Guodong Zhou
Hongling Wang;Guodong Zhou
中科院分区:
其他
文献类型:
--
作者:
Hongling Wang;Guodong Zhou

文献摘要

相似文献

本文提出了一个统一的框架,用于从一组文档中提取标准摘要和更新摘要。特别是,主题建模方法被用于显著性确定和动态建模方法被提出用于冗余控制。在显著性确定的主题建模方法中,我们使用单个向量空间模型通过给定文档或相关语料库的固有主题的相应概率分布来表示各种文本单元,例如词、句子、文档、文档和摘要。因此,我们能够通过它们的主题概率分布来计算任意两个文本单元之间的相似度。在冗余控制的动态建模方法中,我们考虑了摘要与给定文档的相似度,以及句子与摘要的相似度,除了句子与给定文档的相似度外,对于标准摘要和更新摘要,我们还考虑了句子与历史文档或摘要的相似度。对TAC 2008年和2009年的英文版的评估显示了令人鼓舞的结果,特别是动态建模方法在消除给定文档中的冗余方面。最后,我们将该框架扩展到中文多文档摘要中,实验证明了该框架的有效性。
This article presents a unified framework for extracting standard and update summaries from a set of documents. In particular, a topic modeling approach is employed for salience determination and a dynamic modeling approach is proposed for redundancy control. In the topic modeling approach for salience determination, we represent various kinds of text units, such as word, sentence, document, documents, and summary, using a single vector space model via their corresponding probability distributions over the inherent topics of given documents or a related corpus. Therefore, we are able to calculate the similarity between any two text units via their topic probability distributions. In the dynamic modeling approach for redundancy control, we consider the similarity between the summary and the given documents, and the similarity between the sentence and the summary, besides the similarity between the sentence and the given documents, for standard summarization while for update summarization, we also consider the similarity between the sentence and the history documents or summary. Evaluation on TAC 2008 and 2009 in English language shows encouraging results, especially the dynamic modeling approach in removing the redundancy in the given documents. Finally, we extend the framework to Chinese multi-document summarization and experiments show the effectiveness of our framework.