A Novel Biased Diversity Ranking Model for Query-Oriented Multi-Document Summarization

A Novel Biased Diversity Ranking Model for Query-Oriented Multi-Document Summarization
复制标题

DOI:
10.4028/www.scientific.net/amm.380-384.2811
复制
发表时间:
2012-11
期刊:
Applied Mechanics and Materials
影响因子:
--
通讯作者:
Kai Lei;Yin Zeng
Kai Lei;Yin Zeng
中科院分区:
其他
文献类型:
--
作者:
Kai Lei;Yin Zeng

文献摘要

相似文献

面向查询的多文档摘要(QMDS)试图通过从目标文档集合中提取句子来生成一段简明的文本,目的不仅是传达该语料库的关键内容,而且满足该查询所表达的信息需求。由于其巨大的应用价值,近几十年来QMDS得到了广泛的研究。对于一个好的摘要来说,有三个属性是至关重要的,即相关性、威望和低冗余度(或所谓的多样性)。遗憾的是,现有的大多数工作要么忽视了对多样性的关注,要么采用了非优化的启发式算法,通常是基于贪婪的语句选举。在流形排序算法和DivRank算法的启发下,针对查询敏感的摘要任务,提出了一种新的偏向多样性排序模型--ManifoldDivRank。我们的算法发现的排名靠前的句子不仅在面向查询方面享有很高的声望,更重要的是它们彼此之间是不相似的。在DUC2005和DUC2006基准数据集上的实验结果证明了该方法的有效性。
Query-oriented multi-document summarization (QMDS) attempts to generate a concise piece of text byextracting sentences from a target document collection, with the aim of not only conveying the key content of that corpus, also, satisfying the information needs expressed by that query. Due to its great applicable value, QMDS has been intensively studied in recent decades. Three properties are supposed crucial for a good summary, i.e., relevance, prestige and low redundancy (orso-called diversity). Unfortunately, most existing work either disregarded the concern of diversity, or handled it with non-optimized heuristics, usually based on greedy sentences election. Inspired by the manifold-ranking process, which deals with query-biased prestige, and DivRank algorithm which captures query-independent diversity ranking, in this paper, we propose a novel biased diversity ranking model, named ManifoldDivRank, for query-sensitive summarization tasks. The top-ranked sentences discovered by our algorithm not only enjoy query-oriented high prestige, more importantly, they are dissimilar with each other. Experimental results on DUC2005and DUC2006 benchmark data sets demonstrate the effectiveness of our proposal.