LexRank: Graph-based lexical centrality as salience in text summarization

LexRank: Graph-based lexical centrality as salience in text summarization
复制标题

DOI:
10.1613/jair.1523
复制
发表时间:
2004-01-01
影响因子:
5
通讯作者:
Radev, DR
Radev, DR
中科院分区:
计算机科学3区
文献类型:
--
作者:
Erkan, G;Radev, DR

文献摘要

被引文献

相似文献

我们介绍了一种基于随机图的方法来计算自然语言处理中文本单元的相对重要性。我们在文本摘要(TS)问题上测试了该技术。抽取式语音识别依赖于句子显著性的概念来识别一个文档或一组文档中最重要的句子。显著性通常是根据特定重要词的存在或与质心伪句的相似度来定义的。我们考虑了一种新的方法,LexRank,基于特征向量中心性的概念来计算句子的重要性。该模型采用基于句内余弦相似度的连通性矩阵作为句子图表示的邻接矩阵。在最近的DUC 2004评估中,我们基于LexRank的系统在多个任务中排名第一。在本文中,我们详细分析了我们的方法,并将其应用于更大的数据集,包括早期DUC评估的数据。讨论了几种利用相似图计算中心性的方法。结果表明,在大多数情况下,基于度的方法(包括LexRank)优于基于质心的方法和其他参与DUC的系统。此外,具有阈值的LexRank方法优于其他基于度的技术,包括连续的LexRank。我们还表明,我们的方法对数据中的噪声非常不敏感,这些噪声可能是由于文档的主题聚类不完美造成的。
We introduce a stochastic graph-based method for computing relative importance of textual units for Natural Language Processing. We test the technique on the problem of Text Summarization (TS). Extractive TS relies on the concept of sentence salience to identify the most important sentences in a document or set of documents. Salience is typically defined in terms of the presence of particular important words or in terms of similarity to a centroid pseudo-sentence. We consider a new approach, LexRank, for computing sentence importance based on the concept of eigenvector centrality in a graph representation of sentences. In this model, a connectivity matrix based on intra-sentence cosine similarity is used as the adjacency matrix of the graph representation of sentences. Our system, based on LexRank ranked in first place in more than one task in the recent DUC 2004 evaluation. In this paper we present a detailed analysis of our approach and apply it to a larger data set including data from earlier DUC evaluations. We discuss several methods to compute centrality using the similarity graph. The results show that degree-based methods (including LexRank) outperform both centroid-based methods and other systems participating in DUC in most of the cases. Furthermore, the LexRank with threshold method outperforms the other degree-based techniques including continuous LexRank. We also show that our approach is quite insensitive to the noise in the data that may result from an imperfect topical clustering of documents.