Word sense discrimination in information retrieval: A spectral clustering-based approach

Word sense discrimination in information retrieval: A spectral clustering-based approach
复制标题

DOI:
10.1016/j.ipm.2014.10.007
复制
发表时间:
2015-03-01
影响因子:
8.6
通讯作者:
Popescu, Marius
Popescu, Marius
中科院分区:
计算机科学1区
文献类型:
--
作者:
Chifu, Adrian-Gabriel;Hristea, Florentina;Popescu, Marius

文献摘要

被引文献

相似文献

词义歧义已被认为是信息检索(IR)系统精度差的一个原因。词义消歧和区分方法已被定义,以帮助系统选择应检索与不明确查询相关的哪些文档。然而,唯一对 IR 中的词义辨别或消歧显示出真正好处的方法通常是有监督的方法。在本文中,我们提出了一种新的无监督方法,该方法在 IR 中使用词义辨别。我们开发的方法基于谱聚类,并通过增强与目标查询语义相似的文档来重新排序最初检索的文档列表。对于几个 TREC 特别集合,我们表明我们的方法在包含不明确术语的查询的情况下非常有用。我们有兴趣分别在检索 5、10 和 30 个文档(P@5、P@10、P@30)后提高精度水平。我们表明,与当前最先进的基线相比,精度可以提高 8%。我们还关注性能不佳的查询。 (C) 2014 Elsevier Ltd. 保留所有权利。
Word sense ambiguity has been identified as a cause of poor precision in information retrieval (IR) systems. Word sense disambiguation and discrimination methods have been defined to help systems choose which documents should be retrieved in relation to an ambiguous query. However, the only approaches that show a genuine benefit for word sense discrimination or disambiguation in IR are generally supervised ones. In this paper we propose a new unsupervised method that uses word sense discrimination in IR. The method we develop is based on spectral clustering and reorders an initially retrieved document list by boosting documents that are semantically similar to the target query. For several TREC ad hoc collections we show that our method is useful in the case of queries which contain ambiguous terms. We are interested in improving the level of precision after 5, 10 and 30 retrieved documents (P@5, P@10, P@30) respectively. We show that precision can be improved by 8% above current state-of-the-art baselines. We also focus on poor performing queries. (C) 2014 Elsevier Ltd. All rights reserved.