Query-dependent learning to rank for cross-lingual information retrieval

Query-dependent learning to rank for cross-lingual information retrieval
复制标题

跨语言信息检索的查询依赖学习排名

DOI:
--
复制
发表时间:
2018
影响因子:
2.7
通讯作者:
A. Shakery
A. Shakery
中科院分区:
计算机科学4区
文献类型:
--
作者:
Elham Ghanbari;A. Shakery

文献摘要

被引文献

相似文献

学习排序(Learning to Rank,LTR)作为一种用于排序任务的机器学习技术,已成为信息检索领域最热门的研究课题之一。跨语言信息检索(CLIR),其中查询的语言不同于文档的语言,是一个重要的IR任务,可以潜在地受益于LTR。本文的重点是LTR在CLIR中的应用。为了根据源语言的查询对目标语言的文档进行排序,我们提出了一种基于LTR的CLIR局部查询依赖方法,称为LQ-DLTR for CLIR。LQ-DLTR的核心思想是利用相似查询的局部特征来构造LTR模型,而不是对所有查询使用单一的全局排名模型。由于查询和文档使用不同的语言,因此LTR中使用的传统功能不能直接用于CLIR。因此,定义适当的特征是使用LTR进行CLIR的主要步骤。本文定义了三类跨语言特征:查询文档特征、文档特征和查询特征。为了定义跨语言特征,使用翻译资源来填补文档和查询之间的差距。然后,在LQ-DLTR的CLIR,一个邻居的相似查询的基础上跨语言的查询特征,创建一个本地的排名功能的LTR算法为一个给定的查询。LTR算法使用两个跨语言特征集,即文档特征和查询文档特征来学习模型。用于识别邻居的查询特征不参与学习阶段。实验结果表明,CLIR的性能提高与使用跨语言的功能,使用几个翻译和他们的概率来计算的功能,相比,在传统的LTR中使用单语功能,翻译查询根据最佳的翻译和忽略的概率。此外,实验结果表明,LQ-DLTR CLIR优于基线信息检索方法和其他LTR排名模型的MAP和NDCG措施。
Learning to rank (LTR), as a machine learning technique for ranking tasks, has become one of the most popular research topics in the area of information retrieval (IR). Cross-lingual information retrieval (CLIR), in which the language of the query is different from the language of the documents, is one of the important IR tasks that can potentially benefit from LTR. Our focus in this paper is the use of LTR for CLIR. To rank the documents in the target language in response to the query in the source language, we propose a local query-dependent approach based on LTR for CLIR, which is called LQ-DLTR for CLIR. The core idea of LQ-DLTR for CLIR is the use of the local characteristics of similar queries to construct the LTR model, instead of using a single global ranking model for all queries. Since the query and the documents are in different languages, the traditional features that are used in LTR cannot be used directly for CLIR. Thus, defining appropriate features is a major step in the use of LTR for CLIR. In this paper, three categories of cross-lingual features are defined: query–document features, document features, and query features. To define the cross-lingual features, translation resources are used to fill the gap between the documents and the queries. Then, in LQ-DLTR for CLIR, a neighborhood of similar queries based on cross-lingual query features is used to create a local ranking function by the LTR algorithm for a given query. The LTR algorithm uses two cross-lingual feature sets, namely document features and query–document features, to learn the model. The query features that are used to identify the neighbors are not involved in the learning phase. Experimental results indicate that the CLIR performance improves with the use of cross-lingual features that use several translations and their probabilities to compute the features, compared to the use of monolingual features in traditional LTR, which translate a query according to the best translation and ignore the probabilities. Moreover, experimental results show that LQ-DLTR for CLIR outperforms the baseline information retrieval methods and other LTR ranking models in terms of the MAP and NDCG measures.