Exploiting Wikipedia for cross-lingual and multilingual information retrieval

Exploiting Wikipedia for cross-lingual and multilingual information retrieval
复制标题

DOI:
10.1016/j.datak.2012.02.003
复制
发表时间:
2012-04-01
影响因子:
2.5
通讯作者:
Cimiano, P.
Cimiano, P.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Sorg, P.;Cimiano, P.

文献摘要

被引文献

相似文献

在本文中,我们展示了如何利用维基百科作为多语言知识资源进行跨语言和多语言信息检索(CLIR/MLIR)。我们描述了一种称为跨语言显式语义分析 (CL-ESA) 的方法,该方法根据显式语间概念对文档进行索引。这些概念被认为是跨语言的和通用的,在我们的例子中对应于维基百科的文章或类别。每个概念都与每种语言的文本签名相关联,该文本签名可用于估计每个概念的特定于语言的术语分布。然后,该知识可用于计算术语和概念之间的关联强度,该关联强度用于将文档映射到概念空间。因此,通过 CL-ESA,我们从词袋模型转向概念袋模型,该模型允许在跨语言和通用概念的向量空间中进行与语言无关的文档表示。我们展示了如何将不同的基于向量的检索模型和术语加权策略与 CL-ESA 结合使用,并通过实验分析不同选择的性能。我们在两个数据集:JRC-Acquis 和 Multext 上评估了该方法的配对检索任务。我们表明,在 MLIR 设置中,CL-ESA 受益于一定程度的抽象,因为使用类别而不是原始 ESA 模型中的文章可以提供更好的结果。 (C) 2012 Elsevier B.V. 保留所有权利。
In this article we show how Wikipedia as a multilingual knowledge resource can be exploited for Cross-Language and Multilingual Information Retrieval (CLIR/MLIR). We describe an approach we call Cross-Language Explicit Semantic Analysis (CL-ESA) which indexes documents with respect to explicit interlingual concepts. These concepts are considered as interlingual and universal and in our case correspond either to Wikipedia articles or categories. Each concept is associated to a text signature in each language which can be used to estimate language-specific term distributions for each concept. This knowledge can then be used to calculate the strength of association between a term and a concept which is used to map documents into the concept space. With CL-ESA we are thus moving from a Bag-Of-Words model to a Bag-Of-Concepts model that allows language-independent document representations in the vector space spanned by interlingual and universal concepts. We show how different vector-based retrieval models and term weighting strategies can be used in conjunction with CL-ESA and experimentally analyze the performance of the different choices. We evaluate the approach on a mate retrieval task on two datasets: JRC-Acquis and Multext. We show that in the MLIR settings, CL-ESA benefits from a certain level of abstraction in the sense that using categories instead of articles as in the original ESA model delivers better results. (C) 2012 Elsevier B.V. All rights reserved.