Exploiting the Web as the multilingual corpus for unknown query translation

Exploiting the Web as the multilingual corpus for unknown query translation
复制标题

利用网络作为未知查询翻译的多语言语料库

DOI:
10.1002/asi.20328
复制
发表时间:
2006
期刊:
J. Assoc. Inf. Sci. Technol.
影响因子:
--
通讯作者:
Lee
Lee
中科院分区:
--
文献类型:
--
作者:
Jenq;Jei;Wen;Lee

文献摘要

被引文献

相似文献

目标语言为了在数字图书馆系统中提供跨语言信息检索(CLIR)服务,开发一个功能强大的查询翻译引擎是非常重要的。这必须能够自动将用户的查询从多种源语言翻译为系统接受的目标语言。传统的跨语言信息检索方法可以分为两类:基于词典的和基于语料库的。基于字典的方法(Ballesteros & Croft,1997)利用一个简单的翻译字典,其中每个查询项被查找并翻译成目标语言中的相应项。然而,术语可能具有多个模糊的含义,并且词典中未涵盖的许多新的未知术语无法翻译。基于语料库的方法(Nie、Isabelle、Simard & Durand,1999年)包含多语言语料库,其中收集不同语言的文本用于翻译目的。无论是可比文本(Fung & Yee,1998年)还是平行文本(Nie等人,1999)可以作为多语种语料库。平行文本包含双语句子,从中可以通过适当的句子对齐方法提取单词或短语翻译(Gale & Church,1993)。可比较的文本在内容上共享相似的概念,但不一定显示用于翻译的对齐对应。由于存在一个基本的单词不匹配问题,即查询和文档可能不共享相同的单词,因此这些方法的基本假设是随着查询变长,问题的严重性趋于降低。查询扩展技术(Ballesteros & Croft,1997)因此可以用来丰富文档中未涵盖的查询术语。然而,这些方法提出了一些基本的困难,希望支持实际的CLIR服务的数字图书馆。首先,大多数现有的数字图书馆只包含单语文本集合,因此,没有双语语料库进行跨语言训练。真实的利用Web作为多语种语料库进行未知查询翻译
target languages. To facilitate a cross-language information retrieval (CLIR) service in digital library systems, it is important to develop a powerful query translation engine. This must be able to automatically translate users’ queries from multiple source languages to the target languages that the systems accept. Conventional approaches to CLIR can be divided into two groups: dictionary-based and corpora-based. A dictionarybased approach (Ballesteros & Croft, 1997) utilizes a simple translation dictionary in which each query term is looked up and translated into its corresponding terms in the target language. However, terms may have multiple ambiguous meanings and many new unknown terms that are not covered in the dictionary cannot be translated. Corpora-based approach (Nie, Isabelle, Simard & Durand, 1999) incorporates multilingual corpus where texts in different languages are collected for translation purpose. Either comparable texts (Fung & Yee, 1998) or parallel texts (Nie et al., 1999) can be used as the multilingual corpus. Parallel texts contain bilingual sentences, from which word or phrase translations can be extracted by appropriate sentence alignment methods (Gale & Church, 1993). Comparable texts share similar concepts in content, but do not necessarily show alignment correspondence for translation. Because there is a fundamental word mismatch problem that queries and documents might not share the same words, the basic assumption of such approaches is that as queries get longer, the severity of the problem tends to decrease. Query expansion techniques (Ballesteros & Croft, 1997) can therefore be used to enrich query terms not covered in documents. However, these approaches present some fundamental difficulties for digital libraries that wish to support practical CLIR services. First, most of existing digital libraries contain only monolingual text collections; hence, there is no bilingual corpus for cross-lingual training. Second, real Exploiting the Web as the Multilingual Corpus for Unknown Query Translation