Exploiting the Web as the multilingual corpus for unknown query translation
Exploiting the Web as the multilingual corpus for unknown query translation
复制标题
利用网络作为未知查询翻译的多语言语料库
DOI:
10.1002/asi.20328
复制
发表时间:
2006
期刊:
影响因子:
--
通讯作者:
Lee
中科院分区:
文献类型:
--
作者:
Jenq;Jei;Wen;Lee
target languages. To facilitate a cross-language information retrieval (CLIR) service in digital library systems, it is important to develop a powerful query translation engine. This must be able to automatically translate users’ queries from multiple source languages to the target languages that the systems accept. Conventional approaches to CLIR can be divided into two groups: dictionary-based and corpora-based. A dictionarybased approach (Ballesteros & Croft, 1997) utilizes a simple translation dictionary in which each query term is looked up and translated into its corresponding terms in the target language. However, terms may have multiple ambiguous meanings and many new unknown terms that are not covered in the dictionary cannot be translated. Corpora-based approach (Nie, Isabelle, Simard & Durand, 1999) incorporates multilingual corpus where texts in different languages are collected for translation purpose. Either comparable texts (Fung & Yee, 1998) or parallel texts (Nie et al., 1999) can be used as the multilingual corpus. Parallel texts contain bilingual sentences, from which word or phrase translations can be extracted by appropriate sentence alignment methods (Gale & Church, 1993). Comparable texts share similar concepts in content, but do not necessarily show alignment correspondence for translation. Because there is a fundamental word mismatch problem that queries and documents might not share the same words, the basic assumption of such approaches is that as queries get longer, the severity of the problem tends to decrease. Query expansion techniques (Ballesteros & Croft, 1997) can therefore be used to enrich query terms not covered in documents. However, these approaches present some fundamental difficulties for digital libraries that wish to support practical CLIR services. First, most of existing digital libraries contain only monolingual text collections; hence, there is no bilingual corpus for cross-lingual training. Second, real Exploiting the Web as the Multilingual Corpus for Unknown Query Translation