NTCIR-3 CLIR Experiments at MSRA

NTCIR-3 CLIR Experiments at MSRA
复制标题

MSRA 的 NTCIR-3 CLIR 实验

DOI:
--
复制
发表时间:
2002
期刊:
--
影响因子:
--
通讯作者:
Jianfeng Gao
Jianfeng Gao
中科院分区:
--
文献类型:
--
作者:
Hongzhao He;Jianfeng Gao

文献摘要

被引文献

相似文献

本文介绍了三种统计模型,以解决跨语言信息检索(CLIR)的查询翻译歧义。首先,提出了一个衰减共生模型。它是传统同现模型的扩展,因为它包含一个衰减因子,当项之间的距离增加时,该衰减因子会降低互信息。其次,描述了一个短语翻译模型,旨在检测和翻译未存储在词典中的名词短语。最后,提出了一个三重翻译模型,提供了一种利用语言依赖信息的方法。我们使用这些模型TREC和NTCIR语料库的实验改进。
This paper describes three statistical models for the purpose of resolving query translation ambiguity for cross-language information retrieval (CLIR). First, a decaying co-occurrence model is present. It is an extension of traditional co-occurrence models in that it contains a decaying factor which decreases the mutual information when the distance between the terms increases. Second, a phrase translation model is described aiming to detect and translate noun phrases that are not stored in the dictionary. Finally, a triple translation model is proposed which provides a way of exploiting linguistic dependency information. We show experimentally improvements of using these models on TREC and NTCIR corpus.