Estimation of statistical translation models based on mutual information for ad hoc information retrieval

Estimation of statistical translation models based on mutual information for ad hoc information retrieval
复制标题

DOI:
10.1145/1835449.1835505
复制
发表时间:
2010-07
期刊:
Proceedings of the 33rd international ACM SIGIR conference on Research and development in information retrieval
影响因子:
--
通讯作者:
Maryam Karimzadehgan;ChengXiang Zhai
Maryam Karimzadehgan;ChengXiang Zhai
中科院分区:
其他
文献类型:
--
作者:
Maryam Karimzadehgan;ChengXiang Zhai

文献摘要

被引文献

相似文献

作为一种在信息检索中捕获单词语义关系的原则方法,已证明统计翻译模型胜过依赖于查询和文档中单词的完全匹配的简单文档语言模型。将翻译模型应用于临时信息检索的主要挑战是在没有培训数据的情况下估算翻译模型。现有工作依赖于基于文档收集生成的合成查询培训。但是,此方法在计算上很昂贵,并且没有良好的查询单词覆盖范围。在本文中,我们提出了一种基于单词之间的归一化相互信息估算翻译模型的替代方法,这在计算上的昂贵较不昂贵,并且比合成查询的估计方法更好地覆盖了查询单词。我们还建议将估计的翻译概率正规化,以确保足够的概率质量质量进行自我翻译。实验结果表明,所提出的基于信息的估计方法不仅更有效,而且比基于合成查询的方法更有效,并且可以与伪重复反馈相结合,以进一步提高检索准确性。结果还表明,提出的正则化策略是有效的,可以提高基于合成查询的估计和基于信息的估计的检索准确性。
As a principled approach to capturing semantic relations of words in information retrieval, statistical translation models have been shown to outperform simple document language models which rely on exact matching of words in the query and documents. A main challenge in applying translation models to ad hoc information retrieval is to estimate a translation model without training data. Existing work has relied on training on synthetic queries generated based on a document collection. However, this method is computationally expensive and does not have a good coverage of query words. In this paper, we propose an alternative way to estimate a translation model based on normalized mutual information between words, which is less computationally expensive and has better coverage of query words than the synthetic query method of estimation. We also propose to regularize estimated translation probabilities to ensure sufficient probability mass for self-translation. Experiment results show that the proposed mutual information-based estimation method is not only more efficient, but also more effective than the synthetic query-based method, and it can be combined with pseudo-relevance feedback to further improve retrieval accuracy. The results also show that the proposed regularization strategy is effective and can improve retrieval accuracy for both synthetic query-based estimation and mutual information-based estimation.