Beyond bags of words: effectively modeling dependence and features in information retrieval

Beyond bags of words: effectively modeling dependence and features in information retrieval
复制标题

DOI:
10.1145/1394251.1394271
复制
发表时间:
2008-06
期刊:
SIGIR Forum
影响因子:
--
通讯作者:
Donald Metzler
Donald Metzler
中科院分区:
其他
文献类型:
--
作者:
Donald Metzler

文献摘要

被引文献

相似文献

当前最先进的信息检索模型将文档和查询视为词袋。人们已经进行了许多尝试来超越这种简单的表示形式。不幸的是,很少有人在广泛的任务和数据集上显示出检索效率的持续改进。在这里,我们提出了一种基于马尔可夫随机场的新的信息检索统计模型。所提出的模型超越了词袋假设,允许将术语之间的依赖关系纳入模型中。这允许在单个模型的保护下轻松组合各种文本和非文本特征。在此框架内,我们探讨了所涉及的理论问题、参数估计、特征选择和查询扩展。我们给出了许多信息检索任务的实验结果,例如即席检索和网络搜索。
Current state of the art information retrieval models treat documents and queries as bags of words. There have been many attempts to go beyond this simple representation. Unfortunately, few have shown consistent improvements in retrieval effectiveness across a wide range of tasks and data sets. Here, we propose a new statistical model for information retrieval based on Markov random fields. The proposed model goes beyond the bag of words assumption by allowing dependencies between terms to be incorporated into the model. This allows for a variety of textual and non-textual features to be easily combined under the umbrella of a single model. Within this framework, we explore the theoretical issues involved, parameter estimation, feature selection, and query expansion. We give experimental results from a number of information retrieval tasks, such as ad hoc retrieval and web search.