Towards context-aware search by learning a very large variable length hidden markov model from search logs

Towards context-aware search by learning a very large variable length hidden markov model from search logs
复制标题

DOI:
10.1145/1526709.1526736
复制
发表时间:
2009-04
期刊:
IEEE Trans. Signal Process.
影响因子:
--
通讯作者:
Huanhuan Cao;Daxin Jiang;Jian Pei;Enhong Chen;Hang Li
Huanhuan Cao;Daxin Jiang;Jian Pei;Enhong Chen;Hang Li
中科院分区:
其他
文献类型:
--
作者:
Huanhuan Cao;Daxin Jiang;Jian Pei;Enhong Chen;Hang Li

文献摘要

被引文献

相似文献

从同一会话中的先前查询和点击捕获用户查询的上下文可以帮助理解用户的信息需求。一种上下文感知的文档重排序、查询建议和URL推荐方法可以大大改善用户的搜索体验。在本文中,我们提出了一个通用的方法上下文感知搜索。为了捕获查询的上下文,我们从从日志数据中提取的搜索会话中学习可变长度的隐马尔可夫模型(vlHMM)。虽然数学模型是直观的,但如何从数亿个搜索会话中学习具有数百万个状态的大型vlHMM构成了一个巨大的挑战。我们开发了一种策略,参数初始化的vlHMM学习,可以大大减少在实践中估计的参数的数量。我们还设计了一种方法,分布式vlHMM学习下的映射-减少模型。我们测试我们的方法在一个真实的数据集,包括18亿次查询,26亿次点击,和8.4亿次搜索会话,并评估的有效性的vlHMM学习从真实的数据上的三个搜索应用程序:文档重新排名,查询建议,和URL推荐。实验结果表明,我们的方法是有效的和高效的。
Capturing the context of a user's query from the previous queries and clicks in the same session may help understand the user's information need. A context-aware approach to document re-ranking, query suggestion, and URL recommendation may improve users' search experience substantially. In this paper, we propose a general approach to context-aware search. To capture contexts of queries, we learn a variable length Hidden Markov Model (vlHMM) from search sessions extracted from log data. Although the mathematical model is intuitive, how to learn a large vlHMM with millions of states from hundreds of millions of search sessions poses a grand challenge. We develop a strategy for parameter initialization in vlHMM learning which can greatly reduce the number of parameters to be estimated in practice. We also devise a method for distributed vlHMM learning under the map-reduce model. We test our approach on a real data set consisting of 1.8 billion queries, 2.6 billion clicks, and 840 million search sessions, and evaluate the effectiveness of the vlHMM learned from the real data on three search applications: document re-ranking, query suggestion, and URL recommendation. The experimental results show that our approach is both effective and efficient.