Learning to extract cross-session search tasks

Learning to extract cross-session search tasks
复制标题

DOI:
10.1145/2488388.2488507
复制
发表时间:
2013-05
期刊:
Proceedings of the 22nd international conference on World Wide Web
影响因子:
--
通讯作者:
Hongning Wang;Yang Song;Ming-Wei Chang;Xiaodong He;Ryen W. White;Wei Chu
Hongning Wang;Yang Song;Ming-Wei Chang;Xiaodong He;Ryen W. White;Wei Chu
中科院分区:
其他
文献类型:
--
作者:
Hongning Wang;Yang Song;Ming-Wei Chang;Xiaodong He;Ryen W. White;Wei Chu

文献摘要

被引文献

相似文献

搜索任务由一系列满足相同信息需求的搜索查询组成,最近被认为是建模用户搜索意图的精确原子单元。该领域的大多数先前研究都集中在单个搜索会话中的短期搜索任务上,并且严重依赖于人工注释来进行监督分类模型学习。在这项工作中,我们的目标是通过研究从用户搜索行为中学习到的查询间依赖关系来识别长期或跨会话的搜索任务(超越会话边界)。提出了一种基于潜在结构支持向量机框架的半监督聚类模型,并提出了一套有效的自动标注规则作为弱监督,以减轻人工标注的负担。基于Bing.com的大规模搜索日志的实验结果证实了该模型在识别跨会话搜索任务方面的有效性以及所引入的弱监督信号的实用性。我们的学习模型可以通过搜索日志更全面地了解用户的搜索行为,并促进了对长期任务的专用搜索引擎支持的开发。
Search tasks, comprising a series of search queries serving the same information need, have recently been recognized as an accurate atomic unit for modeling user search intent. Most prior research in this area has focused on short-term search tasks within a single search session, and heavily depend on human annotations for supervised classification model learning. In this work, we target the identification of long-term, or cross-session, search tasks (transcending session boundaries) by investigating inter-query dependencies learned from users' searching behaviors. A semi-supervised clustering model is proposed based on the latent structural SVM framework, and a set of effective automatic annotation rules are proposed as weak supervision to release the burden of manual annotation. Experimental results based on a large-scale search log collected from Bing.com confirms the effectiveness of the proposed model in identifying cross-session search tasks and the utility of the introduced weak supervision signals. Our learned model enables a more comprehensive understanding of users' search behaviors via search logs and facilitates the development of dedicated search-engine support for long-term tasks.