Mining query subtopics from search log data

Mining query subtopics from search log data
复制标题

DOI:
10.1145/2348283.2348327
复制
发表时间:
2012-08
期刊:
--
影响因子:
--
通讯作者:
Yunhua Hu;Ya-nan Qian;Hang Li;Daxin Jiang;J. Pei;Q. Zheng
Yunhua Hu;Ya-nan Qian;Hang Li;Daxin Jiang;J. Pei;Q. Zheng
中科院分区:
其他
文献类型:
--
作者:
Yunhua Hu;Ya-nan Qian;Hang Li;Daxin Jiang;J. Pei;Q. Zheng

文献摘要

被引文献

相似文献

网络搜索中的大多数查询都是模糊的和多方面的。从搜索日志数据中识别查询的主要意义和方面,本文称之为查询子主题挖掘,是web搜索中的一个非常重要的问题。通过搜索日志分析,我们发现有两个有趣的用户行为现象可以用来识别查询子主题,即“每次搜索一个子主题”和“按关键字澄清子主题”。每个搜索一个子主题意味着如果用户在一个查询中单击多个url,那么被单击的url往往表示相同的意义或方面。通过关键词澄清子主题是指用户经常添加一个或多个关键字来扩展查询,以澄清其搜索意图。因此,关键词往往是指示意义或面。我们提出了一种聚类算法,可以有效地利用这两种现象来自动挖掘查询的主要子主题,其中每个子主题由包含许多url和关键字的集群表示。挖掘的查询子主题可以用于web搜索中的多个任务,我们从聚类和重新排序等搜索结果呈现方面对它们进行评估。结果表明,该聚类算法可以有效地挖掘查询子主题,其F1测度范围在0.896-0.956之间。我们的实验结果表明,使用我们的方法挖掘的子主题可以显着改进用于搜索结果聚类的最先进方法。基于点击数据的实验结果也表明,基于本文方法的搜索结果重排序可以显著提高用户查找信息能力的效率。
Most queries in web search are ambiguous and multifaceted. Identifying the major senses and facets of queries from search log data, referred to as query subtopic mining in this paper, is a very important issue in web search. Through search log analysis, we show that there are two interesting phenomena of user behavior that can be leveraged to identify query subtopics, referred to as `one subtopic per search' and `subtopic clarification by keyword'. One subtopic per search means that if a user clicks multiple URLs in one query, then the clicked URLs tend to represent the same sense or facet. Subtopic clarification by keyword means that users often add an additional keyword or keywords to expand the query in order to clarify their search intent. Thus, the keywords tend to be indicative of the sense or facet. We propose a clustering algorithm that can effectively leverage the two phenomena to automatically mine the major subtopics of queries, where each subtopic is represented by a cluster containing a number of URLs and keywords. The mined subtopics of queries can be used in multiple tasks in web search and we evaluate them in aspects of the search result presentation such as clustering and re-ranking. We demonstrate that our clustering algorithm can effectively mine query subtopics with an F1 measure in the range of 0.896-0.956. Our experimental results show that the use of the subtopics mined by our approach can significantly improve the state-of-the-art methods used for search result clustering. Experimental results based on click data also show that the re-ranking of search result based on our method can significantly improve the efficiency of users' ability to find information.