Query similarity by projecting the query-flow graph

Query similarity by projecting the query-flow graph
复制标题

DOI:
10.1145/1835449.1835536
复制
发表时间:
2010-07
期刊:
Proceedings of the 33rd international ACM SIGIR conference on Research and development in information retrieval
影响因子:
--
通讯作者:
Ilaria Bordino;C. Castillo;D. Donato;A. Gionis
Ilaria Bordino;C. Castillo;D. Donato;A. Gionis
中科院分区:
其他
文献类型:
--
作者:
Ilaria Bordino;C. Castillo;D. Donato;A. Gionis

文献摘要

被引文献

相似文献

定义查询之间的相似性度量是一个有趣而困难的问题。可靠的查询相似度度量可用于查询推荐、查询扩展和广告等各种应用程序。在本文中,我们利用查询日志中的信息来开发查询之间语义相似度的度量。我们的方法依赖于查询流图的概念。查询流图聚合了来自许多用户的查询重新表述:图中的节点表示查询,如果两个查询可能作为同一搜索目标的一部分出现,则将它们连接起来。我们的查询相似度度量是通过在低维欧几里德空间上投影图(或其适当的子图)来获得的。我们的实验表明,我们获得的度量捕获了查询之间的语义相似度的概念,这对于多样化的查询推荐是有用的。
Defining a measure of similarity between queries is an interesting and difficult problem. A reliable query-similarity measure can be used in a variety of applications such as query recommendation, query expansion, and advertising. In this paper, we exploit the information present in query logs in order to develop a measure of semantic similarity between queries. Our approach relies on the concept of the query-flow graph. The query-flow graph aggregates query reformulations from many users: nodes in the graph represent queries, and two queries are connected if they are likely to appear as part of the same search goal. Our query similarity measure is obtained by projecting the graph (or appropriate subgraphs of it) on a low-dimensional Euclidean space. Our experiments show that the measure we obtain captures a notion of semantic similarity between queries and it is useful for diversifying query recommendations.