Identifying similarities, periodicities and bursts for online search queries

Identifying similarities, periodicities and bursts for online search queries
复制标题

DOI:
10.1145/1007568.1007586
复制
发表时间:
2004-06
期刊:
--
影响因子:
--
通讯作者:
M. Vlachos;Christopher Meek;Zografoula Vagena;D. Gunopulos
M. Vlachos;Christopher Meek;Zografoula Vagena;D. Gunopulos
中科院分区:
其他
文献类型:
--
作者:
M. Vlachos;Christopher Meek;Zografoula Vagena;D. Gunopulos

文献摘要

被引文献

相似文献

提出了几种从MSN搜索引擎的查询日志中挖掘知识的方法。使用查询日志,我们为每个查询词或短语(例如,“感恩节”或“圣诞礼物”)构建时间序列,其中时间序列的元素是一天中发出查询的次数。我们描述的所有方法都使用这种形式的序列,并且通常可以应用于时间序列数据。我们的主要目标是发现语义相似的查询,我们通过识别具有相似需求模式的查询来实现这一目标。利用最佳傅立叶系数和省略分量的能量,我们改进了最先进的时间序列相似性匹配。然后将提取的序列特征组织成有效的度量树索引结构。我们还演示了如何有效和准确地发现时间序列中的重要时期。最后,我们提出了一种简单而有效的识别突发(长期或短期)的方法。利用从序列中提取的突发信息,我们能够有效地对时间序列数据库进行“突发查询”。最后,我们描述了一个工具,该工具使用所描述的方法,并作为MSN查询数据库的交互式探索性数据发现工具。
We present several methods for mining knowledge from the query logs of the MSN search engine. Using the query logs, we build a time series for each query word or phrase (e.g., 'Thanksgiving' or 'Christmas gifts') where the elements of the time series are the number of times that a query is issued on a day. All of the methods we describe use sequences of this form and can be applied to time series data generally. Our primary goal is the discovery of semantically similar queries and we do so by identifying queries with similar demand patterns. Utilizing the best Fourier coefficients and the energy of the omitted components, we improve upon the state-of-the-art in time-series similarity matching. The extracted sequence features are then organized in an efficient metric tree index structure. We also demonstrate how to efficiently and accurately discover the important periods in a time-series. Finally we propose a simple but effective method for identification of bursts (long or short-term). Using the burst information extracted from a sequence, we are able to efficiently perform 'query-by-burst' on the database of time-series. We conclude the presentation with the description of a tool that uses the described methods, and serves as an interactive exploratory data discovery tool for the MSN query database.