Automatic classification of Web queries using very large unlabeled query logs

Automatic classification of Web queries using very large unlabeled query logs
复制标题

DOI:
10.1145/1229179.1229183
复制
发表时间:
2007-04-01
影响因子:
5.6
通讯作者:
Frieder, Ophir
Frieder, Ophir
中科院分区:
计算机科学2区
文献类型:
--
作者:
Beitzel, Steven M.;Jensen, Eric C.;Frieder, Ophir

文献摘要

被引文献

相似文献

对用户查询进行准确的主题分类可以提高通用Web搜索系统的有效性和效率。如果系统必须将查询路由到特定主题和资源受限的后端数据库的子集,则这种分类变得至关重要。成功的查询分类是一个具有挑战性的问题,因为Web查询很短,因此提供的功能很少。这种稀疏性,再加上查询的分布和词汇量不断变化,阻碍了传统的文本分类。我们通过组合多个分类器来解决这个问题,包括在人工分类的频繁查询数据库中的精确查找和部分匹配,通过监督学习训练的线性模型,以及一种基于从大量未标记查询日志中挖掘选择偏好的新方法。我们的方法在不使用外部信息源(如在线Web目录或检索到的页面的内容)的情况下对查询进行分类,使其适用于要求苛刻的操作环境,如大规模Web搜索服务。我们使用一个运行中的Web搜索引擎的大样本查询对我们的方法进行了评估,结果表明,我们的组合方法在保持足够的查准率的同时,比最好的单一方法提高了近40%的召回率。此外,我们将我们的结果与2005年KDD杯的结果进行比较,发现尽管我们的运营受到限制,但我们的表现具有竞争力。这表明,可以在不需要外部信息源的情况下对查询流的很大一部分进行局部分类,从而允许在操作受限的环境中进行部署。
Accurate topical classification of user queries allows for increased effectiveness and efficiency in general-purpose Web search systems. Such classification becomes critical if the system must route queries to a subset of topic-specific and resource-constrained back-end databases. Successful query classification poses a challenging problem, as Web queries are short, thus providing few features. This feature sparseness, coupled with the constantly changing distribution and vocabulary of queries, hinders traditional text classification. We attack this problem by combining multiple classifiers, including exact lookup and partial matching in databases of manually classified frequent queries, linear models trained by supervised learning, and a novel approach based on mining selectional preferences from a large unlabeled query log. Our approach classifies queries without using external sources of information, such as online Web directories or the contents of retrieved pages, making it viable for use in demanding operational environments, such as large-scale Web search services. We evaluate our approach using a large sample of queries from an operational Web search engine and show that our combined method increases recall by nearly 40% over the best single method while maintaining adequate precision. Additionally, we compare our results to those from the 2005 KDD Cup and find that we perform competitively despite our operational restrictions. This suggests it is possible to topically classify a significant portion of the query stream without requiring external sources of information, allowing for deployment in operationally restricted environments.