Understanding user's query intent with wikipedia

Understanding user's query intent with wikipedia
复制标题

DOI:
10.1145/1526709.1526773
复制
发表时间:
2009-04
期刊:
--
影响因子:
--
通讯作者:
Jian Hu;G. Wang;F. Lochovsky;Jian-Tao Sun-;Zheng Chen
Jian Hu;G. Wang;F. Lochovsky;Jian-Tao Sun-;Zheng Chen
中科院分区:
其他
文献类型:
--
作者:
Jian Hu;G. Wang;F. Lochovsky;Jian-Tao Sun-;Zheng Chen

文献摘要

被引文献

相似文献

了解用户查询背后的意图可以帮助搜索引擎自动将查询路由到相应的垂直搜索引擎,以获得特别相关的内容,从而大大提高用户满意度。查询意图分类问题有三个主要挑战:(1)意图表示;(2)域覆盖和(3)语义解释。目前预测用户意图的方法主要利用机器学习技术。然而,这是困难的,往往需要许多人的努力,以满足所有这些挑战的统计机器学习方法。在本文中,我们提出了一个通用的方法来解决问题的查询意图分类。只需很少的人力,我们的方法就可以通过利用维基百科(最好的人类知识库之一)来发现大量的意图概念。Wikipedia概念被用作意图表示空间,因此,每个意图域被表示为一组Wikipedia文章和类别。通过将查询映射到Wikipedia表示空间来识别任何输入查询的意图。与以前的方法相比,我们提出的方法可以实现更好的覆盖率分类查询的意图域,即使种子意图的例子是非常小的。此外,该方法是非常普遍的,可以很容易地应用到各种意图域。我们证明了这种方法在三个不同的应用,即,旅行、工作和人名。在这三种情况下,只提供了几个种子意图查询。通过与两种基准方法的比较,我们进行了定量评估,实验结果表明,我们的方法在每个意图域都明显优于其他方法。
Understanding the intent behind a user's query can help search engine to automatically route the query to some corresponding vertical search engines to obtain particularly relevant contents, thus, greatly improving user satisfaction. There are three major challenges to the query intent classification problem: (1) Intent representation; (2) Domain coverage and (3) Semantic interpretation. Current approaches to predict the user's intent mainly utilize machine learning techniques. However, it is difficult and often requires many human efforts to meet all these challenges by the statistical machine learning approaches. In this paper, we propose a general methodology to the problem of query intent classification. With very little human effort, our method can discover large quantities of intent concepts by leveraging Wikipedia, one of the best human knowledge base. The Wikipedia concepts are used as the intent representation space, thus, each intent domain is represented as a set of Wikipedia articles and categories. The intent of any input query is identified through mapping the query into the Wikipedia representation space. Compared with previous approaches, our proposed method can achieve much better coverage to classify queries in an intent domain even through the number of seed intent examples is very small. Moreover, the method is very general and can be easily applied to various intent domains. We demonstrate the effectiveness of this method in three different applications, i.e., travel, job, and person name. In each of the three cases, only a couple of seed intent queries are provided. We perform the quantitative evaluations in comparison with two baseline methods, and the experimental results shows that our method significantly outperforms other methods in each intent domain.