Learning from users' querying experience on intranets

Learning from users' querying experience on intranets
复制标题

借鉴用户内网查询体验

DOI:
10.1145/2187980.2188197
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
Adeyanju I
Adeyanju I
中科院分区:
--
文献类型:
--
作者:
Adeyanju I

文献摘要

参考文献

相似文献

查询推荐正在成为网络搜索引擎的一个常见功能,尤其是那些上下文更加严格的内联网搜索引擎。这是因为它可以支持用户使用最合适的查询词在更短的时间内找到相关信息。选择推荐查询通常是通过挖掘网络文档或先前用户的搜索日志来完成的。我们建议通过组合两个模型来集成这些方法,即概念层次结构(通常根据内联网文档构建)和查询流程图(通常根据搜索日志构建)。但是,我们根据从搜索日志子集(训练集)中提取的术语构建概念层次结构模型,因为这些术语比从集合中提取的任何概念更能代表该领域的用户视图。然后,我们通过合并来自用户搜索日志的另一个子集(测试集)的查询细化来不断调整模型。此过程意味着学习或重用先前用户的查询经验来为新的但相似的用户查询推荐查询。适应权重是从使用相同日志构建的查询流程图中提取的。我们使用从学术机构内联网及其搜索日志爬取的文档来评估我们的混合模型。然后将混合模型与分别从相同集合和搜索日志构建的概念层次模型和查询流程图进行比较。我们还测试了各种策略,用于结合搜索日志中的信息以及查询细化后点击文档的频率。在学年的两个不同时期进行测试时,我们的混合模型显着优于概念层次模型和查询流程图。我们打算用来自另一个机构的文档和搜索日志进一步验证我们的实验,并设计更好的策略来从混合模型中选择推荐查询。
Query recommendation is becoming a common feature of web search engines especially those for Intranets where the context is more restrictive. This is because of its utility for supporting users to find relevant information in less time by using the most suitable query terms. Selection of queries for recommendation is typically done by mining web documents or search logs of previous users. We propose the integration of these approaches by combining two models namely the concept hierarchy, typically built from an Intranet's documents, and the query flow graph, typically built from search logs. However, we build our concept hierarchy model from terms extracted from a subset (training set) of search logs since these are more representative of the user view of the domain than any concepts extracted from the collection. We then continually adapt the model by incorporating query refinements from another subset (test set) of the user search logs. This process implies learning from or reusing previous users' querying experience to recommend queries for a new but similar user query. The adaptation weights are extracted from a query flow graph built with the same logs. We evaluated our hybrid model using documents crawled from the Intranet of an academic institution and its search logs. The hybrid model was then compared to a concept hierarchy model and query flow graph built from the same collection and search logs respectively. We also tested various strategies for combining information in the search logs with respect to the frequency of clicked documents after query refinement. Our hybrid model significantly outperformed the concept hierarchy model and query flow graph when tested over two different periods of the academic year. We intend to further validate our experiments with documents and search logs from another institution and devise better strategies for selecting queries for recommendation from the hybrid model.
使用具有最大信息增益的词汇亲和力自动查询优化
DOI: --
发表时间: 2002
期刊: Annual International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子: --
作者:
David Carmel;E. Farchi;Yael Petruschka;A. Soffer
通讯作者: A. Soffer
术语建议装置的分层方法
DOI: --
发表时间: 2002
期刊: Annual International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子: --
作者:
Hideo Joho;M. Sanderson;M. Beaulieu
通讯作者: M. Beaulieu
AutoEval:一种使用查询日志评估查询建议的评估方法
DOI: --
发表时间: 2011
期刊: European Conference on Information Retrieval
影响因子: --
作者:
M. Albakour;Udo Kruschwitz;Nikolaos Nanas;Yunhyong Kim;D. Song;Maria Fasli;A. Roeck
通讯作者: A. Roeck
DOI: --
发表时间: 2006
期刊: J. Assoc. Inf. Sci. Technol.
影响因子: --
作者:
Ryen W. White;I. Ruthven
通讯作者: I. Ruthven