Adaptation of the concept hierarchy model with search logs for query recommendation on intranets

Adaptation of the concept hierarchy model with search logs for query recommendation on intranets
复制标题

DOI:
10.1145/2348283.2348288
复制
发表时间:
2012-08
期刊:
--
影响因子:
--
通讯作者:
I. Adeyanju;D. Song;M. Albakour;Udo Kruschwitz;A. Roeck;Maria Fasli
I. Adeyanju;D. Song;M. Albakour;Udo Kruschwitz;A. Roeck;Maria Fasli
中科院分区:
其他
文献类型:
--
作者:
I. Adeyanju;D. Song;M. Albakour;Udo Kruschwitz;A. Roeck;Maria Fasli

文献摘要

被引文献

相似文献

从文档集合中创建的概念层次结构可用于内联网上的查询推荐,方法是根据层次结构中查询的链接强度对术语进行排名。一个主要的限制是,该模型为相同的查询生成相同的建议,并且由于计算成本高,定期从头开始重建它可能效率极低。我们建议通过将搜索日志中的查询细化来适应该模型。我们的直觉是,从集合和搜索日志中构建的概念层次结构提供了相同搜索域的互补概念视图,它们的集成应该不断提高推荐术语的有效性。两种适应方法使用查询日志与点击信息进行比较。我们评估的概念层次结构模型(静态和适应版本)建立从两个学术机构的内联网集合,并将它们与一个国家的最先进的基于日志的查询推荐,查询流图,建立从相同的日志。我们的自适应模型显着优于其静态版本和查询流图测试时,在一段时间内的数据(文件和搜索日志)从两个机构的内联网。
A concept hierarchy created from a document collection can be used for query recommendation on Intranets by ranking terms according to the strength of their links to the query within the hierarchy. A major limitation is that this model produces the same recommendations for identical queries and rebuilding it from scratch periodically can be extremely inefficient due to the high computational costs. We propose to adapt the model by incorporating query refinements from search logs. Our intuition is that the concept hierarchy built from the collection and the search logs provide complementary conceptual views on the same search domain, and their integration should continually improve the effectiveness of recommended terms. Two adaptation approaches using query logs with and without click information are compared. We evaluate the concept hierarchy models (static and adapted versions) built from the Intranet collections of two academic institutions and compare them with a state-of-the-art log-based query recommender, the Query Flow Graph, built from the same logs. Our adaptive model significantly outperforms its static version and the query flow graph when tested over a period of time on data (documents and search logs) from two institutions' Intranets.