Beyond precision@10: clustering the long tail of web search results

Beyond precision@10: clustering the long tail of web search results
复制标题

DOI:
10.1145/2063576.2063910
复制
发表时间:
2011-10
期刊:
--
影响因子:
--
通讯作者:
Benno Stein;Tim Gollub;Dennis Hoppe
Benno Stein;Tim Gollub;Dennis Hoppe
中科院分区:
其他
文献类型:
--
作者:
Benno Stein;Tim Gollub;Dennis Hoppe

文献摘要

相似文献

本文解决了网络搜索结果聚类的用户接受度缺失问题。我们报告了选定的分析,并提出了新的概念来改进现有的结果聚类方法。我们的研究结果概括如下:1。不要与搜索引擎的高点击率竞争。在响应查询时,我们假定搜索引擎根据概率排序原则返回最优结果列表:大多数用户期望的文档被放在顶部并形成结果列表头部。我们认为,对于排名靠前的结果,取代这种既定的结果表示形式是无益的。2. 改进结果列表尾部的文档访问。解决“少数群体”信息需求的文件出现在结果列表尾部的某个位置。特别是对于模棱两可和多方面的查询,我们希望这条尾巴很长,因为许多用户会欣赏不同的文档。在这种情况下,网络搜索结果聚类可以通过将长尾重新组织成特定主题的聚类来提高用户满意度。3. 在构造聚类标签时避免阴影。我们展示了当前聚类技术生成的大多数聚类标签都出现在结果列表头部的片段中——我们称之为阴影效应。在整个结果列表的聚类中,这些标签对于主题组织和导航的价值是有限的。我们提出并分析了一种过滤方法来显著减轻标签阴影效应。
The paper addresses the missing user acceptance of web search result clustering. We report on selected analyses and propose new concepts to improve existing result clustering approaches. Our findings in a nutshell are: 1. Don't compete with a search engine's top hits. In response to a query we presume search engines to return an optimal result list in the sense of the probabilistic ranking principle: documents that are expected by the majority of users are placed on top and form the result list head. We argue that, with respect to the top results, it is not beneficial to replace this established form of result presentation. 2. Improve document access in the result list tail. Documents that address the information need of "minorities" appear at some position in the result list tail. Especially for ambiguous and multi-faceted queries we expect this tail to be long, with many users appreciating different documents. In this situation web search result clustering can improve user satisfaction by reorganizing the long tail into topic-specific clusters. 3. Avoid shadowing when constructing cluster labels. We show that most of the cluster labels that are generated by current clustering technology occur within the snippets of the result list head--an effect which we call shadowing. The value of such labels for topic organization and navigating within a clustering of the entire result list is limited. We propose and analyze a filtering approach to significantly alleviate the label shadowing effect.