The optimum clustering framework: implementing the cluster hypothesis

The optimum clustering framework: implementing the cluster hypothesis
复制标题

DOI:
10.1007/s10791-011-9173-9
复制
发表时间:
2012-04-01
期刊:
INFORMATION RETRIEVAL
影响因子:
--
通讯作者:
Gollub, Tim
Gollub, Tim
中科院分区:
其他
文献类型:
--
作者:
Fuhr, Norbert;Lechtenfeld, Marc;Gollub, Tim

文献摘要

被引文献

相似文献

文档聚类提供了在交互式检索中支持用户的潜力,特别是当用户在精确指定其信息需求方面存在问题时。本文提出了最优文档聚类的理论基础。关键思想是基于一组查询进行聚类分析和评估,方法是将与相同查询相关的文档定义为相似的文档。在我们的最佳聚类框架OCF中,有三个组件是必不可少的:(1)一组查询,(2)一个概率检索方法,(3)一个文档相似度度量。在引入适当的有效性度量后,我们根据所考虑的查询文档对的相关概率估计来定义最佳聚类。此外,我们还表明,众所周知的聚类方法隐式地基于这三个组件,但它们对其中一些组件使用启发式设计决策。我们认为,有了我们的框架,开发更好的文档聚类方法的更有针对性的研究成为可能。实验结果证明了我们考虑的潜力。
Document clustering offers the potential of supporting users in interactive retrieval, especially when users have problems in specifying their information need precisely. In this paper, we present a theoretic foundation for optimum document clustering. Key idea is to base cluster analysis and evalutation on a set of queries, by defining documents as being similar if they are relevant to the same queries. Three components are essential within our optimum clustering framework, OCF: (1) a set of queries, (2) a probabilistic retrieval method, and (3) a document similarity metric. After introducing an appropriate validity measure, we define optimum clustering with respect to the estimates of the relevance probability for the query-document pairs under consideration. Moreover, we show that well-known clustering methods are implicitly based on the three components, but that they use heuristic design decisions for some of them. We argue that with our framework more targeted research for developing better document clustering methods becomes possible. Experimental results demonstrate the potential of our considerations.