Document clustering with committees

Document clustering with committees
复制标题

DOI:
10.1145/564376.564412
复制
发表时间:
2002-08
期刊:
--
影响因子:
--
通讯作者:
Patrick Pantel;Dekang Lin
Patrick Pantel;Dekang Lin
中科院分区:
其他
文献类型:
--
作者:
Patrick Pantel;Dekang Lin

文献摘要

被引文献

相似文献

文档聚类在许多信息检索任务中是有用的:文档浏览、组织和查看检索结果、生成类似Yahoo的文档层次结构等。聚类的总体目标是将数据元素分组,使得组内相似性高,组间相似性低。我们提出了一个聚类算法CBC(聚类委员会),产生更高质量的聚类文件聚类任务相比,几个著名的聚类算法。它首先发现一组紧密的集群(高组内相似性),称为委员会,这些集群在相似性空间中分布良好(低组间相似性)。委员会的联合只是所有要素的一个子集。该算法通过将元素分配给它们最相似的委员会来进行。评估集群质量一直是一项艰巨的任务。我们提出了一种新的评估方法,该方法基于输出聚类和手动构建的类(答案键)之间的编辑距离。这种评价措施比以前的评价措施更直观,更容易解释。
Document clustering is useful in many information retrieval tasks: document browsing, organization and viewing of retrieval results, generation of Yahoo-like hierarchies of documents, etc. The general goal of clustering is to group data elements such that the intra-group similarities are high and the inter-group similarities are low. We present a clustering algorithm called CBC (Clustering By Committee) that is shown to produce higher quality clusters in document clustering tasks as compared to several well known clustering algorithms. It initially discovers a set of tight clusters (high intra-group similarity), called committees, that are well scattered in the similarity space (low inter-group similarity). The union of the committees is but a subset of all elements. The algorithm proceeds by assigning elements to their most similar committee. Evaluating cluster quality has always been a difficult task. We present a new evaluation methodology that is based on the editing distance between output clusters and manually constructed classes (the answer key). This evaluation measure is more intuitive and easier to interpret than previous evaluation measures.