Conceptual Clustering of Text Clusters

Conceptual Clustering of Text Clusters
复制标题

文本簇的概念聚类

DOI:
--
复制
发表时间:
2003
期刊:
--
影响因子:
--
通讯作者:
Gerd Stumme
Gerd Stumme
中科院分区:
--
文献类型:
--
作者:
A. Hotho;Gerd Stumme

文献摘要

被引文献

相似文献

常见的集群技术的缺点是它们不提供对所获得的集群的内涵描述。另一方面,概念聚类技术提供了这样的描述,但众所周知是相当慢的。在本文中,我们讨论了一种将这两种技术结合起来的方法。我们首先使用同义词作为背景知识,通过k-均值的一种变体对文档进行聚类。此聚类将大量文档减少到相对较少的聚类,然后可以在第二步中从概念上对其进行聚类。
Common clustering techniques have the disadvantage that they do not provide intensional descriptions of the clusters obtained. Conceptual Clustering techniques, on the other hand, provide such descriptions, but are known to be rather slow. In this paper, we discuss a way of combining both techniques. We first cluster the documents by a variant of k–Means, using a thesaurus as background knowledge. This clustering reduces the large number of documents to a relatively small number of clusters, which can then be clustered conceptually in the second step.