A top-down approach for density-based clustering using multidimensional indexes

A top-down approach for density-based clustering using multidimensional indexes
复制标题

使用多维索引进行基于密度的聚类的自顶向下方法

DOI:
10.1016/j.jss.2003.08.237
复制
发表时间:
2004
期刊:
J. Syst. Softw.
影响因子:
--
通讯作者:
B. Lee
B. Lee
中科院分区:
--
文献类型:
--
作者:
JaeYoun Hwang;K. Whang;Yang;B. Lee

文献摘要

被引文献

相似文献

随着越来越多的应用程序涉及大量数据,大型数据库上的集群得到了积极的研究。在本文中,我们提出了一种有效的自上而下的基于密度的聚类方法,该方法基于存储在多维索引的索引节点中的密度信息。我们首先基于区域对比度划分的概念提供簇的正式定义。基于这个概念,我们提出了一种新颖的自顶向下聚类算法,通过分支定界剪枝提高效率。对于这种剪枝,我们提出了一种基于稀疏和密集内部区域确定边界的技术,并正式证明了边界的正确性。实验结果表明,与著名的聚类方法BIRCH相比,该方法的运行时间减少了高达96倍。结果还表明,随着数据库大小的增加,性能提升变得更加明显。
Clustering on large databases has been studied actively as an increasing number of applications involve huge amount of data. In this paper, we propose an efficient top-down approach for density-based clustering, which is based on the density information stored in index nodes of a multidimensional index. We first provide a formal definition of the cluster based on the concept of region contrast partition. Based on this notion, we propose a novel top-down clustering algorithm, which improves the efficiency through branch-and-bound pruning. For this pruning, we present a technique for determining the bounds based on sparse and dense internal regions and formally prove the correctness of the bounds. Experimental results show that the proposed method reduces the elapsed time by up to 96 times compared with that of BIRCH, which is a well-known clustering method. The results also show that the performance improvement becomes more marked as the size of the database increases.