An Axis-Shifted Grid-Clustering Algorithm

An Axis-Shifted Grid-Clustering Algorithm
复制标题

轴移网格聚类算法

DOI:
10.6180/jase.2009.12.2.10
复制
发表时间:
2009
期刊:
--
影响因子:
--
通讯作者:
Nien
Nien
中科院分区:
--
文献类型:
--
作者:
Chung;Nancy P. Lin;Nien

文献摘要

被引文献

相似文献

这些空间聚类方法可以分为四类:分区方法,层次方法,基于密度的方法和基于网格的方法。基于网格的聚类算法是一种高效的聚类算法,它将数据空间划分为有限个网格单元,形成网格结构,然后在此网格结构上执行所有聚类操作,将相似的空间对象分组到类中。为了有效地同时进行聚类,减少网格大小和边界对聚类结果的影响,提出了一种新的基于网格的聚类算法--轴移网格聚类算法(Axis-Shifted Grid-Clustering Algorithm,ASGC)。这种新的聚类方法结合了一种新的基于密度网格的聚类与轴移分区策略,以确定在输入数据空间中的高密度区域。其主要思想是在获得由原始网格结构生成的聚类后,在数据空间的每个维度上移动原始网格结构。移动网格结构可以看作是对原始单元大小的动态调整,减少了单元边界的薄弱环节。因此,从该移位网格结构生成的聚类可以用于修正最初获得的聚类。实验结果表明,这种新算法的效果受单元大小的影响比其他基于网格的算法小,并且最多需要一次扫描。
These spatial clustering methods can be classified into four categories: partitioning method, hierarchical method, density-based method and grid-based method. The grid-based clustering algorithm, which partitions the data space into a finite number of cells to form a grid structure and then performs all clustering operations to group similar spatial objects into classes on this obtained grid structure, is an efficient clustering algorithm. To cluster efficiently and simultaneously, to reduce the influences of the size and borders of the cells, a new grid-based clustering algorithm, an Axis-Shifted Grid-Clustering algorithm (ASGC), is proposed in this paper. This new clustering method combines a novel density-grid based clustering with axis-shifted partitioning strategy to identify areas of high density in the input data space. The main idea is to shift the original grid structure in each dimension of the data space after the clusters generated from this original structure have been obtained. The shifted grid structure can be considered as a dynamic adjustment of the size of the original cells and reduce the weakness of borders of cells. And thus, the clusters generated from this shifted grid structure can be used to revise the originally obtained clusters. The experimental results verify that, indeed, the effect of this new algorithm is less influenced by the size of cells than other grid-based ones and requires at most a single scan through the data.