HASTA: A Hierarchical-Grid Clustering Algorithm with Data Field

HASTA: A Hierarchical-Grid Clustering Algorithm with Data Field
复制标题

DOI:
10.4018/ijdwm.2014040103
复制
发表时间:
2014-04
期刊:
Int. J. Data Warehous. Min.
影响因子:
--
通讯作者:
Shuliang Wang;Yasen Chen
Shuliang Wang;Yasen Chen
中科院分区:
其他
文献类型:
--
作者:
Shuliang Wang;Yasen Chen

文献摘要

被引文献

相似文献

本文提出了一种新的聚类算法--基于数据场的HASTA层次网格聚类算法,通过将数据对象分配到网格中,将数据集建模为一个数据场。HASTA的聚类中心被定义为定位局部势的最大值。HASTA算法通过分析势值的一阶偏导数来识别团簇的边缘,从而可以检测出任意形状团簇的完整尺寸。实验结果表明,HASTA算法在不同的数据集上都有很好的性能,可以在噪声环境下发现任意形状的聚类。除此之外,HASTA并不强制用户预设数据集内聚类的确切数量。此外,HASTA对数据输入的顺序不敏感。HASTA算法的时间复杂度达到了On,这些优点将为大数据挖掘带来潜在的好处。
In this paper, a novel clustering algorithm, HASTA HierArchical-grid cluStering based on daTA field, is proposed to model the dataset as a data field by assigning all the data objects into qusantized grids. Clustering centers of HASTA are defined to locate where the maximum value of local potential is. Edges of cluster in HASTA are identified by analyzing the first-order partial derivative of potential value, thus the full size of arbitrary shaped clusters can be detected. The experimented case demonstrates that HASTA performs effectively upon different datasets and can find out clusters of arbitrary shapes in noisy circumstance. Besides those, HASTA does not force users to preset the exact amount of clusters inside dataset. Furthermore, HASTA is insensitive to the order of data input. The time complexity of HASTA achieves On. Those advantages will potentially benefit the mining of big data.