A half-split grid clustering algorithm by simulating cell division

A half-split grid clustering algorithm by simulating cell division
复制标题

DOI:
10.1109/ijcnn.2014.6889720
复制
发表时间:
2014-07
期刊:
2014 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Wenxiang Dou;Jinglu Hu
Wenxiang Dou;Jinglu Hu
中科院分区:
其他
文献类型:
--
作者:
Wenxiang Dou;Jinglu Hu

文献摘要

相似文献

聚类是重要的数据挖掘技术之一,主要有基于数据的相似聚类和基于空间的密度网格聚类两种处理方法。在更大、更多元的形状和密度数据集上,后者比前者更有优势。然而,由于现有基于网格的方法的全局划分,当集群密度差异较大时,它们的性能会较差。本文提出了一种新的算法,通过模拟细胞分裂过程,在不同密度区域产生合适的网格空间。算法的时间复杂度为O(n),其中n为数据集中的点数。该算法将应用于常见的变色龙数据集和我们的密度差较大的合成数据集。结果表明,该算法在任何多密度情况下都是有效的,并且在空间优化问题上具有可扩展性。
Clustering, one of the important data mining techniques, has two main processing methods on data-based similarity clustering and space-based density grid clustering. The latter has more advantage than the former on larger and multiple shape and density dataset. However, due to a global partition of existing grid-based methods, they will perform worse when there is a big difference on the density of clusters. In this paper, we propose a novel algorithm that can produces appropriate grid space in different density regions by simulating cell division process. The time complexity of the algorithm is O(n) in which n is number of points in dataset. The proposed algorithm will be applied on popular chameleon datasets and our synthetic datasets with big density difference. The results show our algorithm is effective on any multi-density situation and has scalability on space optimization problems.