Data Field for Hierarchical Clustering

Data Field for Hierarchical Clustering
复制标题

层次聚类的数据字段

DOI:
10.4018/jdwm.2011100103
复制
发表时间:
2011-10-01
影响因子:
1.2
通讯作者:
Li, Deren
Li, Deren
中科院分区:
计算机科学4区
文献类型:
--
作者:
Wang, Shuliang;Gan, Wenyan;Li, Deren

文献摘要

被引文献

相似文献

提出了通过模拟数据对象之间的相互作用和相对运动来对数据对象进行分组的层次聚类方法。受物理空间场的启发,提出了模拟核场的数据场,用以说明数据空间中对象之间的相互作用。在数据场中,许多数据对象上的等势线的自组织过程发现了它们的层次聚类特征。在聚类过程中,首先生成一个随机样本来优化影响因子。然后估计数据对象的质量以选择具有非零质量的核心数据对象。该算法以核心数据对象为初始聚类,逐层迭代合并,具有良好的性能。实例研究的结果表明,该数据场能够在不需要用户指定参数的情况下对不同大小、形状或粒度的对象进行层次聚类,并考虑聚类内的对象特征并去除噪声数据中的离群点。比较表明,数据域聚类比K-Means、Birch、CURE和VARERON算法表现更好。
In this paper, data field is proposed to group data objects via simulating their mutual interactions and opposite movements for hierarchical clustering. Enlightened by the field in physical space, data field to simulate nuclear field is presented to illuminate the interaction between objects in data space. In the data field, the self-organized process of equipotential lines on many data objects discovers their hierarchical clustering-characteristics. During the clustering process, a random sample is first generated to optimize the impact factor. The masses of data objects are then estimated to select core data object with nonzero masses. Taking the core data objects as the initial clusters, the clusters are iteratively merged hierarchy by hierarchy with good performance. The results of a case study show that the data field is capable of hierarchical clustering on objects varying size, shape or granularity without user-specified parameters, as well as considering the object features inside the clusters and removing the outliers from noisy data. The comparisons illustrate that the data field clustering performs better than K-means, BIRCH, CURE, and CHAMELEON.