Performance evaluation of some clustering algorithms and validity indices

Performance evaluation of some clustering algorithms and validity indices
复制标题

DOI:
10.1109/tpami.2002.1114856
复制
发表时间:
2002-12-01
影响因子:
23.6
通讯作者:
Bandyopadhyay, S
Bandyopadhyay, S
中科院分区:
计算机科学1区
文献类型:
--
作者:
Maulik, U;Bandyopadhyay, S

文献摘要

被引文献

相似文献

在这篇文章中,我们评估了三个聚类算法的性能,硬K均值,单链接,和模拟退火(SA)为基础的技术,结合四个集群有效性指标,即Davies-Bouldin指数,邓恩指数,Calinski-Harabasz指数,和最近开发的指数I。基于指数I与Dunn指数之间的关系,理论上给出了指数I的下界估计,从而在具有不同子结构的数据集上得到唯一的硬K-划分.不同的有效性指标和聚类方法在自动发展适当数量的集群的有效性实验证明了人工和现实生活中的数据集与集群的数量从两个到十个不等。一旦确定了适当数量的聚类,基于SA的聚类技术用于将数据适当地划分为所述数量的聚类。
In this article, we evaluate the performance of three clustering algorithms, hard K-Means, single linkage, and a simulated annealing (SA) based technique, in conjunction with four cluster validity indices, namely Davies-Bouldin index, Dunn's index, Calinski-Harabasz index, and a recently developed index I. Based on a relation between the index I and the Dunn's index, a lower bound of the value of the former is theoretically estimated in order to get unique hard K-partition when the data set has distinct substructures. The effectiveness of the different validity indices and clustering methods in automatically evolving the appropriate number of clusters is demonstrated experimentally for both artificial and real-life data sets with the number of clusters varying from two to ten. Once the appropriate number of clusters is determined, the SA-based clustering technique is used for proper partitioning of the data into the said number of clusters.