Validity index for crisp and fuzzy clusters

Validity index for crisp and fuzzy clusters
复制标题

DOI:
10.1016/j.patcog.2003.06.005
复制
发表时间:
2004-03-01
影响因子:
8
通讯作者:
Maulik, U
Maulik, U
中科院分区:
计算机科学1区
文献类型:
--
作者:
Pakhira, MK;Bandyopadhyay, S;Maulik, U

文献摘要

被引文献

相似文献

本文提出了一种聚类有效性指标及其模糊化方法,该指标可以用来衡量数据集不同分区上聚类的优劣。这个索引的最大值,称为PBM-index,在整个层次结构中提供最佳分区。该指数被定义为一个产品的三个因素,最大化,确保形成一个小数目的紧凑的集群之间的至少两个集群大的分离。我们使用了k-means和期望最大化算法作为底层的清晰聚类技术。对于模糊聚类,我们利用了著名的模糊c均值算法。结果表明,PBM指数在适当地确定集群的数量的优越性,相比其他三个众所周知的措施,戴维斯-博尔丁指数,邓恩指数和谢-贝尼指数,提供了几个人工和现实生活中的数据集。(C)2003模式识别学会。由爱思唯尔有限公司出版。保留所有权利。
In this article, a cluster validity index and its fuzzification is described, which can provide a measure of goodness of clustering on different partitions of a data set. The maximum value of this index, called the PBM-index, across the hierarchy provides the best partitioning. The index is defined as a product of three factors, maximization of which ensures the formation of a small number of compact clusters with large separation between at least two clusters. We have used both the k-means and the expectation maximization algorithms as underlying crisp clustering techniques. For fuzzy clustering, we have utilized the well-known fuzzy c-means algorithm. Results demonstrating the superiority of the PBM-index in appropriately determining the number of clusters, as compared to three other well-known measures, the Davies-Bouldin index, Dunn's index and the Xie-Beni index, are provided for several artificial and real-life data sets. (C) 2003 Pattern Recognition Society. Published by Elsevier Ltd. All rights reserved.