A dissimilarity measure for the k-Modes clustering algorithm

A dissimilarity measure for the k-Modes clustering algorithm
复制标题

k-Modes 聚类算法的相异性度量

DOI:
10.1016/j.knosys.2011.07.011
复制
发表时间:
2012-02-01
影响因子:
8.8
通讯作者:
Dang, Chuangyin
Dang, Chuangyin
中科院分区:
计算机科学1区
文献类型:
--
作者:
Cao, Fuyuan;Liang, Jiye;Dang, Chuangyin

文献摘要

被引文献

相似文献

聚类是一种重要的数据挖掘技术,它根据相似性准则对数据进行划分。分类数据的聚类问题近年来引起了数据挖掘研究界的广泛关注。k-Modes算法作为k-Means算法的扩展,通过用模式代替均值,被广泛应用于分类数据聚类。本文通过实例分析了简单匹配相异性测度和Ng相异性测度的局限性。基于生物和遗传分类学的思想,结合粗糙隶属函数,定义了一种新的k-Modes算法相异性测度。新的相异度度量的一个显著特点是考虑了属性值在整个论域上的分布。对基于新相异性测度的k-Modes算法的收敛性和时间复杂度进行了研究,结果表明该算法能有效地用于大数据集。在人工合成数据集和UCI提供的5个真实的数据集上的对比实验结果表明了新的相异度测度的有效性,特别是在具有生物和遗传分类信息的数据集上。(C)2011 Elsevier B. V.保留所有权利。
Clustering is one of the most important data mining techniques that partitions data according to some similarity criterion. The problems of clustering categorical data have attracted much attention from the data mining research community recently. As the extension of the k-Means algorithm, the k-Modes algorithm has been widely applied to categorical data clustering by replacing means with modes. In this paper, the limitations of the simple matching dissimilarity measure and Ng's dissimilarity measure are analyzed using some illustrative examples. Based on the idea of biological and genetic taxonomy and rough membership function, a new dissimilarity measure for the k-Modes algorithm is defined. A distinct characteristic of the new dissimilarity measure is to take account of the distribution of attribute values on the whole universe. A convergence study and time complexity of the k-Modes algorithm based on new dissimilarity measure indicates that it can be effectively used for large data sets. The results of comparative experiments on synthetic data sets and five real data sets from UCI show the effectiveness of the new dissimilarity measure, especially on data sets with biological and genetic taxonomy information. (C) 2011 Elsevier B.V. All rights reserved.