Extensions to the k-means algorithm for clustering large data sets with categorical values

Extensions to the k-means algorithm for clustering large data sets with categorical values
复制标题

DOI:
10.1023/a:1009769707641
复制
发表时间:
1998-09-01
影响因子:
4.8
通讯作者:
Huang, ZX
Huang, ZX
中科院分区:
计算机科学3区
文献类型:
--
作者:
Huang, ZX

文献摘要

被引文献

相似文献

k-means 算法以其在大型数据集聚类方面的效率而闻名。但是,仅处理数值会导致其无法用于对包含分类值的现实世界数据进行聚类。在本文中,我们提出了两种算法,将 k 均值算法扩展到分类域以及具有混合数值和分类值的域。 k-modes算法使用简单的匹配相异性度量来处理分类对象,用众数代替聚类的均值,并在聚类过程中使用基于频率的方法更新众数,以最小化聚类成本函数。通过这些扩展,k 模式算法能够以类似于 k 均值的方式对分类数据进行聚类。 k-prototypes算法通过定义组合相异性度量,进一步集成了k-means和k-modes算法,以允许对由混合数字和分类属性描述的对象进行聚类。我们使用众所周知的大豆病害和信用审批数据集来演示两种算法的聚类性能。我们对两个包含 50 万个对象的现实世界数据集进行的实验表明,这两种算法在对大型数据集进行聚类时非常有效,这对于数据挖掘应用程序至关重要。
The k-means algorithm is well known for its efficiency in clustering large data sets. However, working only on numeric values prohibits it from being used to cluster real world data containing categorical values. In this paper we present two algorithms which extend the k-means algorithm to categorical domains and domains with mixed numeric and categorical values. The k-modes algorithm uses a simple matching dissimilarity measure to deal with categorical objects, replaces the means of clusters with modes, and uses a frequency-based method to update modes in the clustering process to minimise the clustering cost function. With these extensions the k-modes algorithm enables the clustering of categorical data in a fashion similar to k-means. The k-prototypes algorithm, through the definition of a combined dissimilarity measure, further integrates the k-means and k-modes algorithms to allow for clustering objects described by mixed numeric and categorical attributes. We use the well known soybean disease and credit approval data sets to demonstrate the clustering performance of the two algorithms. Our experiments on two real world data sets with half a million objects each show that the two algorithms are efficient when clustering large data sets, which is critical to data mining applications.