Finding Consistent Clusters in Data Partitions

Finding Consistent Clusters in Data Partitions
复制标题

DOI:
10.1007/3-540-48219-9_31
复制
发表时间:
2001-07
期刊:
--
影响因子:
--
通讯作者:
A. Fred
A. Fred
中科院分区:
其他
文献类型:
--
作者:
A. Fred

文献摘要

被引文献

相似文献

给定任意数据集,其中不能假定其具有特定的参数、统计或几何结构,不同的聚类算法通常将产生不同的数据分区。事实上,由于依赖于初始化或某些设计参数值的选择,使用单一的聚类算法也可以获得多个分区。本文讨论了在数据分区中发现一致性簇的问题,提出了对多数投票方案中最常见的关联进行分析的方法。通过将数据分区转换成映射相关关联的共关联样本矩阵来执行对聚类结果的组合。然后使用该矩阵来提取底层一致的集群。提出了一种新的聚类算法--投票k-均值算法,并在k-均值聚类的背景下对该方法进行了评估。使用模拟和真实数据的例子显示了该多数投票组合方案如何同时处理选择簇数和对初始化的依赖性的问题。此外,生成的簇不受超球体形状的约束。
Given an arbitrary data set, to which no particular parametrical, statistical or geometrical structure can be assumed, different clustering algorithms will in general produce different data partitions. In fact, several partitions can also be obtained by using a single clustering algorithm due to dependencies on initialization or the selection of the value of some design parameter. This paper addresses the problem of finding consistent clusters in data partitions, proposing the analysis of the most common associations performed in a majority voting scheme. Combination of clustering results are performed by transforming data partitions into a co-association sample matrix, which maps coherent associations. This matrix is then used to extract the underlying consistent clusters. The proposed methodology is evaluated in the context of k-means clustering, a new clustering algorithm –voting-k-means, being presented. Examples, using both simulated and real data, show how this majority voting combination scheme simultaneously handles the problems of selecting the number of clusters, and dependency on initialization. Furthermore, resulting clusters are not constrained to be hyper-spherically shaped.