Combining multiple clusterings using evidence accumulation

Combining multiple clusterings using evidence accumulation
复制标题

DOI:
10.1109/tpami.2005.113
复制
发表时间:
2005-06-01
影响因子:
23.6
通讯作者:
Jain, AK
Jain, AK
中科院分区:
计算机科学1区
文献类型:
--
作者:
Fred, ALN;Jain, AK

文献摘要

被引文献

相似文献

我们探讨了证据积累的想法(EAC),以结合多个聚类的结果。首先,产生了集合集合A集合A集合。给定数据集(d维度中的n个对象或模式),生成数据分区的不同方式是:1)应用不同的聚类算法和2)应用具有不同参数或初始化值的相同聚类算法。此外,不同数据表示(特征空间)和聚类算法的组合也可以提供多种显着不同的数据分区。鉴于聚类集合中的各个分区,我们提出了一个简单的框架来提取一致的聚类。根据EAC概念,每个分区都被视为数据组织的独立证据,基于投票机制将单个数据分区组合在一起,以在N模式之间生成新的N X N相似性矩阵。 N模式的最终数据分区是通过在此矩阵上应用层次集聚集算法获得的。我们基于数据分区之间的共同信息的概念,开发了一个理论框架,用于分析提出的聚类组合策略及其评估。使用自举技术评估结果的稳定性。介绍了基于K-均值聚类算法的拆分和合并策略的基于证据积累的聚类算法的详细讨论。该方法对几个合成和实际数据集的实验结果与其他组合策略进行了比较,并与众所周知的聚类算法产生的个别聚类结果进行了比较。
We explore the idea of evidence accumulation (EAC) for combining the results of multiple clusterings. First, a clustering ensemble-a set of object partitions, is produced. Given a data set (n objects or patterns in d dimensions), different ways of producing data partitions are: 1) applying different clustering algorithms and 2) applying the same clustering algorithm with different values of parameters or initializations. Further, combinations of different data representations (feature spaces) and clustering algorithms can also provide a multitude of significantly different data partitionings. We propose a simple framework for extracting a consistent clustering, given the various partitions in a clustering ensemble. According to the EAC concept, each partition is viewed as an independent evidence of data organization, individual data partitions being combined, based on a voting mechanism, to generate a new n x n similarity matrix between the n patterns. The final data partition of the n patterns is obtained by applying a hierarchical agglomerative clustering algorithm on this matrix. We have developed a theoretical framework for the analysis of the proposed clustering combination strategy and its evaluation, based on the concept of mutual information between data partitions. Stability of the results is evaluated using bootstrapping techniques. A detailed discussion of an evidence accumulation-based clustering algorithm, using a split and merge strategy based on the K-means clustering algorithm, is presented. Experimental results of the proposed method on several synthetic and real data sets are compared with other combination strategies, and with individual clustering results produced by well-known clustering algorithms.