K-Means-Based Consensus Clustering: A Unified View

K-Means-Based Consensus Clustering: A Unified View
复制标题

基于 K 均值的共识聚类:统一视图

DOI:
10.1109/tkde.2014.2316512
复制
发表时间:
2015-01-01
影响因子:
8.9
通讯作者:
Chen, Jian
Chen, Jian
中科院分区:
计算机科学2区
文献类型:
--
作者:
Wu, Junjie;Liu, Hongfu;Chen, Jian

文献摘要

被引文献

相似文献

共识聚类的目标是找到一个与现有基本划分尽可能一致的单一划分。共识聚类是一种很有前途的从异构数据中寻找聚类结构的解决方案。作为一种高效的共识聚类方法,基于k均值的聚类方法在文献中得到了广泛的关注,但现有的研究工作仍处于初级阶段,且比较零散。为此,本文对基于k均值的共识聚类(KCC)进行了系统的研究。具体地说,我们首先揭示了效用函数对KCC有效的一个充分必要条件。这有助于在完整和不完整数据集上建立统一的KCC框架。此外,我们还研究了影响KCC性能的一些重要因素,如基本分区的质量和多样性。在各种真实数据集上的实验结果表明,KCC是高效的,在聚类质量方面可以与最先进的方法相媲美。此外,KCC对缺失值较多的不完全基本分区具有较高的鲁棒性。
The objective of consensus clustering is to find a single partitioning which agrees as much as possible with existing basic partitionings. Consensus clustering emerges as a promising solution to find cluster structures from heterogeneous data. As an efficient approach for consensus clustering, the K-means based method has garnered attention in the literature, however the existing research efforts are still preliminary and fragmented. To that end, in this paper, we provide a systematic study of K-means-based consensus clustering (KCC). Specifically, we first reveal a necessary and sufficient condition for utility functions which work for KCC. This helps to establish a unified framework for KCC on both complete and incomplete data sets. Also, we investigate some important factors, such as the quality and diversity of basic partitionings, which may affect the performances of KCC. Experimental results on various realworld data sets demonstrate that KCC is highly efficient and is comparable to the state-of-the-art methods in terms of clustering quality. In addition, KCC shows high robustness to incomplete basic partitionings with many missing values.