A Mixture Model for Clustering Ensembles

A Mixture Model for Clustering Ensembles
复制标题

DOI:
10.1137/1.9781611972740.35
复制
发表时间:
2004
期刊:
--
影响因子:
--
通讯作者:
A. Topchy;Anil K. Jain;W. Punch
A. Topchy;Anil K. Jain;W. Punch
中科院分区:
其他
文献类型:
--
作者:
A. Topchy;Anil K. Jain;W. Punch

文献摘要

被引文献

相似文献

聚类集成已经成为提高无监督分类解的稳健性和稳定性的一种有效方法。然而,从多个分区中找到一致聚类是一个困难的问题,可以从基于图的、组合的或统计的角度来处理。我们提供了一个概率模型的共识,使用有限的混合多项式分布在一个集群空间。使用EM算法找到一个组合划分作为相应的最大似然问题的解。该算法具有良好的可扩展性和易于理解的底层模型,对于大数据集的聚类尤为重要。这项研究比较了EM共识算法和其他融合方法在聚类集成方面的性能。我们还分析了信息不完全情况下的聚类集成,以及丢失聚类标签对总体共识质量的影响。实验结果证明了该方法在大规模真实数据集上的有效性。关键词:无监督学习、聚类集成、一致性函数、混合模型、EM算法。
Clustering ensembles have emerged as a powerful method for improving both the robustness and the stability of unsupervised classification solutions. However, finding a consensus clustering from multiple partitions is a difficult problem that can be approached from graph-based, combinatorial or statistical perspectives. We offer a probabilistic model of consensus using a finite mixture of multinomial distributions in a space of clusterings. A combined partition is found as a solution to the corresponding maximum likelihood problem using the EM algorithm. The excellent scalability of this algorithm and comprehensible underlying model are particularly important for clustering of large datasets. This study compares the performance of the EM consensus algorithm with other fusion approaches for clustering ensembles. We also analyze clustering ensembles with incomplete information and the effect of missing cluster labels on the quality of overall consensus. Experimental results demonstrate the effectiveness of the proposed method on large real-world datasets. keywords: unsupervised learning, clustering ensemble, consensus function, mixture model, EM algorithm.