Simultaneous Clustering and Ensemble

Simultaneous Clustering and Ensemble
复制标题

DOI:
10.1609/aaai.v31i1.10720
复制
发表时间:
2017-02
期刊:
--
影响因子:
--
通讯作者:
Zhiqiang Tao;Hongfu Liu;Y. Fu
Zhiqiang Tao;Hongfu Liu;Y. Fu
中科院分区:
其他
文献类型:
--
作者:
Zhiqiang Tao;Hongfu Liu;Y. Fu

文献摘要

被引文献

相似文献

集成聚类(EC)作为一种有效且稳健的聚类框架出现后,在数据挖掘和机器学习领域受到了广泛关注。通常,EC方法试图将多个基本划分(BP)融合成一个共识划分,其中每个BP是通过对同一数据集执行传统聚类方法获得的。集成聚类的一个有前景的方向是从BP中推导出成对相似性,然后将其转化为一个图划分问题。然而,这些基于图的方法在计算数据点之间的相似性时可能会遭受信息损失,因为它们仅利用了多个BP提供的分类数据,而忽略了原始特征中的丰富信息。这个问题会严重破坏原始特征空间中的潜在聚类结构,从而降低聚类性能。鉴于此,我们提出了一种新颖的同时聚类与集成(SCE)框架来减轻这种不利影响,该框架利用原始特征的相似性矩阵来增强由多个BP汇总的共关联矩阵。针对SCE,通过特征值分解给出了两个简洁的闭式解。在16个真实世界数据集上进行的实验证明了所提出的SCE相对于传统聚类和最先进的集成聚类方法的有效性。此外,还对可能影响我们方法的几个影响因素进行了广泛的探索。
Ensemble Clustering (EC) has gained a great deal of attention throughout the fields of data mining and machine learning, since it emerged as an effective and robust clustering framework. Typically, EC methods try to fuse multiple basic partitions (BPs) into a consensus one, of which each BP is obtained by performing traditional clustering method on the same dataset. One promising direction for ensemble clustering is to derive pairwise similarity from BPs, and then transform it as a graph partition problem. However, these graph based methods may suffer from an information loss when computing the similarity between data points, because they only utilize the categorical data provided by multiple BPs, yet neglect rich information from raw features. This problem can badly undermine the underlying cluster structure in the original feature space, and thus degrade the clustering performance. In light of this, we propose a novel Simultaneous Clustering and Ensemble (SCE) framework to alleviate such detrimental effect, which employs the similarity matrix from raw features to enhance the co-association matrix summarized by multiple BPs. Two neat closed-form solutions given by eigenvalue decomposition are provided for SCE. Experiments conducted on 16 real-world datasets demonstrate the effectiveness of the proposed SCE over the traditional clustering and state-of-the-art ensemble clustering methods. Moreover, several impact factors that may affect our method are also explored extensively.