Finding Stable Clustering for Noisy Data via Structure-Aware Representation

Finding Stable Clustering for Noisy Data via Structure-Aware Representation
复制标题

DOI:
10.1109/bigdata47090.2019.9006431
复制
发表时间:
2019-12
期刊:
2019 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Huiyuan Chen;Jing Li
Huiyuan Chen;Jing Li
中科院分区:
其他
文献类型:
--
作者:
Huiyuan Chen;Jing Li

文献摘要

相似文献

聚类是机器学习中最突出的问题之一。人们提出了多种聚类方法,其中谱聚类受到了极大的关注。然而,在实践中,谱聚类对噪声数据高度敏感,并且通常需要后处理步骤(例如,特征向量的k-均值)来获得聚类指标,这可能不是最优的。此外,由于其特征分解过程,它不能很好地扩展到大规模数据。这里,我们提出了一个结构感知的聚类模型来解决这些问题。为了达到这个目的,通过稀疏加性分解从原始噪声数据中提取高质量的亲和度矩阵,用来逼近理想的聚类结构。然后,我们联合学习高质量的亲和度矩阵以及在统一模型中嵌入的谱-因此,对噪声具有鲁棒性,并且无需任何后处理步骤就可以获得最优的聚类指标。通过考虑亲和力矩阵的拉普拉斯本征角,进一步提高了聚类的稳定性。结果表明,拉普拉斯本征隙越大,聚类结果越稳定。为了有效地计算大型矩阵的特征向量,我们引入了一种加速策略。实验结果表明,该模型在处理含噪数据方面优于已有方法。
Clustering is one of the most prominent topics in machine learning. A multitude of clustering methods have been proposed, among which the spectral clustering has attracted much attention. However, in practice, spectral clustering is highly sensitive to noise data and a post-processing step (e.g., k-means for eigenvectors) is often required to obtain clustering indicators, which may be not optimal. Also, it does not scale well to large-scale data due to its eigen-decomposition procedures.Here we propose a structure-aware clustering model to address those issues. To achieve our goal, a high-quality affinity matrix is extracted from the original noisy data by a sparse additive decomposition, which is used to approximate the ideal clustering structure. We then jointly learn the high-quality affinity matrix as well as the spectral embedding in a unified model— thus, being robust to noise and obtaining the optimal clustering indicators without any post-processing steps. We further improve the clustering stability by considering the Laplacian eigengap of the affinity matrix. We show that the larger the Laplacian eigengap, the more stable the clustering results. We introduce a speedup strategy to effectively compute eigenvectors of large matrices. Experimental results demonstrate that the proposed model outperforms existing approaches for noisy data.