A new similarity combining reconstruction coefficient with pairwise distance for agglomerative clustering

A new similarity combining reconstruction coefficient with pairwise distance for agglomerative clustering
复制标题

一种结合重建系数和成对距离的凝聚聚类新相似度

DOI:
10.1016/j.ins.2019.08.048
复制
发表时间:
2020-01-01
影响因子:
8.1
通讯作者:
Zhu, William
Zhu, William
中科院分区:
计算机科学1区
文献类型:
--
作者:
Cai, Zhiling;Yang, Xiaofei;Zhu, William

文献摘要

被引文献

相似文献

凝聚聚类是一种主流的聚类方法,可以产生信息丰富的聚类层次结构。凝聚聚类中现有的相似性通常基于成对距离。虽然这种类型的相似性很好地捕获了数据的局部结构,但它对噪声和异常值很敏感,因为它只考虑数据点之间的距离。在本文中,我们通过将对噪声和异常值具有鲁棒性的重建系数与凝聚聚类的成对距离相结合,提出了一种称为 RCPD 的新相似性。我们的新相似性利用了数据点之间的距离和数据点之间的线性表示。因此,RCPD 不仅可以很好地捕获数据的局部结构,而且对噪声和异常值也具有鲁棒性。在 11 个真实世界基准数据集上的实验结果表明,我们的新聚类方法始终优于许多最先进的聚类方法。 (C) 2019 Elsevier Inc. 保留所有权利。
Agglomerative clustering is a mainstream clustering method that can produce an informative hierarchical structure of clusters. Existing similarities in agglomerative clustering are typically based on the pairwise distance. Although this type of similarity captures the local structure of data well, it is sensitive to noise and outliers because it considers only the distance between data points. In this paper, we propose a new similarity called RCPD by combining the reconstruction coefficient, which is robust to noise and outliers, with the pairwise distance for agglomerative clustering. Our new similarity takes advantage of both the distance between data points and the linear representation among data points. Thus, RCPD not only captures the local structure of data well but is also robust to noise and outliers. The experimental results on 11 real-world benchmark datasets show that our new clustering method consistently outperforms many state-of-the-art clustering approaches. (C) 2019 Elsevier Inc. All rights reserved.