A novel spectral clustering algorithm based on neighbor relation and Gaussian kernel function with only one parameter

A novel spectral clustering algorithm based on neighbor relation and Gaussian kernel function with only one parameter
复制标题

DOI:
10.1007/s00500-023-09309-z
复制
发表时间:
2023-10
期刊:
影响因子:
4.1
通讯作者:
Hao Zhou;Zekun Wang;Hongjia Chen;Xiang Wang
Hao Zhou;Zekun Wang;Hongjia Chen;Xiang Wang
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hao Zhou;Zekun Wang;Hongjia Chen;Xiang Wang

文献摘要

相似文献

谱聚类在复杂结构的非凸空间数据集上表现良好,已成为数据聚类的一种普遍选择。谱聚类的效果取决于相似图矩阵的构造。为了进一步提高聚类性能,本文提出了一种新的基于邻域关系的相似性度量函数。该方法被称为SC-NR。它使用高斯核函数来衡量两个对象之间的相似性。由于欧氏距离不能完全反映数据之间的关系,该方法在两点之间的距离上增加了一个与最近邻顺序相关的权重。加权欧氏距离更好地表达了相似性。在实验中,我们通过外部指标,即聚类准确度(ACC),归一化互信息(NMI),和F-测度的方法与以前的作品进行了比较。通过与现有方法的比较,证明了该算法的优越性。实验包括六个合成数据集和十二个真实世界的数据集。例如,在PenDigits数据集中,F-measure指标比当前算法高16.50%。
Spectral clustering has become a prevalent option for data clustering as it performs well on non-convex spatial datasets with sophisticated structures. The spectral clustering effects depend on the construction of the similarity graph matrix. In this paper, in order to further enhance the clustering performance, we propose a novel similarity measure function based on neighbor relations. The proposed method is called SC-NR. It uses the Gaussian kernel function to measure the similarity between two objects. Since Euclidean distance cannot fully reflect the relation between data, this method adds a weight related to the order of nearest neighbors to the distance between two points. The similarity is better expressed by weighted-Euclidean distance. In experiments, we compared the proposed method with the previous works via the external indexes, that is, clustering accuracy (ACC), normalized mutual information (NMI), and F-measure. The comparison of indexes with state-of-the-art methods demonstrates the superiority of our algorithm. The experiment includes six synthetic datasets and twelve real-world datasets. For instance, in the PenDigits dataset F-measure metric is 16.50% higher than the current algorithms.