SpecNet2: Orthogonalization-free spectral embedding by neural networks

SpecNet2: Orthogonalization-free spectral embedding by neural networks
复制标题

DOI:
10.48550/arxiv.2206.06644
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Ziyu Chen;Yingzhou Li;Xiuyuan Cheng
Ziyu Chen;Yingzhou Li;Xiuyuan Cheng
中科院分区:
其他
文献类型:
--
作者:
Ziyu Chen;Yingzhou Li;Xiuyuan Cheng

文献摘要

相似文献

谱方法通过核矩阵或图拉普拉斯矩阵的特征向量来表示数据点,是无监督数据分析的主要工具。在许多应用场景中,通过可以在批量数据样本上训练的神经网络来参数化谱嵌入,这为实现自动样本外扩展以及计算可扩展性提供了一种有前途的方法。SpectralNet的原始论文(Shaham et al. 2018)采用了这种方法,我们称之为SpecNet1。本文介绍了一种新的神经网络方法,命名为SpecNet2,来计算谱嵌入,优化了特征问题的等价目标,并删除了SpecNet1中的正交化层。SpecNet2还允许通过梯度公式跟踪每个数据点的邻居来分离图形亲和矩阵的行和列的采样。从理论上讲,我们表明,任何新的正交化自由目标的局部极小揭示了领先的特征向量。此外,这个新的正交化自由的目标,使用基于批处理的梯度下降方法的全局收敛性被证明。数值实验证明了SpecNet2在模拟数据和图像数据集上的性能和计算效率的提高。
Spectral methods which represent data points by eigenvectors of kernel matrices or graph Laplacian matrices have been a primary tool in unsupervised data analysis. In many application scenarios, parametrizing the spectral embedding by a neural network that can be trained over batches of data samples gives a promising way to achieve automatic out-of-sample extension as well as computational scalability. Such an approach was taken in the original paper of SpectralNet (Shaham et al. 2018), which we call SpecNet1. The current paper introduces a new neural network approach, named SpecNet2, to compute spectral embedding which optimizes an equivalent objective of the eigen-problem and removes the orthogonalization layer in SpecNet1. SpecNet2 also allows separating the sampling of rows and columns of the graph affinity matrix by tracking the neighbors of each data point through the gradient formula. Theoretically, we show that any local minimizer of the new orthogonalization-free objective reveals the leading eigenvectors. Furthermore, global convergence for this new orthogonalization-free objective using a batch-based gradient descent method is proved. Numerical experiments demonstrate the improved performance and computational efficiency of SpecNet2 on simulated data and image datasets.