On the Equivalence between Neural Network and Support Vector Machine

On the Equivalence between Neural Network and Support Vector Machine
复制标题

DOI:
--
复制
发表时间:
2021-11
期刊:
--
影响因子:
--
通讯作者:
Yilan Chen;Wei Huang-;Lam M. Nguyen;Tsui-Wei Weng
Yilan Chen;Wei Huang-;Lam M. Nguyen;Tsui-Wei Weng
中科院分区:
其他
文献类型:
--
作者:
Yilan Chen;Wei Huang-;Lam M. Nguyen;Tsui-Wei Weng

文献摘要

被引文献

相似文献

最近的研究表明,通过梯度下降训练的无限宽神经网络(NN)的动力学可以用神经切线核(NTK)\citep{jacot 2018 neural}来表征。在平方损失下,通过梯度下降以无限小的学习率训练的无限宽NN相当于NTK \citep{arora 2019 exact}的核回归。然而,等价性目前仅为岭回归所知,而NN与其他核机器(KM)(例如支持向量机(SVM))之间的等价性仍然未知。因此,在这项工作中,我们建议建立NN和SVM之间的等价性,特别是通过软间隔损失训练的无限宽NN和通过次梯度下降训练的NTK的标准软间隔SVM。我们的主要理论结果包括建立神经网络和一个广泛的家庭的$\ell_2$正则化KM有限宽度的界限,这不能处理以前的工作之间的等价关系,并表明,每一个有限宽度的神经网络训练这样的正则化损失函数是大约一个KM。此外,我们证明了我们的理论可以实现三个实际应用,包括(i)通过相应的KM的NN的\textit{non-vacuous}泛化界;(ii)无限宽度NN的\textit{non-trivial}鲁棒性证书(而现有的鲁棒性验证方法将提供vacuous界);(iii)本质上比以前的核回归更强大的无限宽度NN。我们的实验代码可以在\url{https://github.com/leslie-CH/equiv-nn-svm}上找到。
Recent research shows that the dynamics of an infinitely wide neural network (NN) trained by gradient descent can be characterized by Neural Tangent Kernel (NTK) \citep{jacot2018neural}. Under the squared loss, the infinite-width NN trained by gradient descent with an infinitely small learning rate is equivalent to kernel regression with NTK \citep{arora2019exact}. However, the equivalence is only known for ridge regression currently \citep{arora2019harnessing}, while the equivalence between NN and other kernel machines (KMs), e.g. support vector machine (SVM), remains unknown. Therefore, in this work, we propose to establish the equivalence between NN and SVM, and specifically, the infinitely wide NN trained by soft margin loss and the standard soft margin SVM with NTK trained by subgradient descent. Our main theoretical results include establishing the equivalences between NNs and a broad family of $\ell_2$ regularized KMs with finite-width bounds, which cannot be handled by prior work, and showing that every finite-width NN trained by such regularized loss functions is approximately a KM. Furthermore, we demonstrate our theory can enable three practical applications, including (i) \textit{non-vacuous} generalization bound of NN via the corresponding KM; (ii) \textit{non-trivial} robustness certificate for the infinite-width NN (while existing robustness verification methods would provide vacuous bounds); (iii) intrinsically more robust infinite-width NNs than those from previous kernel regression. Our code for the experiments is available at \url{https://github.com/leslie-CH/equiv-nn-svm}.