Last iterate convergence of SGD for Least-Squares in the Interpolation regime
Last iterate convergence of SGD for Least-Squares in the Interpolation regime
复制标题
插值体系中最小二乘法 SGD 的最后一次迭代收敛
DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Nicolas Flammarion
中科院分区:
文献类型:
--
作者:
Aditya Varre;Loucas Pillaud;Nicolas Flammarion
Motivated by the recent successes of neural networks that have the ability to fit the data perfectly and generalize well, we study the noiseless model in the fundamental least-squares setup. We assume that an optimum predictor fits perfectly inputs and outputs $langle heta_* , phi(X)
angle = Y$, where $phi(X)$ stands for a possibly infinite dimensional non-linear feature map. To solve this problem, we consider the estimator given by the last iterate of stochastic gradient descent (SGD) with constant step-size. In this context, our contribution is two fold: (i) from a (stochastic) optimization perspective, we exhibit an archetypal problem where we can show explicitly the convergence of SGD final iterate for a non-strongly convex problem with constant step-size whereas usual results use some form of average and (ii) from a statistical perspective, we give explicit non-asymptotic convergence rates in the over-parameterized setting and leverage a fine-grained parameterization of the problem to exhibit polynomial rates that can be faster than $O(1/T)$. The link with reproducing kernel Hilbert spaces is established.
DOI:
--
发表时间:
2019-05
期刊:
ArXiv
影响因子:
--
作者:
Kwang-Sung Jun;Ashok Cutkosky;Francesco Orabona
通讯作者:
Kwang-Sung Jun;Ashok Cutkosky;Francesco Orabona
DOI:
10.1073/pnas.1903070116
发表时间:
2019-08-06
影响因子:
11.1
作者:
Belkin, Mikhail;Hsu, Daniel;Mandal, Soumik
通讯作者:
Mandal, Soumik