Noise Regularizes Over-parameterized Rank One Matrix Recovery, Provably

Noise Regularizes Over-parameterized Rank One Matrix Recovery, Provably
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
--
影响因子:
--
通讯作者:
Tianyi Liu;Yan Li;Enlu Zhou;Tuo Zhao
Tianyi Liu;Yan Li;Enlu Zhou;Tuo Zhao
中科院分区:
其他
文献类型:
--
作者:
Tianyi Liu;Yan Li;Enlu Zhou;Tuo Zhao

文献摘要

被引文献

相似文献

我们研究了噪声在学习过参数化模型的优化算法中的作用。具体地说,我们考虑使用过参数化模型从噪声观测$Y$中恢复R^{d\times d}$中的秩一矩阵$Y^*\。我们通过$XX^\top$来参数化秩一矩阵$Y^*$,其中$X\in R^{d\times d}$。然后,我们表明,在温和的条件下,估计,获得随机扰动梯度下降算法使用平方损失函数,达到$O(\sigma^2/d)$的均方误差,其中$\sigma^2 $是观测噪声的方差。相比之下,没有随机扰动梯度下降获得的估计只有达到$O(\sigma^2)$的均方误差。我们的结果部分地证明了噪声在学习过参数化模型时的隐式正则化效应,并为训练过参数化神经网络提供了新的理解。
We investigate the role of noise in optimization algorithms for learning over-parameterized models. Specifically, we consider the recovery of a rank one matrix $Y^*\in R^{d\times d}$ from a noisy observation $Y$ using an over-parameterization model. We parameterize the rank one matrix $Y^*$ by $XX^\top$, where $X\in R^{d\times d}$. We then show that under mild conditions, the estimator, obtained by the randomly perturbed gradient descent algorithm using the square loss function, attains a mean square error of $O(\sigma^2/d)$, where $\sigma^2$ is the variance of the observational noise. In contrast, the estimator obtained by gradient descent without random perturbation only attains a mean square error of $O(\sigma^2)$. Our result partially justifies the implicit regularization effect of noise when learning over-parameterized models, and provides new understanding of training over-parameterized neural networks.