Implicit Sparse Regularization: The Impact of Depth and Early Stopping

Implicit Sparse Regularization: The Impact of Depth and Early Stopping
复制标题

DOI:
--
复制
发表时间:
2021-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Jiangyuan Li;Thanh V. Nguyen;C. Hegde;R. K. Wong
Jiangyuan Li;Thanh V. Nguyen;C. Hegde;R. K. Wong
中科院分区:
其他
文献类型:
--
作者:
Jiangyuan Li;Thanh V. Nguyen;C. Hegde;R. K. Wong

文献摘要

被引文献

相似文献

本文研究了稀疏回归中梯度下降的隐式偏差。我们将二次参数化回归的结果,即深度-2对角线性网络,扩展到更一般的深度- n网络,在更现实的噪声和相关设计设置下。我们证明了提前停止对于梯度下降收敛到一个稀疏模型是至关重要的,这种现象我们称之为隐式稀疏正则化。这一结果与无噪声和不相关设计案例的已知结果形成鲜明对比。我们描述了深度和提前停止的影响,并证明了对于一般深度参数N,在初始化和步长足够小的情况下,提前停止的梯度下降可以实现最小最大最优稀疏恢复。特别是,我们表明,增加深度扩大了工作初始化的规模和早期停止窗口,使这种隐式稀疏正则化效果更有可能发生。
In this paper, we study the implicit bias of gradient descent for sparse regression. We extend results on regression with quadratic parametrization, which amounts to depth-2 diagonal linear networks, to more general depth-N networks, under more realistic settings of noise and correlated designs. We show that early stopping is crucial for gradient descent to converge to a sparse model, a phenomenon that we call implicit sparse regularization. This result is in sharp contrast to known results for noiseless and uncorrelated-design cases. We characterize the impact of depth and early stopping and show that for a general depth parameter N, gradient descent with early stopping achieves minimax optimal sparse recovery with sufficiently small initialization and step size. In particular, we show that increasing depth enlarges the scale of working initialization and the early-stopping window so that this implicit sparse regularization effect is more likely to take place.