On the generalization of learning algorithms that do not converge

On the generalization of learning algorithms that do not converge
复制标题

关于不收敛学习算法的泛化

DOI:
10.48550/arxiv.2208.07951
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
S. Jegelka
S. Jegelka
中科院分区:
--
文献类型:
--
作者:
N. Chandramoorthy;Andreas Loukas;Khashayar Gatmiry;S. Jegelka

文献摘要

参考文献

相似文献

深度学习的泛化分析通常假设训练收敛到一个固定点。但是,最近的结果表明,在实践中,使用随机梯度下降优化的深度神经网络的权重经常无限期振荡。为了减少理论与实践之间的这种差异,本文重点讨论了神经网络的泛化,其训练动态不一定收敛到固定点。我们的主要贡献是提出了统计算法稳定性(SAS)的概念,扩展了经典算法的稳定性,非收敛算法,并研究其连接到泛化。与传统的优化和学习理论观点相比,这种遍历理论方法带来了新的见解。我们证明了学习算法的时间渐近行为的稳定性与其泛化能力有关,并以经验证明了损失动态如何为泛化性能提供线索。我们的研究结果提供了证据,证明即使训练无限期地持续下去,权重也不会收敛,“稳定训练的网络也能更好地泛化”。
Generalization analyses of deep learning typically assume that the training converges to a fixed point. But, recent results indicate that in practice, the weights of deep neural networks optimized with stochastic gradient descent often oscillate indefinitely. To reduce this discrepancy between theory and practice, this paper focuses on the generalization of neural networks whose training dynamics do not necessarily converge to fixed points. Our main contribution is to propose a notion of statistical algorithmic stability (SAS) that extends classical algorithmic stability to non-convergent algorithms and to study its connection to generalization. This ergodic-theoretic approach leads to new insights when compared to the traditional optimization and learning theory perspectives. We prove that the stability of the time-asymptotic behavior of a learning algorithm relates to its generalization and empirically demonstrate how loss dynamics can provide clues to generalization performance. Our findings provide evidence that networks that"train stably generalize better"even when the training continues indefinitely and the weights do not converge.
DOI: --
发表时间: 2020-02
期刊: ArXiv
影响因子: --
作者:
Lingkai Kong;Molei Tao
通讯作者: Lingkai Kong;Molei Tao
两层神经网络的平均场理论:无维数界限和核极限
DOI: --
发表时间: 2019
期刊: Proceedings of the Thirty-Second Conference on Learning Theory
影响因子: --
作者:
Mei, Song;Misiakiewicz, Theodor;Montanari, Andrea
通讯作者: Montanari, Andrea
DOI: 10.1016/j.acha.2021.12.009
发表时间: 2022-04-25
影响因子: 2.5
作者:
Liu, Chaoyue;Zhu, Libin;Belkin, Mikhail
通讯作者: Belkin, Mikhail
DOI: --
发表时间: 2021-06
期刊: ArXiv
影响因子: --
作者:
Andreas Loukas;Marinos Poiitis;S. Jegelka
通讯作者: Andreas Loukas;Marinos Poiitis;S. Jegelka
大学习率抑制同质性:收敛和平衡效应
DOI: --
发表时间: 2022
期刊: The International Conference on Learning Representations
影响因子: --
作者:
Wang, Yuqing;Chen, Minshuo;Zhao, Tuo;Tao, Molei
通讯作者: Tao, Molei