Demystify Hyperparameters for Stochastic Optimization with Transferable Representations

Demystify Hyperparameters for Stochastic Optimization with Transferable Representations
复制标题

DOI:
10.1145/3534678.3539298
复制
发表时间:
2022-08
期刊:
Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
Jianhui Sun;Mengdi Huai;Kishlay Jha;Aidong Zhang
Jianhui Sun;Mengdi Huai;Kishlay Jha;Aidong Zhang
中科院分区:
其他
文献类型:
--
作者:
Jianhui Sun;Mengdi Huai;Kishlay Jha;Aidong Zhang

文献摘要

相似文献

本文研究了一类随机梯度下降(SGD)动量格式的收敛性和推广性。基于动量的SGD加速是许多深度学习模型的默认优化器。然而,有许多现有的动量变量结合随机梯度缺乏一般的收敛保证。目前还不清楚动量方法如何影响泛化误差。在本文中,我们给出了几个流行的优化,例如,波利亚克的重球动量和内斯特罗夫的加速梯度。我们的贡献是三方面的。首先,我们给出了一个统一的收敛保证的一大类动量变量的随机设置。值得注意的是,我们的结果涵盖凸和非凸目标。其次,我们证明了动量变量训练的神经网络的泛化界。我们分析了超参数如何影响推广界,并因此提出了如何在各种动量方案中调整这些超参数以良好推广的指导方针。我们提供了广泛的经验证据,我们提出的指导方针。第三,本研究填补了文献中对微调进行形式化分析的空白。据我们所知,我们的工作是第一个系统的概括性分析的势头方法,包括从头开始学习和微调。我们的代码可以在https://github.com/jsycsjh/Demystify-Hyperparameters-for-Stochastic-Optimization-with-Transferable-Representations上找到。
This paper studies the convergence and generalization of a large class of Stochastic Gradient Descent (SGD) momentum schemes, in both learning from scratch and transferring representations with fine-tuning. Momentum-based acceleration of SGD is the default optimizer for many deep learning models. However, there is a lack of general convergence guarantees for many existing momentum variants in conjunction withstochastic gradient. It is also unclear how the momentum methods may affect thegeneralization error. In this paper, we give a unified analysis of several popular optimizers, e.g., Polyak's heavy ball momentum and Nesterov's accelerated gradient. Our contribution is threefold. First, we give a unified convergence guarantee for a large class of momentum variants in thestochastic setting. Notably, our results cover both convex and nonconvex objectives. Second, we prove a generalization bound for neural networks trained by momentum variants. We analyze how hyperparameters affect the generalization bound and consequently propose guidelines on how to tune these hyperparameters in various momentum schemes to generalize well. We provide extensive empirical evidence to our proposed guidelines. Third, this study fills the vacancy of a formal analysis of fine-tuning in literature. To our best knowledge, our work is the first systematic generalizability analysis on momentum methods that cover both learning from scratch and fine-tuning. Our codes are available https://github.com/jsycsjh/Demystify-Hyperparameters-for-Stochastic-Optimization-with-Transferable-Representations .