The Step Decay Schedule: A Near Optimal, Geometrically Decaying Learning Rate Procedure

The Step Decay Schedule: A Near Optimal, Geometrically Decaying Learning Rate Procedure
复制标题

步长衰减计划:近乎最优的几何衰减学习率过程

DOI:
--
复制
发表时间:
2019
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
Praneeth Netrapalli
Praneeth Netrapalli
中科院分区:
--
文献类型:
--
作者:
Rong Ge;S. Kakade;Rahul Kidambi;Praneeth Netrapalli

文献摘要

参考文献

被引文献

相似文献

实用大规模机器学习中使用的步骤时间表与被随机近似理论认为最佳的步长计划之间存在明显的差异。从理论上讲,大多数结果都利用多项式衰减的学习率时间表,而在实践中,“阶跃衰减”时间表是最受欢迎的时间表之一,其中学习率是削减每个恒定数量的时代(即这是一个几何衰减的时间表)) 。 这项工作研究了流式传输最小二乘回归的随机优化问题的台阶时间表(无论是在非严格的凸面和强烈凸出的情况下),我们表明,最佳学习率时间表的理论表征急剧更加细微差别。比以前的工作所建议。我们专门关注使用随机梯度下降的最终迭代时可以达到的速率,就像在实践中通常所做的那样。事实证明,我们的主要结果表明,正确调整的几何衰减学习率计划为任何多项式衰减的学习率计划提供了指数改进(就条件数而言)。我们还为这些结果的更广泛适用性提供了实验支持,包括用于培训现代深度神经网络。
There is a stark disparity between the step size schedules used in practical large scale machine learning and those that are considered optimal by the theory of stochastic approximation. In theory, most results utilize polynomially decaying learning rate schedules, while, in practice, the "Step Decay" schedule is among the most popular schedules, where the learning rate is cut every constant number of epochs (i.e. this is a geometrically decaying schedule). This work examines the step-decay schedule for the stochastic optimization problem of streaming least squares regression (both in the non-strongly convex and strongly convex case), where we show that a sharp theoretical characterization of an optimal learning rate schedule is far more nuanced than suggested by previous work. We focus specifically on the rate that is achievable when using the final iterate of stochastic gradient descent, as is commonly done in practice. Our main result provably shows that a properly tuned geometrically decaying learning rate schedule provides an exponential improvement (in terms of the condition number) over any polynomially decaying learning rate schedule. We also provide experimental support for wider applicability of these results, including for training modern deep neural networks.
DOI: --
发表时间: 2019
期刊: Advances in neural information processing systems
影响因子: --
作者:
Aybat, Necdet Serhat;Fallah, Alireza;Gurbuzbalaban, Mert;Ozdaglar, Asuman
通讯作者: Ozdaglar, Asuman