Selecting and Composing Learning Rate Policies for Deep Neural Networks

Selecting and Composing Learning Rate Policies for Deep Neural Networks
复制标题

选择和制定深度神经网络的学习率策略

DOI:
10.1145/3570508
复制
发表时间:
2023
影响因子:
5
通讯作者:
Liu, Ling
Liu, Ling
中科院分区:
计算机科学3区
文献类型:
--
作者:
Wu, Yanzhao;Liu, Ling

文献摘要

参考文献

被引文献

相似文献

学习率(LR)函数和策略的选择已经从简单的固定LR发展到衰减LR和循环LR,旨在提高深度神经网络(DNN)的准确性并减少训练时间。本文提出了一种系统的方法来选择和组成LR策略,以进行有效的DNN训练,以满足所需的目标精度并减少预定义训练迭代内的训练时间。它有三个原创性的贡献。首先,我们开发了一种LR调整机制,用于在预定义的训练时间约束下自动验证给定的LR策略是否达到所需的准确性目标。其次,我们开发了一个LR策略推荐系统(LRBench),通过动态调整从相同和/或不同的LR函数中选择和组合好的LR策略,并避免错误的选择,对于给定的学习任务,DNN模型和数据集。第三,我们通过支持不同的DNN优化器来扩展LRBench,并显示不同LR策略和不同优化器的显著相互影响。使用流行的基准数据集和不同的DNN模型(LeNet、CNN 3、ResNet)进行评估,我们表明我们的方法可以有效地提供高DNN测试准确性,优于现有推荐的默认LR策略,并将DNN训练时间减少1.6-6.7倍以满足目标模型准确度。
The choice of learning rate (LR) functions and policies has evolved from a simple fixed LR to the decaying LR and the cyclic LR, aiming to improve the accuracy and reduce the training time of Deep Neural Networks (DNNs). This article presents a systematic approach to selecting and composing an LR policy for effective DNN training to meet desired target accuracy and reduce training time within the pre-defined training iterations. It makes three original contributions. First, we develop an LR tuning mechanism for auto-verification of a given LR policy with respect to the desired accuracy goal under the pre-defined training time constraint. Second, we develop an LR policy recommendation system (LRBench) to select and compose good LR policies from the same and/or different LR functions through dynamic tuning, and avoid bad choices, for a given learning task, DNN model, and dataset. Third, we extend LRBench by supporting different DNN optimizers and show the significant mutual impact of different LR policies and different optimizers. Evaluated using popular benchmark datasets and different DNN models (LeNet, CNN3, ResNet), we show that our approach can effectively deliver high DNN test accuracy, outperform the existing recommended default LR policies, and reduce the DNN training time by 1.6-6.7× to meet a targeted model accuracy.
DOI: --
发表时间: 2016-11
期刊: ArXiv
影响因子: --
作者:
Barret Zoph;Quoc V. Le
通讯作者: Barret Zoph;Quoc V. Le
ADASECANT:随机梯度的鲁棒自适应割线法
DOI: --
发表时间: 2014
期刊: arXiv.org
影响因子: --
作者:
Çaglar Gülçehre;Yoshua Bengio
通讯作者: Yoshua Bengio
开发 Adaptive Boost 方法并使用该方法分析铁中的碳扩散 [MD 奖获得者]
DOI: --
发表时间: 2011
期刊:
影响因子: --
作者:
石井明男;尾方成信;君塚肇
通讯作者: 君塚肇