Training Stronger Baselines for Learning to Optimize

Training Stronger Baselines for Learning to Optimize
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Tianlong Chen;Weiyi Zhang;Jingyang Zhou;Shiyu Chang;Sijia Liu;Lisa Amini;Zhangyang Wang
Tianlong Chen;Weiyi Zhang;Jingyang Zhou;Shiyu Chang;Sijia Liu;Lisa Amini;Zhangyang Wang
中科院分区:
其他
文献类型:
--
作者:
Tianlong Chen;Weiyi Zhang;Jingyang Zhou;Shiyu Chang;Sijia Liu;Lisa Amini;Zhangyang Wang

文献摘要

被引文献

相似文献

学习优化(L2O)得到了越来越多的关注,因为经典的优化器需要费力的特定问题的设计和超参数调整。然而,现有的L2O模型的实际需求和可实现的性能之间存在差距。具体来说,这些学习优化器仅适用于有限的一类问题,并且通常表现出不稳定性。随着许多努力致力于设计更复杂的L2O模型,我们主张另一个正交的,未充分探索的主题:这些L2O模型的训练技术。我们证明,即使是最简单的L2O模型也可以训练得更好。我们首先提出了一个渐进的训练方案,以逐渐增加优化器展开长度,以减轻截断偏差(较短的展开)与梯度爆炸(较长的展开)的众所周知的L2O困境。我们进一步利用政策外模仿学习来指导L2O学习,通过参考分析优化器的行为。我们改进的训练技术被插入到各种最先进的L2O模型中,并立即提高其性能,而无需对其模型结构进行任何更改。特别是,通过我们提出的技术,可以训练最早和最简单的L2O模型,使其在许多任务上优于最新的复杂L2O模型。我们的研究结果表明,L2O的更大潜力尚未释放,并敦促重新思考最近的进展。我们的代码可在以下网址公开获取:this https URL。
Learning to optimize (L2O) has gained increasing attention since classical optimizers require laborious problem-specific design and hyperparameter tuning. However, there is a gap between the practical demand and the achievable performance of existing L2O models. Specifically, those learned optimizers are applicable to only a limited class of problems, and often exhibit instability. With many efforts devoted to designing more sophisticated L2O models, we argue for another orthogonal, under-explored theme: the training techniques for those L2O models. We show that even the simplest L2O model could have been trained much better. We first present a progressive training scheme to gradually increase the optimizer unroll length, to mitigate a well-known L2O dilemma of truncation bias (shorter unrolling) versus gradient explosion (longer unrolling). We further leverage off-policy imitation learning to guide the L2O learning, by taking reference to the behavior of analytical optimizers. Our improved training techniques are plugged into a variety of state-of-the-art L2O models, and immediately boost their performance, without making any change to their model structures. Especially, by our proposed techniques, an earliest and simplest L2O model can be trained to outperform the latest complicated L2O models on a number of tasks. Our results demonstrate a greater potential of L2O yet to be unleashed, and urge to rethink the recent progress. Our codes are publicly available at: this https URL.