Guarantees for Tuning the Step Size using a Learning-to-Learn Approach

Guarantees for Tuning the Step Size using a Learning-to-Learn Approach
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
--
影响因子:
--
通讯作者:
Xiang Wang;Shuai Yuan;Chenwei Wu;Rong Ge
Xiang Wang;Shuai Yuan;Chenwei Wu;Rong Ge
中科院分区:
其他
文献类型:
--
作者:
Xiang Wang;Shuai Yuan;Chenwei Wu;Rong Ge

文献摘要

相似文献

学习-学习(使用优化算法学习新的优化器)在实践中成功地培养了高效的优化器。这种方法依赖于基于优化器生成的轨迹在元目标上的亚梯度下降。然而,对于如何避免亚梯度爆炸/消失问题,以及如何训练具有良好泛化性能的优化器,几乎没有理论上的保证。在这篇文章中,我们研究了一个简单的问题,即调节二次损失的步长,学习到学习的方法。我们的结果表明,尽管有一种方法可以设计元目标,使元梯度保持多项式有界,但直接使用反向传播计算元梯度会导致类似于梯度爆炸/消失问题的数值问题。我们还描述了何时需要在单独的验证集而不是原始训练集上计算元目标。最后,我们对我们的结果进行了实证验证,结果表明,即使对于由神经网络参数化的更复杂的学习优化器,也会出现类似的现象。
Learning-to-learn (using optimization algorithms to learn a new optimizer) has successfully trained efficient optimizers in practice. This approach relies on meta-gradient descent on a meta-objective based on the trajectory that the optimizer generates. However, there were few theoretical guarantees on how to avoid meta-gradient explosion/vanishing problems, or how to train an optimizer with good generalization performance. In this paper, we study the learning-to-learn approach on a simple problem of tuning the step size for quadratic loss. Our results show that although there is a way to design the meta-objective so that the meta-gradient remain polynomially bounded, computing the meta-gradient directly using backpropagation leads to numerical issues that look similar to gradient explosion/vanishing problems. We also characterize when it is necessary to compute the meta-objective on a separate validation set instead of the original training set. Finally, we verify our results empirically and show that a similar phenomenon appears even for more complicated learned optimizers parametrized by neural networks.