Meta-Learning with Implicit Gradients

Meta-Learning with Implicit Gradients
复制标题

DOI:
--
复制
发表时间:
2019-09
期刊:
--
影响因子:
--
通讯作者:
A. Rajeswaran;Chelsea Finn;S. Kakade;S. Levine
A. Rajeswaran;Chelsea Finn;S. Kakade;S. Levine
中科院分区:
其他
文献类型:
--
作者:
A. Rajeswaran;Chelsea Finn;S. Kakade;S. Levine

文献摘要

被引文献

相似文献

智能系统的核心能力是能够通过借鉴先前的经验来快速学习新任务的能力。基于梯度(或优化)基于元学习的方法最近已成为几次学习的有效方法。在此公式中,在外循环中学习了元参数,而在内循环中仅使用当前任务中的少量数据来学习特定于任务的模型。扩展这些方法的关键挑战是需要通过内部循环学习过程进行区分,这可以施加相当大的计算和内存负担。通过借鉴隐式分化,我们开发了隐式MAML算法,该算法仅取决于内部级优化的解决方案,而不是内部循环优化器所采用的路径。这有效地将元梯度计算与内部循环优化器的选择分解。结果,我们的方法对内部循环优化器的选择不可知,并且可以优雅地处理许多梯度步骤,而不会消失梯度或记忆约束。从理论上讲,我们证明隐式MAML可以用内存足迹来计算准确的元梯度,而记忆足迹的恒定因素不超过计算单个内部循环梯度所需的,并且总体计算成本的总体上没有总体上升。 。在实验上,我们表明,隐式MAML的这些好处转化为几乎没有图像识别基准的经验收益。
A core capability of intelligent systems is the ability to quickly learn new tasks by drawing on prior experience. Gradient (or optimization) based meta-learning has recently emerged as an effective approach for few-shot learning. In this formulation, meta-parameters are learned in the outer loop, while task-specific models are learned in the inner-loop, by using only a small amount of data from the current task. A key challenge in scaling these approaches is the need to differentiate through the inner loop learning process, which can impose considerable computational and memory burdens. By drawing upon implicit differentiation, we develop the implicit MAML algorithm, which depends only on the solution to the inner level optimization and not the path taken by the inner loop optimizer. This effectively decouples the meta-gradient computation from the choice of inner loop optimizer. As a result, our approach is agnostic to the choice of inner loop optimizer and can gracefully handle many gradient steps without vanishing gradients or memory constraints. Theoretically, we prove that implicit MAML can compute accurate meta-gradients with a memory footprint that is, up to small constant factors, no more than that which is required to compute a single inner loop gradient and at no overall increase in the total computational cost. Experimentally, we show that these benefits of implicit MAML translate into empirical gains on few-shot image recognition benchmarks.