On the conditions of MAML convergence

On the conditions of MAML convergence
复制标题

DOI:
--
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
S. Takagi;Yoshihiro Nagano;Yuki Yoshida;M. Okada
S. Takagi;Yoshihiro Nagano;Yuki Yoshida;M. Okada
中科院分区:
其他
文献类型:
--
作者:
S. Takagi;Yoshihiro Nagano;Yuki Yoshida;M. Okada

文献摘要

相似文献

模型不可知元学习(MAML)是一种功能强大的元学习方法。在本文中,我们得到了内部学习率(cid:11)和元学习率(cid:12)必须满足的条件,使简化的艾德MAML从任何点局部收敛到局部极小值。我们发现(cid:12)的上界依赖于(cid:11),与使用正常梯度下降法的情况相反。此外,我们还证明了(cid:12)的阈值随着(cid:11)接近其自身的上界而增加。这一结果通过各种少数任务和架构的实验得到了艾德;具体来说,我们使用多层感知器和CNN对Omniglot和MiniImagenet数据集进行正弦回归和分类。基于这个结果,我们提出了一个确定学习率的准则:首先,搜索最大可能的(cid:11);其次,根据(cid:11)的选择值调整(cid:12)。
Model-agnostic meta-learning (MAML) is known as a powerful meta-learning method. In this paper, we derive the conditions that inner learning rate (cid:11) and meta-learning rate (cid:12) must satisfy for a simplified MAML to locally converge to local minima from any point. We find that the upper bound of (cid:12) depends on (cid:11) , in contrast to the case of using the normal gradient descent method. Moreover, we show that the threshold of (cid:12) increases as (cid:11) approaches its own upper bound. This result is verified by experiments on various few-shot tasks and architectures; specifically, we perform sinusoid regression and classification of Omniglot and MiniImagenet datasets with a multilayer perceptron and a CNN. Based on this outcome, we present a guideline for determining the learning rates: first, search for the largest possible (cid:11) ; next, tune (cid:12) based on the chosen value of (cid:11) .