Convergence of Meta-Learning with Task-Specific Adaptation over Partial Parameters

Convergence of Meta-Learning with Task-Specific Adaptation over Partial Parameters
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Kaiyi Ji;J. Lee;Yingbin Liang;H. Poor
Kaiyi Ji;J. Lee;Yingbin Liang;H. Poor
中科院分区:
其他
文献类型:
--
作者:
Kaiyi Ji;J. Lee;Yingbin Liang;H. Poor

文献摘要

相似文献

虽然模型不可知元学习(MAML)是元学习实践中非常成功的算法,但它可能具有很高的计算成本,因为它在特定任务自适应的内环和Meta初始化训练的外环上更新所有模型参数。Raghu et al. 2019最近提出了一种更有效的算法ANIL(几乎没有内部循环),它只适应内部循环中的一小部分参数,因此计算成本比MAML低得多。然而,ANIL的理论收敛性尚未得到研究。在本文中,我们描述了收敛速度和计算复杂性的ANIL下两个代表性的内环损失的几何形状,即,强凸性和非凸性我们的研究结果表明,这样的几何属性可以显着影响ANIL的整体收敛性能。例如,ANIL实现了一个更快的收敛速度为一个强凸的内环损失的内环梯度下降步骤的数量$N$增加,但一个较慢的收敛速度为一个非凸的内环损失的$N$增加。此外,我们的复杂性分析提供了一个理论量化的ANIL的效率提高MAML。在标准的少量元学习基准上的实验验证了我们的理论发现。
Although model-agnostic meta-learning (MAML) is a very successful algorithm in meta-learning practice, it can have high computational cost because it updates all model parameters over both the inner loop of task-specific adaptation and the outer-loop of meta initialization training. A more efficient algorithm ANIL (which refers to almost no inner loop) was proposed recently by Raghu et al. 2019, which adapts only a small subset of parameters in the inner loop and thus has substantially less computational cost than MAML as demonstrated by extensive experiments. However, the theoretical convergence of ANIL has not been studied yet. In this paper, we characterize the convergence rate and the computational complexity for ANIL under two representative inner-loop loss geometries, i.e., strongly-convexity and nonconvexity. Our results show that such a geometric property can significantly affect the overall convergence performance of ANIL. For example, ANIL achieves a faster convergence rate for a strongly-convex inner-loop loss as the number $N$ of inner-loop gradient descent steps increases, but a slower convergence rate for a nonconvex inner-loop loss as $N$ increases. Moreover, our complexity analysis provides a theoretical quantification on the improved efficiency of ANIL over MAML. The experiments on standard few-shot meta-learning benchmarks validate our theoretical findings.