Provable Generalization of Overparameterized Meta-learning Trained with SGD

Provable Generalization of Overparameterized Meta-learning Trained with SGD
复制标题

DOI:
10.48550/arxiv.2206.09136
复制
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Yu Huang;Yingbin Liang;Longbo Huang-
Yu Huang;Yingbin Liang;Longbo Huang-
中科院分区:
其他
文献类型:
--
作者:
Yu Huang;Yingbin Liang;Longbo Huang-

文献摘要

相似文献

尽管深度元学习在经验上取得了上级成功,但对过度参数化元学习的理论理解仍然有限。本文研究了一种广泛使用的元学习方法,模型不可知元学习(MAML),其目的是找到一个良好的初始化快速适应新的任务。在混合线性回归模型下,我们分析了在过参数化机制下使用SGD训练的MAML的泛化性能。我们提供了MAML的超额风险的上限和下限,它捕捉了SGD动态如何影响这些泛化界限。有了这些尖锐的特征,我们进一步探讨了各种学习参数如何影响过参数化MAML的泛化能力,包括明确识别典型的数据和任务分布,可以实现减少泛化误差与过参数化,并表征适应学习率对过度风险和提前停止时间的影响。我们的理论研究结果得到了实验的进一步验证。
Despite the superior empirical success of deep meta-learning, theoretical understanding of overparameterized meta-learning is still limited. This paper studies the generalization of a widely used meta-learning approach, Model-Agnostic Meta-Learning (MAML), which aims to find a good initialization for fast adaptation to new tasks. Under a mixed linear regression model, we analyze the generalization properties of MAML trained with SGD in the overparameterized regime. We provide both upper and lower bounds for the excess risk of MAML, which captures how SGD dynamics affect these generalization bounds. With such sharp characterizations, we further explore how various learning parameters impact the generalization capability of overparameterized MAML, including explicitly identifying typical data and task distributions that can achieve diminishing generalization error with overparameterization, and characterizing the impact of adaptation learning rate on both excess risk and the early stopping time. Our theoretical findings are further validated by experiments.