Theory on Forgetting and Generalization of Continual Learning

Theory on Forgetting and Generalization of Continual Learning
复制标题

DOI:
10.48550/arxiv.2302.05836
复制
发表时间:
2023-02
期刊:
--
影响因子:
--
通讯作者:
Sen Lin;Peizhong Ju;Yitao Liang;N. Shroff
Sen Lin;Peizhong Ju;Yitao Liang;N. Shroff
中科院分区:
其他
文献类型:
--
作者:
Sen Lin;Peizhong Ju;Yitao Liang;N. Shroff

文献摘要

相似文献

旨在学习一系列任务的持续学习(CL)引起了最近的关注。但是,大多数工作都集中在CL的实验性能上,而CL的理论研究仍然有限。特别是,缺乏了解哪些因素很重要,以及它们如何影响“灾难性遗忘”和泛化表现。为了填补这一空白,我们的理论分析(在过度参数化的线性模型下)提供了预期遗忘和概括误差的首先明确形式。对这样的关键结果的进一步分析产生了许多理论解释,该解释是关于过度参数化,任务相似性和任务排序如何影响CL的遗忘和概括错误。更有趣的是,通过使用深层神经网络(DNN)对实际数据集进行实验,我们表明其中一些见解甚至超出了线性模型,并且可以将其转移到实用的设置中。特别是,我们使用具体示例来表明我们的结果不仅解释了最近的研究中一些有趣的经验观察,而且还激发了CL的更好实用算法设计。
Continual learning (CL), which aims to learn a sequence of tasks, has attracted significant recent attention. However, most work has focused on the experimental performance of CL, and theoretical studies of CL are still limited. In particular, there is a lack of understanding on what factors are important and how they affect"catastrophic forgetting"and generalization performance. To fill this gap, our theoretical analysis, under overparameterized linear models, provides the first-known explicit form of the expected forgetting and generalization error. Further analysis of such a key result yields a number of theoretical explanations about how overparameterization, task similarity, and task ordering affect both forgetting and generalization error of CL. More interestingly, by conducting experiments on real datasets using deep neural networks (DNNs), we show that some of these insights even go beyond the linear models and can be carried over to practical setups. In particular, we use concrete examples to show that our results not only explain some interesting empirical observations in recent studies, but also motivate better practical algorithm designs of CL.