Learning curves for continual learning in neural networks: Self-knowledge transfer and forgetting

Learning curves for continual learning in neural networks: Self-knowledge transfer and forgetting
复制标题

DOI:
--
复制
发表时间:
2021-12
期刊:
--
影响因子:
--
通讯作者:
Ryo Karakida;S. Akaho
Ryo Karakida;S. Akaho
中科院分区:
其他
文献类型:
--
作者:
Ryo Karakida;S. Akaho

文献摘要

相似文献

从任务到任务的顺序训练正在成为深度学习应用(如持续学习和迁移学习)的主要目标之一。然而,目前还不清楚在什么条件下训练模型的性能会提高或下降。为了加深我们对顺序训练的理解,本研究提供了一个理论分析的泛化性能在一个可解决的情况下的持续学习。我们认为神经网络的神经切线内核(NTK)制度,不断学习目标函数从任务到任务,并通过使用一个既定的统计力学分析内核无脊回归研究的泛化。我们首先展示了从正迁移到负迁移的特征性转变。在特定的临界值以上,更相似的目标可以为随后的任务实现积极的知识转移,而灾难性的遗忘即使非常相似的目标也会发生。接下来,我们研究了一种持续学习的变体,它假设在多个任务中具有相同的目标函数。即使对于相同的目标,训练后的模型也会根据每个任务的样本大小显示出一些转移和遗忘。我们可以保证,推广误差单调递减任务相同的样本量,而不平衡的样本量恶化的推广。我们将这些改善和恶化分别称为自我知识转移和遗忘,并在深度神经网络的实际训练中进行了实证验证。
Sequential training from task to task is becoming one of the major objects in deep learning applications such as continual learning and transfer learning. Nevertheless, it remains unclear under what conditions the trained model's performance improves or deteriorates. To deepen our understanding of sequential training, this study provides a theoretical analysis of generalization performance in a solvable case of continual learning. We consider neural networks in the neural tangent kernel (NTK) regime that continually learn target functions from task to task, and investigate the generalization by using an established statistical mechanical analysis of kernel ridge-less regression. We first show characteristic transitions from positive to negative transfer. More similar targets above a specific critical value can achieve positive knowledge transfer for the subsequent task while catastrophic forgetting occurs even with very similar targets. Next, we investigate a variant of continual learning which supposes the same target function in multiple tasks. Even for the same target, the trained model shows some transfer and forgetting depending on the sample size of each task. We can guarantee that the generalization error monotonically decreases from task to task for equal sample sizes while unbalanced sample sizes deteriorate the generalization. We respectively refer to these improvement and deterioration as self-knowledge transfer and forgetting, and empirically confirm them in realistic training of deep neural networks as well.