Achieving Forgetting Prevention and Knowledge Transfer in Continual Learning

Achieving Forgetting Prevention and Knowledge Transfer in Continual Learning
复制标题

DOI:
--
复制
发表时间:
2021-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Zixuan Ke;Bing Liu;Nianzu Ma;Hu Xu;Lei Shu
Zixuan Ke;Bing Liu;Nianzu Ma;Hu Xu;Lei Shu
中科院分区:
其他
文献类型:
--
作者:
Zixuan Ke;Bing Liu;Nianzu Ma;Hu Xu;Lei Shu

文献摘要

相似文献

持续学习(CL)以递增的方式学习一系列任务,目标是实现两个主要目标:克服灾难性遗忘(CF)和鼓励跨任务的知识转移(KT)。然而,现有的大多数技术只关注于克服知识转移,没有激励知识转移的机制,因此在知识转移中做得不好。虽然已经有几篇论文试图同时处理任务和知识转换,但我们的实验表明,当任务没有太多共享知识时,它们会受到严重的任务冲突的影响。另一个观察结果是,大多数当前的协作学习方法不使用预先训练的模型,但已有研究表明,这种模型可以显著提高结束任务的性能。例如,在自然语言处理中,微调类似Bert的预训练语言模型是最有效的方法之一。然而,对于CL来说,这种方法受到了严重的CF的影响。一个有趣的问题是,如何最大限度地利用预先训练好的模型。本文提出了一种称为CTR的新模型来解决这些问题。实验结果证明了CTR的有效性
Continual learning (CL) learns a sequence of tasks incrementally with the goal of achieving two main objectives: overcoming catastrophic forgetting (CF) and encouraging knowledge transfer (KT) across tasks. However, most existing techniques focus only on overcoming CF and have no mechanism to encourage KT, and thus do not do well in KT. Although several papers have tried to deal with both CF and KT, our experiments show that they suffer from serious CF when the tasks do not have much shared knowledge. Another observation is that most current CL methods do not use pre-trained models, but it has been shown that such models can significantly improve the end task performance. For example, in natural language processing, fine-tuning a BERT-like pre-trained language model is one of the most effective approaches. However, for CL, this approach suffers from serious CF. An interesting question is how to make the best use of pre-trained models for CL. This paper proposes a novel model called CTR to solve these problems. Our experimental results demonstrate the effectiveness of CTR