Linear Mode Connectivity in Multitask and Continual Learning

Linear Mode Connectivity in Multitask and Continual Learning
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Seyed Iman Mirzadeh;Mehrdad Farajtabar;Dilan Gorur;Razvan Pascanu;H. Ghasemzadeh
Seyed Iman Mirzadeh;Mehrdad Farajtabar;Dilan Gorur;Razvan Pascanu;H. Ghasemzadeh
中科院分区:
其他
文献类型:
--
作者:
Seyed Iman Mirzadeh;Mehrdad Farajtabar;Dilan Gorur;Razvan Pascanu;H. Ghasemzadeh

文献摘要

被引文献

相似文献

连续(顺序)训练和多任务(同时)训练通常试图解决相同的总体目标:找到一个在所有考虑的任务上都表现良好的解决方案。主要区别在于训练机制,持续学习一次只能访问一个任务,这对于神经网络来说通常会导致灾难性的遗忘。也就是说,为后续任务找到的解决方案在先前的任务上不再表现良好。然而,这两种训练制度达到的不同最小值之间的关系还没有得到很好的理解。是什么让他们与众不同?是否有一个当地的结构可以解释这两个不同的计划所取得的业绩差异?最近的工作表明,不同的最小值相同的任务通常连接非常简单的曲线,低错误的动机,我们调查是否多任务和连续的解决方案是类似的连接。我们根据经验发现,这种连接确实可以可靠地实现,更有趣的是,它可以通过线性路径来实现,条件是两者具有相同的初始化。我们深入分析了这一观察结果,并讨论了它对持续学习过程的意义。此外,我们利用这一发现,提出了一个有效的算法,约束顺序学习的最小行为的多任务解决方案。我们表明,我们的方法优于几个国家的最先进的持续学习算法在各种视觉基准。
Continual (sequential) training and multitask (simultaneous) training are often attempting to solve the same overall objective: to find a solution that performs well on all considered tasks. The main difference is in the training regimes, where continual learning can only have access to one task at a time, which for neural networks typically leads to catastrophic forgetting. That is, the solution found for a subsequent task does not perform well on the previous ones anymore. However, the relationship between the different minima that the two training regimes arrive at is not well understood. What sets them apart? Is there a local structure that could explain the difference in performance achieved by the two different schemes? Motivated by recent work showing that different minima of the same task are typically connected by very simple curves of low error, we investigate whether multitask and continual solutions are similarly connected. We empirically find that indeed such connectivity can be reliably achieved and, more interestingly, it can be done by a linear path, conditioned on having the same initialization for both. We thoroughly analyze this observation and discuss its significance for the continual learning process. Furthermore, we exploit this finding to propose an effective algorithm that constrains the sequentially learned minima to behave as the multitask solution. We show that our method outperforms several state of the art continual learning algorithms on various vision benchmarks.