Enhancing network modularity to mitigate catastrophic forgetting

Enhancing network modularity to mitigate catastrophic forgetting
复制标题

DOI:
10.1007/s41109-020-00332-9
复制
发表时间:
2020-11
影响因子:
2.2
通讯作者:
Lu Chen;M. Murata
Lu Chen;M. Murata
中科院分区:
--
文献类型:
--
作者:
Lu Chen;M. Murata

文献摘要

被引文献

相似文献

当学习算法改变用于编码先前获得的技能的连接以学习新技能时,就会发生灾难性遗忘。最近,随着学习问题的规模和复杂性的增长,神经网络的模块化方法被认为是必要的,因为它直观地应该通过将功能分离到物理上不同的网络模块中来减少学习干扰。然而,算法方法在实践中是困难的,因为它涉及专家设计和试错。Kashtan等人发现,在以模块化方式变化的环境下的进化导致模块化网络结构的自发进化。在本文中,我们的目标是解决模块化变化的目标(MVG)的逆问题,以获得一个高度模块化的结构,可以减轻灾难性遗忘,使它也可以适用于现实的数据。首先,我们通过对现实数据集应用MVG来确认具有高度模块化结构的配置存在,并确认该神经网络可以减轻灾难性遗忘。接下来,我们解决了反向问题,也就是说,我们提出了一种方法,可以获得一个高度模块化的结构,能够减轻灾难性遗忘。由于MVG获得的神经网络可以相对地保持模块内元素,同时使模块间元素相对可变,因此我们提出了一种方法来限制模块间权重元素,使它们可以相对于模块内元素相对可变。从结果来看,所获得的神经网络具有高度模块化的结构,并且可以比没有这种方法更快地学习未学习的目标。
Catastrophic forgetting occurs when learning algorithms change connections used to encode previously acquired skills to learn a new skill. Recently, a modular approach for neural networks was deemed necessary as learning problems grow in scale and complexity since it intuitively should reduce learning interference by separating functionality into physically distinct network modules. However, an algorithmic approach is difficult in practice since it involves expert design and trial and error. Kashtan et al. finds that evolution under an environment that changes in a modular fashion leads to the spontaneous evolution of a modular network structure. In this paper, we aim to solve the reverse problem of modularly varying goal (MVG) to obtain a highly modular structure that can mitigate catastrophic forgetting so that it can also apply to realistic data. First, we confirm that a configuration with a highly modular structure exists by applying an MVG against a realistic dataset and confirm that this neural network can mitigate catastrophic forgetting. Next, we solve the reverse problem, that is, we propose a method that can obtain a highly modular structure able to mitigate catastrophic forgetting. Since the MVG-obtained neural network can relatively maintain the intra-module elements while leaving the inter-module elements relatively variable, we propose a method to restrict the inter-module weight elements so that they can be relatively variable against the intra-module ones. From the results, the obtained neural network has a highly modular structure and can learn an unlearned goal faster than without this method.