Learning in deep neural networks and brains with similarity-weighted interleaved learning.
Learning in deep neural networks and brains with similarity-weighted interleaved learning.
复制标题
DOI:
10.1073/pnas.2115229119
复制
发表时间:
2022-07-05
影响因子:
11.1
通讯作者:
McNaughton, Bruce L.
中科院分区:
文献类型:
--
作者:
Saxena, Rajat;Shobe, Justin L.;McNaughton, Bruce L.
Unlike humans, artificial neural networks rapidly forget previously learned information when learning something new and must be retrained by interleaving the new and old items; however, interleaving all old items is time-consuming and might be unnecessary. It might be sufficient to interleave only old items having substantial similarity to new ones. We show that training with similarity-weighted interleaving of old items with new ones allows deep networks to learn new items rapidly without forgetting, while using substantially less data. We hypothesize how similarity-weighted interleaving might be implemented in the brain using persistent excitability traces on recently active neurons and attractor dynamics. These findings may advance both neuroscience and machine learning. Understanding how the brain learns throughout a lifetime remains a long-standing challenge. In artificial neural networks (ANNs), incorporating novel information too rapidly results in catastrophic interference, i.e., abrupt loss of previously acquired knowledge. Complementary Learning Systems Theory (CLST) suggests that new memories can be gradually integrated into the neocortex by interleaving new memories with existing knowledge. This approach, however, has been assumed to require interleaving all existing knowledge every time something new is learned, which is implausible because it is time-consuming and requires a large amount of data. We show that deep, nonlinear ANNs can learn new information by interleaving only a subset of old items that share substantial representational similarity with the new information. By using such similarity-weighted interleaved learning (SWIL), ANNs can learn new information rapidly with a similar accuracy level and minimal interference, while using a much smaller number of old items presented per epoch (fast and data-efficient). SWIL is shown to work with various standard classification datasets (Fashion-MNIST, CIFAR10, and CIFAR100), deep neural network architectures, and in sequential learning frameworks. We show that data efficiency and speedup in learning new items are increased roughly proportionally to the number of nonoverlapping classes stored in the network, which implies an enormous possible speedup in human brains, which encode a high number of separate categories. Finally, we propose a theoretical model of how SWIL might be implemented in the brain.
登录
查看更多内容
影响因子:
5.4
作者:
Gepperth, Alexander;Karaoguz, Cem
通讯作者:
Karaoguz, Cem
影响因子:
14.4
作者:
McNaughton, Bruce L.
通讯作者:
McNaughton, Bruce L.
DOI:
10.1073/pnas.1820226116
发表时间:
2019-06-04
影响因子:
11.1
作者:
Saxe, Andrew M.;McClelland, James L.;Ganguli, Surya
通讯作者:
Ganguli, Surya
影响因子:
56.9
作者:
Tse, Dorothy;Langston, Rosamund F.;Morris, Richard G. M.
通讯作者:
Morris, Richard G. M.
影响因子:
2
作者:
DEJONGE, MC;BLACK, J;DISTERHOFT, JF
通讯作者:
DISTERHOFT, JF