Neural Routing by Memory

Neural Routing by Memory
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Kaipeng Zhang;Zhenqiang Li;Zhifeng Li;Wei Liu;Yoichi Sato
Kaipeng Zhang;Zhenqiang Li;Zhifeng Li;Wei Liu;Yoichi Sato
中科院分区:
其他
文献类型:
--
作者:
Kaipeng Zhang;Zhenqiang Li;Zhifeng Li;Wei Liu;Yoichi Sato

文献摘要

相似文献

最近的卷积神经网络(CNN)通过堆叠多个卷积块(本文中称为过程)来提取语义特征,取得了重大成功。然而,它们对所有输入使用相同的过程序列,而不管中间特征。本文提出了一个简单而有效的思想,构造并行过程,并分配类似的中间功能,以分而治之的方式相同的专业程序。它减轻了每个程序的学习难度,从而导致上级性能。具体来说,我们为现有的CNN架构提出了一种按内存路由的机制。在网络的每个阶段,我们引入并行的程序单元(PU)。PU由一个内存头和一个过程组成。存储器头维护一种类型的特征的概要。对于中间特征,我们搜索其最近的记忆并将其转发到训练和测试中的相应过程。通过这种方式,不同的程序可以针对不同的特点进行调整,从而更好地解决这些问题。具有所提出的机制的网络可以使用四步训练策略来有效地训练。实验结果表明,我们的方法提高了VGGNet,ResNet和EfficientNet在Tiny ImageNet,ImageNet和CIFAR-100基准测试中的准确性,而额外的计算成本可以忽略不计。
Recent Convolutional Neural Networks (CNNs) have achieved significant success by stacking multiple convolutional blocks, named procedures in this paper, to extract semantic features. However, they use the same procedure sequence for all inputs, regardless of the intermediate features. This paper proffers a simple yet effective idea of constructing parallel procedures and assigning similar intermediate features to the same specialized procedures in a divide-and-conquer fashion. It relieves each procedure’s learning difficulty and thus leads to superior performance. Specifically, we propose a routing-by-memory mechanism for existing CNN architectures. In each stage of the network, we introduce parallel Procedural Units (PUs). A PU consists of a memory head and a procedure. The memory head maintains a summary of a type of features. For an intermediate feature, we search its closest memory and forward it to the corresponding procedure in both training and testing. In this way, different procedures are tailored to different features and therefore tackle them better. Networks with the proposed mechanism can be trained efficiently using a four-step training strategy. Experimental results show that our method improves VGGNet, ResNet, and EfficientNet’s accuracies on Tiny ImageNet, ImageNet, and CIFAR-100 benchmarks with a negligible extra computational cost.