ModuleNet: Knowledge-Inherited Neural Architecture Search

ModuleNet: Knowledge-Inherited Neural Architecture Search
复制标题

DOI:
10.1109/tcyb.2021.3078573
复制
发表时间:
2020-04
影响因子:
11.8
通讯作者:
Yaran Chen;Ruiyuan Gao;Fenggang Liu;Dongbin Zhao
Yaran Chen;Ruiyuan Gao;Fenggang Liu;Dongbin Zhao
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yaran Chen;Ruiyuan Gao;Fenggang Liu;Dongbin Zhao

文献摘要

被引文献

相似文献

神经网络结构搜索(NAS)虽然能对深层模型带来改进,但往往忽略了已有模型的宝贵知识。NAS中的计算和时间消耗属性也意味着我们不应该从头开始搜索,而是尽一切努力重用现有的知识。在本文中,我们讨论了模型中的哪些知识可以并且应该用于新的架构设计。然后,我们提出了一种新的NAS算法,即ModuleNet,它可以完全继承现有卷积神经网络的知识。为了充分利用现有的模型,我们将现有的模型分解成不同的模块,这些模块也保持它们的权重,组成一个知识库。然后,我们采样和搜索一个新的体系结构,根据知识库。与以往的搜索算法,并受益于继承的知识,我们的方法是能够直接搜索架构在宏空间的NSGA-II算法,而无需调整参数,在这些模块。实验表明,我们的策略可以有效地评估一个新的架构的性能,即使没有调整卷积层的权重。在我们继承的知识的帮助下,我们的搜索结果总是可以在各种数据集(CIFAR 10,CIFAR 100和ImageNet)上实现比原始架构更好的性能。
Although neural the architecture search (NAS) can bring improvement to deep models, it always neglects precious knowledge of existing models. The computation and time costing property in NAS also means that we should not start from scratch to search, but make every attempt to reuse the existing knowledge. In this article, we discuss what kind of knowledge in a model can and should be used for a new architecture design. Then, we propose a new NAS algorithm, namely, ModuleNet, which can fully inherit knowledge from the existing convolutional neural networks. To make full use of the existing models, we decompose existing models into different modules, which also keep their weights, consisting of a knowledge base. Then, we sample and search for a new architecture according to the knowledge base. Unlike previous search algorithms, and benefiting from inherited knowledge, our method is able to directly search for architectures in the macrospace by the NSGA-II algorithm without tuning parameters in these modules. Experiments show that our strategy can efficiently evaluate the performance of a new architecture even without tuning weights in convolutional layers. With the help of knowledge we inherited, our search results can always achieve better performance on various datasets (CIFAR10, CIFAR100, and ImageNet) over original architectures.