One-Shot Neural Architecture Search: Maximising Diversity to Overcome Catastrophic Forgetting

One-Shot Neural Architecture Search: Maximising Diversity to Overcome Catastrophic Forgetting
复制标题

DOI:
10.1109/tpami.2020.3035351
复制
发表时间:
2020-11
影响因子:
23.6
通讯作者:
Miao Zhang;Huiqi Li;Shirui Pan;Xiaojun Chang;Chuan Zhou;ZongYuan Ge;Steven W. Su
Miao Zhang;Huiqi Li;Shirui Pan;Xiaojun Chang;Chuan Zhou;ZongYuan Ge;Steven W. Su
中科院分区:
计算机科学1区
文献类型:
--
作者:
Miao Zhang;Huiqi Li;Shirui Pan;Xiaojun Chang;Chuan Zhou;ZongYuan Ge;Steven W. Su

文献摘要

被引文献

相似文献

一次性神经架构搜索(NAS)最近成为 NAS 社区的主流,因为它通过权重共享显着提高了计算效率。然而,一次性 NAS 中的超网训练范例引入了灾难性遗忘,其中训练的每一步都会降低包含与当前架构部分共享权重的其他架构的性能。为了克服这一灾难性遗忘问题,我们将一次性 NAS 的超网训练制定为一个受约束的持续学习优化问题,这样学习当前架构就不会降低先前架构的验证准确性。解决这个约束优化问题的关键是基于新颖性搜索的架构选择(NSAS)损失函数,该函数通过使用贪婪新颖性搜索方法来寻找最具代表性的子集来规范超网训练。我们将 NSAS 损失函数应用于两个一次性 NAS 基线,并在公共搜索空间和 NAS 基准数据集上对它们进行了广泛的测试。我们进一步基于 NSAS 损失函数推导了三种变体,具有深度约束的 NSAS (NSAS-C) 来提高可转移性,NSAS-G 和 NSAS-LG 来处理具有有限数量约束的情况。在公共 NAS 搜索空间上的实验表明,NSAS 及其变体提高了一次性 NAS 中超网训练的预测能力,在 CIFAR-10、CIFAR-100 和 ImageNet 数据集上具有显着且高效的性能。 NAS 基准数据集的结果也证实了这些一次性 NAS 基准可以做出的显着改进。
One-shot neural architecture search (NAS) has recently become mainstream in the NAS community because it significantly improves computational efficiency through weight sharing. However, the supernet training paradigm in one-shot NAS introduces catastrophic forgetting, where each step of the training can deteriorate the performance of other architectures that contain partially-shared weights with current architecture. To overcome this problem of catastrophic forgetting, we formulate supernet training for one-shot NAS as a constrained continual learning optimization problem such that learning the current architecture does not degrade the validation accuracy of previous architectures. The key to solving this constrained optimization problem is a novelty search based architecture selection (NSAS) loss function that regularizes the supernet training by using a greedy novelty search method to find the most representative subset. We applied the NSAS loss function to two one-shot NAS baselines and extensively tested them on both a common search space and a NAS benchmark dataset. We further derive three variants based on the NSAS loss function, the NSAS with depth constrain (NSAS-C) to improve the transferability, and NSAS-G and NSAS-LG to handle the situation with a limited number of constraints. The experiments on the common NAS search space demonstrate that NSAS and it variants improve the predictive ability of supernet training in one-shot NAS with remarkable and efficient performance on the CIFAR-10, CIFAR-100, and ImageNet datasets. The results with the NAS benchmark dataset also confirm the significant improvements these one-shot NAS baselines can make.