ZiCo: Zero-shot NAS via Inverse Coefficient of Variation on Gradients

ZiCo: Zero-shot NAS via Inverse Coefficient of Variation on Gradients
复制标题

DOI:
10.48550/arxiv.2301.11300
复制
发表时间:
2023-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Guihong Li;Yuedong Yang-;Kartikeya Bhardwaj;R. Marculescu
Guihong Li;Yuedong Yang-;Kartikeya Bhardwaj;R. Marculescu
中科院分区:
其他
文献类型:
--
作者:
Guihong Li;Yuedong Yang-;Kartikeya Bhardwaj;R. Marculescu

文献摘要

相似文献

神经网络结构搜索(NAS)被广泛用于在大量候选结构中自动获得性能最佳的神经网络。为了减少搜索时间,zero-shot NAS旨在设计无需训练的代理,可以预测给定架构的测试性能。然而,如最近所示,迄今为止提出的零触发代理中没有一个实际上可以始终比朴素代理(即网络参数的数量(#Params))更好地工作。为了改善这种状况,作为主要的理论贡献,我们首先揭示了不同样本之间的某些特定梯度特性如何影响神经网络的收敛速度和泛化能力。基于这一理论分析,我们提出了一个新的zero-shot代理,ZiCo,第一个始终比#Params更好的代理。我们证明了ZiCo在几个流行的NAS基准测试(NASBench 101,NATSBench-SSS/TSS,TransNASBench-101)上比最先进的(SOTA)代理更好地用于多个应用程序(例如,图像分类/重建和像素级预测)。最后,我们证明了通过ZiCo找到的最佳架构与通过一次和多次NAS方法找到的架构一样具有竞争力,但搜索时间要少得多。例如,基于ZiCo的NAS可以在ImageNet上在0.4 GPU天内分别在450 M,600 M和1000 M FLOP的推理预算下找到78.1%,79.4%和80.4%的测试准确度的最佳架构。我们的代码可在https://github.com/SLDGroup/ZiCo上获得。
Neural Architecture Search (NAS) is widely used to automatically obtain the neural network with the best performance among a large number of candidate architectures. To reduce the search time, zero-shot NAS aims at designing training-free proxies that can predict the test performance of a given architecture. However, as shown recently, none of the zero-shot proxies proposed to date can actually work consistently better than a naive proxy, namely, the number of network parameters (#Params). To improve this state of affairs, as the main theoretical contribution, we first reveal how some specific gradient properties across different samples impact the convergence rate and generalization capacity of neural networks. Based on this theoretical analysis, we propose a new zero-shot proxy, ZiCo, the first proxy that works consistently better than #Params. We demonstrate that ZiCo works better than State-Of-The-Art (SOTA) proxies on several popular NAS-Benchmarks (NASBench101, NATSBench-SSS/TSS, TransNASBench-101) for multiple applications (e.g., image classification/reconstruction and pixel-level prediction). Finally, we demonstrate that the optimal architectures found via ZiCo are as competitive as the ones found by one-shot and multi-shot NAS methods, but with much less search time. For example, ZiCo-based NAS can find optimal architectures with 78.1%, 79.4%, and 80.4% test accuracy under inference budgets of 450M, 600M, and 1000M FLOPs, respectively, on ImageNet within 0.4 GPU days. Our code is available at https://github.com/SLDGroup/ZiCo.