Traditional and Accelerated Gradient Descent for Neural Architecture Search
Traditional and Accelerated Gradient Descent for Neural Architecture Search
复制标题
用于神经架构搜索的传统和加速梯度下降
DOI:
10.1007/978-3-030-80209-7_55
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Morales, Javier
中科院分区:
文献类型:
--
作者:
Garcia Trillos, Nicolas;Morales, Felix;Morales, Javier
In this paper we introduce two algorithms for neural architecture search (NASGD and NASAGD) following the theoretical work by two of the authors [4] which used the geometric structure of optimal transport to introduce the conceptual basis for new notions of traditional and accelerated gradient descent algorithms for the optimization of a function on a semi-discrete space. Our algorithms, which use the network morphism framework introduced in [1] as a baseline, can analyze forty times as many architectures as the hill climbing methods [1, 10] while using the same computational resources and time and achieving comparable levels of accuracy. For example, using NASGD on CIFAR-10, our method designs and trains networks with an error rate of 4.06 in only 12 h on a single GPU.