Traditional and Accelerated Gradient Descent for Neural Architecture Search

Traditional and Accelerated Gradient Descent for Neural Architecture Search
复制标题

用于神经架构搜索的传统和加速梯度下降

DOI:
10.1007/978-3-030-80209-7_55
复制
发表时间:
2021
期刊:
978-3-030-80208-0
影响因子:
--
通讯作者:
Morales, Javier
Morales, Javier
中科院分区:
--
文献类型:
--
作者:
Garcia Trillos, Nicolas;Morales, Felix;Morales, Javier

文献摘要

相似文献

在本文中,我们介绍了两种算法的神经结构搜索(NASGD和NASAGD)以下的理论工作,由两位作者[4],其中使用的几何结构的最佳运输介绍的概念基础的新概念的传统和加速梯度下降算法的优化功能的半离散空间。我们的算法使用[1]中引入的网络态射框架作为基线,可以分析40倍于爬山方法的架构[1,10],同时使用相同的计算资源和时间,并达到可比的准确度水平。例如,在CIFAR-10上使用NASGD,我们的方法在单个GPU上仅用12小时就设计和训练了错误率为4.06的网络。
In this paper we introduce two algorithms for neural architecture search (NASGD and NASAGD) following the theoretical work by two of the authors [4] which used the geometric structure of optimal transport to introduce the conceptual basis for new notions of traditional and accelerated gradient descent algorithms for the optimization of a function on a semi-discrete space. Our algorithms, which use the network morphism framework introduced in [1] as a baseline, can analyze forty times as many architectures as the hill climbing methods [1, 10] while using the same computational resources and time and achieving comparable levels of accuracy. For example, using NASGD on CIFAR-10, our method designs and trains networks with an error rate of 4.06 in only 12 h on a single GPU.