One Proxy Device Is Enough for Hardware-Aware Neural Architecture Search

One Proxy Device Is Enough for Hardware-Aware Neural Architecture Search
复制标题

一台代理设备足以进行硬件感知神经架构搜索

DOI:
10.1145/3489048.3522631
复制
发表时间:
2022
期刊:
ACM SIGMETRICS
影响因子:
--
通讯作者:
Ren, Shaolei
Ren, Shaolei
中科院分区:
--
文献类型:
--
作者:
Lu, Bingqian;Yang, Jianyi;Jiang, Weiwen;Shi, Yiyu;Ren, Shaolei

文献摘要

参考文献

被引文献

相似文献

卷积神经网络(cnn)被用于许多现实世界的应用,如基于视觉的自动驾驶和视频内容分析。为了在各种目标设备上运行CNN推理,硬件感知神经结构搜索(NAS)至关重要。高效的硬件感知NAS的一个关键要求是快速评估推理延迟,以便对不同的体系结构进行排序。虽然为每个目标设备构建一个延迟预测器是目前最常用的方法,但这是一个非常耗时的过程,而且在设备种类繁多的情况下缺乏可伸缩性。在这项工作中,我们通过利用延迟单调性来解决可扩展性挑战——不同设备上的架构延迟排名通常是相关的。当存在强延迟单调性时,我们可以在新的目标设备上重用搜索一个代理设备的架构,而不会失去最优性。在缺乏强延迟单调性的情况下,我们提出了一种有效的代理自适应技术来显著提高延迟单调性。最后,我们验证了我们的方法,并在多个主流搜索空间(包括MobileNet-V2、MobileNet-V3、NAS-Bench-201、ProxylessNAS和FBNet)上使用不同平台的设备进行了实验。我们的结果强调,通过只使用一个代理设备,我们可以找到与现有的每设备NAS几乎相同的pareto最优架构,同时避免了为每个设备构建延迟预测器的高昂成本。
Convolutional neural networks (CNNs) are used in numerous real-world applications such as vision-based autonomous driving and video content analysis. To run CNN inference on various target devices, hardware-aware neural architecture search (NAS) is crucial. A key requirement of efficient hardware-aware NAS is the fast evaluation of inference latencies in order to rank different architectures. While building a latency predictor for each target device has been commonly used in state of the art, this is a very time-consuming process, lacking scalability in the presence of extremely diverse devices. In this work, we address the scalability challenge by exploiting latency monotonicity --- the architecture latency rankings on different devices are often correlated. When strong latency monotonicity exists, we can re-use architectures searched for one proxy device on new target devices, without losing optimality. In the absence of strong latency monotonicity, we propose an efficient proxy adaptation technique to significantly boost the latency monotonicity. Finally, we validate our approach and conduct experiments with devices of different platforms on multiple mainstream search spaces, including MobileNet-V2, MobileNet-V3, NAS-Bench-201, ProxylessNAS and FBNet. Our results highlight that, by using just one proxy device, we can find almost the same Pareto-optimal architectures as the existing per-device NAS, while avoiding the prohibitive cost of building a latency predictor for each device.
DOI: --
发表时间: 2021-03
期刊: ArXiv
影响因子: --
作者:
Chaojian Li;Zhongzhi Yu;Yonggan Fu;Yongan Zhang;Yang Zhao;Haoran You;Qixuan Yu;Yue Wang;Yingyan Lin
通讯作者: Chaojian Li;Zhongzhi Yu;Yonggan Fu;Yongan Zhang;Yang Zhao;Haoran You;Qixuan Yu;Yue Wang;Yingyan Lin
nn-Meter:在各种边缘设备上实现深度学习模型推理的准确延迟预测
DOI: 10.1145/3458864.3467882
发表时间: 2021
期刊: Proceedings of the 19th Annual International Conference on Mobile Systems, Applications, and Services
影响因子: --
作者:
L. Zhang;S. Han;Jianyu Wei;Ningxin Zheng;Ting Cao;Yuqing Yang;Yunxin Liu
通讯作者: Yunxin Liu
DOI: --
发表时间: 2016-11
期刊: ArXiv
影响因子: --
作者:
Barret Zoph;Quoc V. Le
通讯作者: Barret Zoph;Quoc V. Le
DOI: 10.1109/tcad.2020.3012863
发表时间: 2020-07
影响因子: 2.9
作者:
Weiwen Jiang;Lei Yang;Sakyasingha Dasgupta;J. Hu;Yiyu Shi
通讯作者: Weiwen Jiang;Lei Yang;Sakyasingha Dasgupta;J. Hu;Yiyu Shi
Gables:移动 SoC 的 Roofline 模型
DOI: 10.1109/hpca.2019.00047
发表时间: 2019
期刊: 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA
影响因子: --
作者:
Hill, Mark;Janapa Reddi, Vijay
通讯作者: Janapa Reddi, Vijay