Standing on the Shoulders of Giants: Hardware and Neural Architecture Co-Search With Hot Start

Standing on the Shoulders of Giants: Hardware and Neural Architecture Co-Search With Hot Start
复制标题

DOI:
10.1109/tcad.2020.3012863
复制
发表时间:
2020-07
影响因子:
2.9
通讯作者:
Weiwen Jiang;Lei Yang;Sakyasingha Dasgupta;J. Hu;Yiyu Shi
Weiwen Jiang;Lei Yang;Sakyasingha Dasgupta;J. Hu;Yiyu Shi
中科院分区:
计算机科学3区
文献类型:
--
作者:
Weiwen Jiang;Lei Yang;Sakyasingha Dasgupta;J. Hu;Yiyu Shi

文献摘要

被引文献

相似文献

从给定数据集自动生成人工智能(AI)解决方案的硬件和神经架构协同搜索有望促进AI的普及;然而,当前协同搜索框架针对一种目标硬件所需的时间约为数百个GPU小时。这限制了此类框架在普通硬件上的使用。现有协同搜索框架效率低下的根本原因是它们从“冷”状态开始(即从头开始搜索)。在本文中,我们提出了一种新颖的框架,即HotNAS,它基于一组现有的预训练模型(也称为模型库)从“热”状态开始,以避免冗长的训练时间。这样,搜索时间可以从200个GPU小时减少到不到3个GPU小时。在HotNAS中,除了硬件设计空间和神经架构搜索空间外,我们还进一步整合了一个压缩空间,以便在协同搜索过程中进行模型压缩,这为降低延迟创造了新的机会,但也带来了挑战。关键挑战之一是上述所有搜索空间相互耦合,例如,没有硬件设计支持,压缩可能无法实现。为了解决这个问题,HotNAS构建了一系列工具来设计支持压缩的硬件,并在此基础上开发了一个全局优化器,以自动对所有相关搜索空间进行协同搜索。在ImageNet数据集和赛灵思FPGA上进行的实验表明,在5毫秒的时间限制内,与现有模型相比,HotNAS生成的神经架构在Top - 1准确率上可提高多达5.79%,在Top - 5准确率上可提高3.97%。
Hardware and neural architecture co-search that automatically generates artificial intelligence (AI) solutions from a given dataset are promising to promote AI democratization; however, the amount of time that is required by current co-search frameworks is in the order of hundreds of GPU hours for one target hardware. This inhibits the use of such frameworks on commodity hardware. The root cause of the low efficiency in existing co-search frameworks is the fact that they start from a “cold” state (i.e., search from scratch). In this article, we propose a novel framework, namely, HotNAS, that starts from a “hot” state based on a set of existing pretrained models (also known as model zoo) to avoid lengthy training time. As such, the search time can be reduced from 200 GPU hours to less than 3 GPU hours. In HotNAS, in addition to hardware design space and neural architecture search space, we further integrate a compression space to conduct model compressing during the co-search, which creates new opportunities to reduce latency, but also brings challenges. One of the key challenges is that all of the above search spaces are coupled with each other, e.g., compression may not work without hardware design support. To tackle this issue, HotNAS builds a chain of tools to design hardware to support compression, based on which a global optimizer is developed to automatically co-search all the involved search spaces. Experiments on ImageNet dataset and Xilinx FPGA show that, within the timing constraint of 5 ms, neural architectures generated by HotNAS can achieve up to 5.79% Top-1 and 3.97% Top-5 accuracy gain, compared with the existing ones.