TEA-DNN: the Quest for Time-Energy-Accuracy Co-optimized Deep Neural Networks

TEA-DNN: the Quest for Time-Energy-Accuracy Co-optimized Deep Neural Networks
复制标题

TEA-DNN:寻求时间-能量-精度协同优化的深度神经网络

DOI:
10.1109/islped.2019.8824934
复制
发表时间:
2018
期刊:
2019 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED)
影响因子:
--
通讯作者:
M. Sabry
M. Sabry
中科院分区:
--
文献类型:
--
作者:
Lile Cai;Anne;Arthur Herbout;Chuan;Jie Lin;V. Chandrasekhar;M. Sabry

文献摘要

参考文献

被引文献

相似文献

嵌入式深度学习平台见证了两个同时改进。首先,通过使用自动化的神经结构搜索(NAS)算法来确定CNN结构,卷积神经网络(CNN)的准确性已得到显着提高。其次,与GPU相比,为CNN开发硬件加速器的兴趣越来越大。这种嵌入的深度学习平台在计算资源和记忆访问带宽的数量上有所不同,这将影响CNN的性能和能耗。因此,在网络体系结构搜索中考虑可用的硬件资源至关重要。为此,我们介绍了针对嵌入式体系结构中CNN工作负载的执行时间,能耗和分类精度的多目标优化的NAS算法。 Tea-dnn在跨精度,执行时间和能源消耗探索帕累托最佳曲线时,利用嵌入式硬件的能量和执行时间测量,并且不需要额外的努力来对基础硬件进行建模。我们在实际嵌入式平台(Nvidia Jetson TX2和Intel Movidius Neural Compute Stick)上应用茶-DNN进行图像分类。我们强调了帕累托最佳的操作点,这些操作点强调了在搜索过程中明确考虑硬件特征的必要性。据我们所知,这是使用硬件上的实际测量值来获得客观值的一系列硬件平台上对帕累托最佳模型的最全面研究。
Embedded deep learning platforms have witnessed two simultaneous improvements. First, the accuracy of convolutional neural networks (CNNs) has been significantly improved through the use of automated neural-architecture search (NAS) algorithms to determine CNN structure. Second, there has been increasing interest in developing hardware accelerators for CNNs that provide improved inference performance and energy consumption compared to GPUs. Such embedded deep learning platforms differ in the amount of compute resources and memory-access bandwidth, which would affect performance and energy consumption of CNNs. It is therefore critical to consider the available hardware resources in the network architecture search. To this end, we introduce TEA-DNN, a NAS algorithm targeting multi-objective optimization of execution time, energy consumption, and classification accuracy of CNN workloads on embedded architectures. TEA-DNN leverages energy and execution time measurements on embedded hardware when exploring the Pareto-optimal curves across accuracy, execution time, and energy consumption and does not require additional effort to model the underlying hardware. We apply TEA-DNN for image classification on actual embedded platforms (NVIDIA Jetson TX2 and Intel Movidius Neural Compute Stick). We highlight the Pareto-optimal operating points that emphasize the necessity to explicitly consider hardware characteristics in the search process. To the best of our knowledge, this is the most comprehensive study of Pareto-optimal models across a range of hardware platforms using actual measurements on hardware to obtain objective values.
DOI: --
发表时间: 2016-11
期刊: ArXiv
影响因子: --
作者:
Barret Zoph;Quoc V. Le
通讯作者: Barret Zoph;Quoc V. Le