High-Throughput CNN Inference on Embedded ARM Big.LITTLE Multicore Processors

High-Throughput CNN Inference on Embedded ARM Big.LITTLE Multicore Processors
复制标题

DOI:
10.1109/tcad.2019.2944584
复制
发表时间:
2019-03
影响因子:
2.9
通讯作者:
Siqi Wang;Gayathri Ananthanarayanan;Yifan Zeng;Neeraj Goel;A. Pathania;T. Mitra
Siqi Wang;Gayathri Ananthanarayanan;Yifan Zeng;Neeraj Goel;A. Pathania;T. Mitra
中科院分区:
计算机科学3区
文献类型:
--
作者:
Siqi Wang;Gayathri Ananthanarayanan;Yifan Zeng;Neeraj Goel;A. Pathania;T. Mitra

文献摘要

被引文献

相似文献

物联网边缘智能需要在边缘设备本身进行卷积神经网络 (CNN) 推理。 ARM big.LITTLE 架构是流行的商业边缘设备的核心。它由分组为多个同构集群的单 ISA 异构核心组成,可实现功耗和性能的权衡。所有核心预计将同时用于推理,以获得最大吞吐量。然而,跨集群的卷积核并行计算涉及的高通信开销不利于吞吐量。我们提出了一个名为 Pipe-it 的替代框架,它采用流水线设计来跨集群分割卷积层,同时限制它们各自内核的并行化到指定的集群。我们开发了一个性能预测模型,仅利用卷积层描述符来预测所有允许的核心配置(类型和数量)上每层的执行时间。然后,Pipe-it 利用预测结果,使用高效的设计空间探索算法创建平衡管道。 Pipe-it 的平均吞吐量比之前的最高吞吐量高出 39%。
Internet of Things edge intelligence requires convolutional neural network (CNN) inference to take place in the edge devices itself. ARM big.LITTLE architecture is at the heart of prevalent commercial edge devices. It comprises of single-ISA heterogeneous cores grouped into multiple homogeneous clusters that enable power and performance tradeoffs. All cores are expected to be simultaneously employed in inference to attain maximal throughput. However, high communication overhead involved in parallelization of computations from convolution kernels across clusters is detrimental to throughput. We present an alternative framework called Pipe-it that employs pipelined design to split convolutional layers across clusters while limiting parallelization of their respective kernels to the assigned cluster. We develop a performance-prediction model that utilizes only the convolutional layer descriptors to predict the execution time of each layer individually on all permitted core configurations (type and count). Pipe-it then exploits the predictions to create a balanced pipeline using an efficient design space exploration algorithm. Pipe-it on average results in a 39% higher throughput than the highest antecedent throughput.