Sparse-YOLO: Hardware/Software Co-Design of an FPGA Accelerator for YOLOv2

Sparse-YOLO: Hardware/Software Co-Design of an FPGA Accelerator for YOLOv2
复制标题

DOI:
10.1109/access.2020.3004198
复制
发表时间:
2020-01-01
期刊:
影响因子:
3.9
通讯作者:
Wang, Dong
Wang, Dong
中科院分区:
计算机科学3区
文献类型:
--
作者:
Wang, Zixiao;Xu, Ke;Wang, Dong

文献摘要

被引文献

相似文献

基于卷积的神经网络(CNN)基于对象检测算法在许多应用领域中占主导地位,因为它们的准确性优于传统方案。其中,您只看一次(Yolo)是最受欢迎的检测框架之一,在速度和准确性之间表现出最佳的折衷。但是,由于CNN的固有较高的计算工作量,在靶向高通量处理以低成本的能源消耗时,它仍然具有挑战性。在本文中,我们提出了针对CPU+基于FPGA的异构平台的硬件/软件(HW/SW)的共同设计方法。首先,我们将一种新颖的稀疏卷积算法扩展到YOLOV2框架,然后基于异步执行的并行卷积核心开发资源有效的FPGA加速器体系结构。其次,引入了算法级优化方案,包括硬件感知的神经网络修剪,聚类和量化,从而成功地将Yolov2算法的计算工作量减少了7倍。最后,提出了针对基于FPGA的加速器设计的端到端设计空间探索流,并研究了两种HW/SW分区策略。实验结果表明,在211 MHz的工作频率下,我们的设计可以在Intel Arria-10 GX1150 FPGA上实现2.13顶部(72.5 fps)的峰值吞吐量,而Pascal VOC2007数据集的检测准确性为74.45。
Convolutional neural network (CNN) based object detection algorithms are becoming dominant in many application fields due to their superior accuracy advantage over traditional schemes. Among them, You Look Only Once (YOLO) is one of the most popular detection frameworks that show best trade-offs between speed and accuracy. However, due to the intrinsic high computational workload of CNN, it is still challenging when targeting high-throughput processing with low cost in energy consumption. In this paper, we propose a hardware/software (HW/SW) co-design methodology targeting CPU+FPGA-based heterogeneous platforms. Firstly, we extend a novel sparse convolution algorithm to the YOLOv2 framework, and then develop a resource-efficient FPGA accelerator architecture based on asynchronously executed parallel convolution cores. Secondly, algorithm-level optimization schemes, including hardware-aware neural network pruning, clustering and quantization are introduced, which successfully save the computational workload of the YOLOv2 algorithm by 7 times. Finally, an end-to-end design space exploration flow for FPGA-based accelerator design is presented and two HW/SW partition strategies are studied and implemented. Experimental results show that our design can achieve a peak throughput of 2.13 TOPS (72.5 fps) on an Intel Arria-10 GX1150 FPGA under the working frequency of 211 MHz, while the detection accuracy is 74.45 on the PASCAL VOC2007 dataset.