FPGA-Based Inter-layer Pipelined Accelerators for Filter-Wise Weight-Balanced Sparse Fully Convolutional Networks with Overlapped Tiling

FPGA-Based Inter-layer Pipelined Accelerators for Filter-Wise Weight-Balanced Sparse Fully Convolutional Networks with Overlapped Tiling
复制标题

DOI:
10.1007/s11265-021-01642-6
复制
发表时间:
2021-02
期刊:
Journal of Signal Processing Systems
影响因子:
--
通讯作者:
Masayuki Shimoda;Youki Sada;Hiroki Nakahara
Masayuki Shimoda;Youki Sada;Hiroki Nakahara
中科院分区:
其他
文献类型:
--
作者:
Masayuki Shimoda;Youki Sada;Hiroki Nakahara

文献摘要

相似文献

卷积神经网络 (CNN) 在执行计算机视觉任务时表现出最先进的性能。 CNN 需要高速、低功耗、高精度的硬件来适应各种场景,例如边缘环境。然而,权重的数量如此之大,以至于嵌入式系统由于片上内存有限而无法存储它们。使用不同的方法来最小化输入图像大小以进行实时处理,但这会导致精度大幅下降。尽管提出了剪枝稀疏CNN和特殊加速器,但随机访问的要求需要大量的宽复用器来实现高度并行性,这变得更加复杂并且不适合FPGA实现。为了解决这个问题,我们提出使用基于蒸馏和块 RAM (BRAM) 的零权重跳跃加速器进行过滤式剪枝。它消除了权重,使每个过滤器具有相同数量的非零权重,通过蒸馏执行再训练,同时保持相当的精度。此外,过滤式修剪使我们的加速器能够利用过滤器间的并行性,其中一层的处理块以简单的架构同时执行过滤器。我们还提出了一种重叠切片算法,其中重叠提取切片,以防止精度下降和存储高分辨率图像的 BRAM 的高利用率。我们使用语义分割任务进行的评估表明,与桌面 GPU 相比,我们的 FPGA 设计的速度提高了 1.8 倍,功效提高了 18.0 倍。此外,与传统的FPGA实现相比,加速比和精度分别提高了1.09倍和6.6个点。因此,我们的方法对于 FPGA 实现非常有用,并且对于嵌入式系统中的应用表现出相当高的准确性。
Convolutional neural networks (CNNs) exhibit state-of-the-art performance while performing computer-vision tasks. CNNs require high-speed, low-power, and high-accuracy hardware for various scenarios, such as edge environments. However, the number of weights is so large that embedded systems cannot store them owing to their limited on-chip memory. A different method is used to minimize the input image size, for real-time processing, but it causes a considerable drop in accuracy. Although pruned sparse CNNs and special accelerators are proposed, the requirement of random access incurs a large number of wide multiplexers for a high degree of parallelism, which becomes more complicated and unsuitable for FPGA implementation. To address this problem, we proposefilter-wise pruning with distillationand block RAM (BRAM)-based zero-weight skipping accelerator. It eliminates weights such that each filter has the same number of nonzero weights, performing retraining with distillation, while retaining comparable accuracy. Further, filter-wise pruning enables our accelerator to exploitinter-filter parallelism, where a processing block for a layer executes filters concurrently, with a straightforward architecture. We also propose anoverlapped tiling algorithm, where tiles are extracted with overlap to prevent both accuracy degradation and high utilization of BRAMs storing high-resolution images. Our evaluation using semantic-segmentation tasks showed a 1.8 times speedup and 18.0 times increase in power efficiency of our FPGA design compared with a desktop GPU. Additionally, compared with the conventional FPGA implementation, the speedup and accuracy improvement were 1.09 times and 6.6 points, respectively. Therefore, our approach is useful for FPGA implementation and exhibits considerable accuracy for applications in embedded systems.