Throughput oriented FPGA overlays using DSP blocks

Throughput oriented FPGA overlays using DSP blocks
复制标题

使用 DSP 块的面向吞吐量的 FPGA 叠加

DOI:
10.3850/9783981537079_0685
复制
发表时间:
2016
期刊:
2016 Design, Automation & Test in Europe Conference & Exhibition (DATE)
影响因子:
--
通讯作者:
Suhaib A. Fahmy
Suhaib A. Fahmy
中科院分区:
--
文献类型:
--
作者:
A. Jain;D. Maskell;Suhaib A. Fahmy

文献摘要

被引文献

相似文献

设计生产率是阻碍FPGA主流采用的主要问题。覆盖架构已经成为应对这一挑战的一种可能的解决方案,它提供了快速编译和类似软件的可编程性。然而,由于对底层FPGA架构的考虑有限,覆盖层通常会受到面积和性能开销的影响。这些覆盖通常具有有限的大小,仅支持相对较小的计算内核。本文探讨了开发更大,更有效的,覆盖使用多个DSP块,然后最大限度地利用内核的多个实例同时映射到覆盖利用内核级并行的可能性。我们展示了可实现的覆盖尺寸和覆盖利用率的显着改善,与现有的覆盖架构相比,覆盖瓦片要求减少了近70%,工作频率超过300 MHz,内核吞吐量近60 GOPS。
Design productivity is a major concern preventing the mainstream adoption of FPGAs. Overlay architectures have emerged as one possible solution to this challenge, offering fast compilation and software-like programmability. However, overlays typically suffer from area and performance overheads due to limited consideration for the underlying FPGA architecture. These overlays have often been of limited size, supporting only relatively small compute kernels. This paper examines the possibility of developing larger, more efficient, overlays using multiple DSP blocks and then maximising utilisation by mapping multiple instances of kernels simultaneously onto the overlay to exploit kernel level parallelism. We show a significant improvement in achievable overlay size and overlay utilisation, with a reduction of almost 70% in the overlay tile requirement compared to existing overlay architectures, an operating frequency in excess of 300 MHz, and kernel throughputs of almost 60 GOPS.