Constructive Synthesis of Memory-Intensive Accelerators for FPGA From Nested Loop Kernels

Constructive Synthesis of Memory-Intensive Accelerators for FPGA From Nested Loop Kernels
复制标题

从嵌套循环内核构造性综合 FPGA 的内存密集型加速器

DOI:
--
复制
发表时间:
2016
影响因子:
5.4
通讯作者:
J. McAllister
J. McAllister
中科院分区:
工程技术1区
文献类型:
--
作者:
Matthew Milford;J. McAllister

文献摘要

被引文献

相似文献

现场编程的门阵列是自定义加速器的理想主机,用于信号,图像和数据处理,但需求手动寄存器传输级别的设计(如果需要高性能和低成本)。高级合成减轻了这种设计负担,但需要手动设计复杂的片上和芯片内存储器体系结构,这是视频处理等应用程序的主要限制。本文提出了一种解决这一缺点的方法。描述了一个可以得出此类加速器的建设性过程,包括从C描述中的芯片和外存储器存储,以便满足用户定义的吞吐量约束。通过采用一种新颖的面向语句的方法,得出数据流中间模型并用于支持简单的方法用于在芯片缓冲区分区,自定义芯片上存储器层次结构的推导和体系结构转换,以确保满足用户定义的吞吐量约束最低成本。当应用于加速器以进行完整的搜索运动估计,矩阵乘法,SOBEL边缘检测和快速傅立叶变换时,它显示了如何在现有商业HLS工具之前达到实时性能到一个数量级,而包括所有必需内存基础设施。此外,进行了优化,将芯片缓冲能力和物理资源成本降低了96%和75%,同时保持实时性能。
Field-programmable gate arrays are ideal hosts to custom accelerators for signal, image, and data processing but demand manual register transfer level design if high performance and low cost are desired. High-level synthesis reduces this design burden but requires manual design of complex on-chip and off-chip memory architectures, a major limitation in applications such as video processing. This paper presents an approach to resolve this shortcoming. A constructive process is described that can derive such accelerators, including on- and off-chip memory storage from a C description such that a user-defined throughput constraint is met. By employing a novel statement-oriented approach, dataflow intermediate models are derived and used to support simple approaches for on-/off-chip buffer partitioning, derivation of custom on-chip memory hierarchies and architecture transformation to ensure user-defined throughput constraints are met with minimum cost. When applied to accelerators for full search motion estimation, matrix multiplication, Sobel edge detection, and fast Fourier transform, it is shown how real-time performance up to an order of magnitude in advance of existing commercial HLS tools is enabled whilst including all requisite memory infrastructure. Further, optimizations are presented that reduce the on-chip buffer capacity and physical resource cost by up to 96% and 75%, respectively, whilst maintaining real-time performance.