A Common Backend for Hardware Acceleration on FPGA

A Common Backend for Hardware Acceleration on FPGA
复制标题

FPGA 硬件加速的通用后端

DOI:
10.1109/iccd.2017.75
复制
发表时间:
2017
期刊:
2017 IEEE International Conference on Computer Design (ICCD)
影响因子:
--
通讯作者:
M. Santambrogio
M. Santambrogio
中科院分区:
--
文献类型:
--
作者:
Emanuele Del Sozzo;Riyadh Baghdadi;Saman P. Amarasinghe;M. Santambrogio

文献摘要

被引文献

相似文献

现场可编程门阵列(现场可编程门阵列)是可配置的集成电路,能够在性能、功耗和灵活性方面与其他架构(如CPU、GPU和ASIC)进行良好的折衷。然而,使用现场可编程门阵列的主要缺点是其学习曲线陡峭。这个问题的一个新出现的解决方案是用域特定语言(DSL)编写算法,并让DSL编译器生成针对FGA的高效代码。这项工作提出了Frost,这是一个统一的后端,使不同的DSL编译器能够针对FPGA架构。与其他面向FPGA的代码生成框架不同,Frost利用了一种调度共语言,使用户能够完全控制要应用哪些优化以生成高效的代码(例如循环流水线、数组分区、矢量化)。Frost首先对输入抽象语法树(AST)进行分析和操作,以便应用面向FPGA的转换和优化,然后生成适合于高级综合(HLS)工具的C/C++实现。最后,利用Xilinx SDAccel工具链对合肥光源的输出进行了综合,并在目标现场可编程门阵列上实现。实验结果表明,与相同算法在CPU上的臭氧优化实现相比,加速比提高了15&#。
Field Programmable Gate Arrays (FPGAs) are configurable integrated circuits able to provide a good trade-off in terms of performance, power consumption, and flexibility with respect to other architectures, like CPUs, GPUs and ASICs. The main drawback in using FPGAs, however, is their steep learning curve. An emerging solution to this problem is to write algorithms in a Domain Specific Language (DSL) and to let the DSL compiler generate efficient code targeting FPGAs. This work proposes FROST, a unified backend that enables different DSL compilers to target FPGA architectures. Differently from other code generation frameworks targeting FPGA, FROST exploits a scheduling co-language that enables users to have full control over which optimizations to apply in order to generate efficient code (e.g. loop pipelining, array partitioning, vectorization). At first, FROST analyzes and manipulates the input Abstract Syntax Tree (AST) in order to apply FPGA-oriented transformations and optimizations, then generates a C/C++ implementation suitable for High-Level Synthesis (HLS) tools. Finally, the output of HLS phase is synthesized and implemented on the target FPGA using Xilinx SDAccel toolchain. The experimental results show a speedup up of 15× with respect to O3-optimized implementations of the same algorithms on CPU.