Loop coarsening in C-based High-Level Synthesis

Loop coarsening in C-based High-Level Synthesis
复制标题

DOI:
10.1109/asap.2015.7245730
复制
发表时间:
2015-07
期刊:
2015 IEEE 26th International Conference on Application-specific Systems, Architectures and Processors (ASAP)
影响因子:
--
通讯作者:
Moritz Schmid;Oliver Reiche;Frank Hannig;J. Teich
Moritz Schmid;Oliver Reiche;Frank Hannig;J. Teich
中科院分区:
其他
文献类型:
--
作者:
Moritz Schmid;Oliver Reiche;Frank Hannig;J. Teich

文献摘要

被引文献

相似文献

现有的高级综合工具在利用指令级并行性方面做得很好,而现场可编程门阵列(FPGA)对数据级并行性的支持却非常有限。这项工作探讨了利用FPGA上的DLP使用基于C的HLS的图像过滤器和流管道,包括点和本地运营商的代码生成。除了众所周知的循环切片技术之外,我们还提出了循环粗化,它可以提供上级性能和可扩展性。循环平铺对应于将图像分割成单独的区域,然后由复制的加速器并行处理。对于数据流,这还需要生成用于图像数据分发的粘合逻辑。相反,循环粗化允许并行处理多个像素,从而仅在单个加速器内复制内核运算符。我们通过循环粗化增强了异构域特定语言(DSL)框架HIPAcc的FPGA后端,并将由此产生的FPGA加速器与图形处理单元(GPU)的高度优化的软件实现进行比较,所有这些都是从完全相同的代码库生成的。此外,我们展示了算法开发的代码生成的优势,概述了如何通过HIPAcc启用设计空间探索可以产生比手工编码的VHDL更有效的实现。
Current tools for High-Level Synthesis (HLS) excel at exploiting Instruction-Level Parallelism (ILP), the support for Data-Level Parallelism (DLP), one of the key advantages of Field Programmable Gate Arrays (FPGAs), is in contrast very limited. This work examines the exploitation of DLP on FPGAs using code generation for C-based HLS of image filters and streaming pipelines, consisting of point and local operators. In addition to well known loop tiling techniques, we propose loop coarsening, which delivers superior performance and scalability. Loop tiling corresponds to splitting an image into separate regions, which are then processed in parallel by replicated accelerators. For data streaming, this also requires the generation of glue logic for the distribution of image data. Conversely, loop coarsening allows to process multiple pixels in parallel, whereby only the kernel operator is replicated within a single accelerator. We augment the FPGA back end of the heterogeneous Domain-Specific Language (DSL) framework HIPAcc by loop coarsening and compare the resulting FPGA accelerators to highly optimized software implementations for Graphics Processing Units (GPUs), all generated from the exact same code base. Moreover, we demonstrate the advantages of code generation for algorithm development by outlining how design space exploration enabled by HIPAcc can yield a more efficient implementation than hand-coded VHDL.