MKPipe: a compiler framework for optimizing multi-kernel workloads in OpenCL for FPGA

MKPipe: a compiler framework for optimizing multi-kernel workloads in OpenCL for FPGA
复制标题

MKPipe:用于优化 OpenCL for FPGA 中的多内核工作负载的编译器框架

DOI:
10.1145/3392717.3392757
复制
发表时间:
2020
期刊:
The 34th ACM International Conference on Supercomputing
影响因子:
--
通讯作者:
Zhou, Huiyang
Zhou, Huiyang
中科院分区:
--
文献类型:
--
作者:
Liu, Ji;Kafi, Abdullah-Al;Shen, Xipeng;Zhou, Huiyang

文献摘要

参考文献

被引文献

相似文献

OpenCL for FPGA 使开发人员能够使用类似于处理器的编程模型来设计 FPGA。最近的工作表明,OpenCL 级别的代码优化对于实现高计算效率非常重要。然而,现有的工作要么主要集中在优化单个内核上,要么仅仅依赖于通道来设计多内核管道。在本文中,我们提出了一种源到源编译器框架 MKPipe,用于优化 FPGA 中 OpenCL 中的多内核工作负载。除了通道之外,我们还提出了支持多内核管道的新方案。我们的优化编译器采用系统方法来探索这些优化方法的权衡。为了在内核执行之间实现更有效的重叠,我们还提出了一种新颖的工作项/工作组 ID 重新映射技术。此外,我们提出了吞吐量平衡和资源平衡的新算法,以调整多内核工作负载中各个内核的优化。我们的结果表明,我们的编译器优化多内核比基准实现了高达 3.6 倍(平均 1.4 倍)的加速,其中内核已经单独优化。
OpenCL for FPGA enables developers to design FPGAs using a programming model similar for processors. Recent works have shown that code optimization at the OpenCL level is important to achieve high computational efficiency. However, existing works either focus primarily on optimizing single kernels or solely depend on channels to design multi-kernel pipelines. In this paper, we propose a source-to-source compiler framework, MKPipe, for optimizing multi-kernel workloads in OpenCL for FPGA. Besides channels, we propose new schemes to enable multi-kernel pipelines. Our optimizing compiler employs a systematic approach to explore the tradeoffs of these optimizations methods. To enable more efficient overlapping between kernel execution, we also propose a novel workitem/workgroup-id remapping technique. Furthermore, we propose new algorithms for throughput balancing and resource balancing to tune the optimizations upon individual kernels in the multi-kernel workloads. Our results show that our compiler-optimized multi-kernels achieve up to 3.6x (1.4x on average) speedup over the baseline, in which the kernels have already been optimized individually.
在 OpenCL 中针对 FPGA 调整 Stencil 代码
DOI: --
发表时间: 2016
期刊: ICCD
影响因子: --
作者:
Q. Jia;Huiyang Zhou
通讯作者: Huiyang Zhou
通过 OpenCL 加速 FPGA 上的工作负载:OpenDwarfs 案例研究
DOI: --
发表时间: 2016
期刊:
影响因子: --
作者:
Anshuman Verma;A. Helal;K. Krommydas;Wu
通讯作者: Wu
DOI: 10.1109/fpt.2018.00018
发表时间: 2018
期刊: 2018 International Conference on Field-Programmable Technology (FPT)
影响因子: --
作者:
A. Sanaullah;Rushi Patel;M. Herbordt
通讯作者: M. Herbordt
Stratix® 10:14nm FPGA,提供 1GHz
DOI: 10.1109/hotchips.2015.7477458
发表时间: 2015
期刊: 2015 IEEE Hot Chips 27 Symposium (HCS)
影响因子: --
作者:
M. Hutton
通讯作者: M. Hutton
OpenCL for HPC with FPGA:分子静电学案例研究
DOI: 10.1109/hpec.2017.8091078
发表时间: 2017
期刊: 2017 IEEE High Performance Extreme Computing Conference (HPEC)
影响因子: --
作者:
Chen Yang;Jiayi Sheng;Rushi Patel;A. Sanaullah;Vipin Sachdeva;M. Herbordt
通讯作者: M. Herbordt