Automatic Kernel Fusion for Image Processing DSLs

Automatic Kernel Fusion for Image Processing DSLs
复制标题

用于图像处理 DSL 的自动内核融合

DOI:
10.1145/3207719.3207723
复制
发表时间:
2018
期刊:
Proceedings of the 21st International Workshop on Software and Compilers for Embedded Systems
影响因子:
--
通讯作者:
J. Teich
J. Teich
中科院分区:
--
文献类型:
--
作者:
B. Qiao;O. Reiche;F. Hannig;J. Teich

文献摘要

参考文献

被引文献

相似文献

在诸如图形处理单元(GPU)之类的硬件加速器上编程图像处理算法通常表现出软件可移植性和性能可移植性之间的权衡。领域特定语言(DSLs)已被证明是一个很有前途的补救措施,它使优化和高效的代码从一个简洁的,高层次的算法representation.The本文的范围是一个优化框架的图像处理DSLs的形式的源到源编译器。为了科普GPU应用程序中内核间通信受限的问题,研究了内核融合作为提高时间局部性的主要优化技术。为了实现自动内核融合,我们分析了算法中每个内核的融合性,在数据依赖性,资源利用率和并行粒度。通过将获得的信息与DSL中捕获的特定于领域的知识相结合,提出了一种自动融合合适的内核的方法,并将其集成到开源DSL框架中。新的内核融合技术进行了评估,两个基于滤波器的图像处理应用程序,其中获得高达1.60的加速比为NVIDIA Geforce 745显卡的目标。
Programming image processing algorithms on hardware accelerators such as graphics processing units (GPUs) often exhibits a trade-off between software portability and performance portability. Domain-specific languages (DSLs) have proven to be a promising remedy, which enable optimizations and generation of efficient code from a concise, high-level algorithm representation.The scope of this paper is an optimization framework for image processing DSLs in the form of a source-to-source compiler. To cope with the inter-kernel communication bound via global memory for GPU applications, kernel fusion is investigated as a primary optimization technique to improve temporal locality. In order to enable automatic kernel fusion, we analyze the fusibility of each kernel in the algorithm, in terms of data dependencies, resource utilization, and parallelism granularity. By combining the obtained information with the domain-specific knowledge captured in the DSL, a method to automatically fuse the suitable kernels is proposed and integrated into an open source DSL framework. The novel kernel fusion technique is evaluated on two filter-based image processing applications, for which speedups of up to 1.60 are obtained for an NVIDIA Geforce 745 graphics card target.
从特定于领域的语言为基于 C 的 HLS 硬件加速器生成代码
DOI: 10.1145/2656075.2656081
发表时间: 2014
期刊: 2014 International Conference on Hardware/Software Codesign and System Synthesis (CODES+ISSS)
影响因子: --
作者:
O. Reiche;M. Schmid;F. Hannig;R. Membarth;J. Teich
通讯作者: J. Teich
HIPAcc:用于图像处理的特定领域语言和编译器
DOI: 10.1109/tpds.2015.2394802
发表时间: 2016
影响因子: 5.3
作者:
R. Membarth;O. Reiche;F. Hannig;J. Teich;M. Körner;W. Eckert
通讯作者: W. Eckert
使用 Hipacc 生成基于 FPGA 的图像处理加速器:(特邀论文)
DOI: 10.1109/iccad.2017.8203894
发表时间: 2017
期刊: 2017 IEEE/ACM International Conference on Computer-Aided Design (ICCAD)
影响因子: --
作者:
O. Reiche. M. A. Özkan;R. Membarth;J. Teich;F. Hannig
通讯作者: F. Hannig