Generating GPU Code from a High-Level Representation for Image Processing Kernels

Generating GPU Code from a High-Level Representation for Image Processing Kernels
复制标题

DOI:
10.1007/978-3-642-29737-3_31
复制
发表时间:
2011-08
期刊:
--
影响因子:
--
通讯作者:
Richard Membarth;Anton Lokhmotov;J. Teich
Richard Membarth;Anton Lokhmotov;J. Teich
中科院分区:
其他
文献类型:
--
作者:
Richard Membarth;Anton Lokhmotov;J. Teich

文献摘要

被引文献

相似文献

我们提出了一个基于解耦的访问/执行元数据的图像处理内核表示框架,允许程序员指定内核的执行约束和内存访问模式。该框架通过有效的设备相关优化,例如用于内存合并的全局内存填充和最优内存带宽利用率,执行以高级框架特定C++类表示的内核到低级CUDA或OpenCL代码的源代码到源代码的转换。我们在几个图像过滤器上对该框架进行了评估,将生成的代码与流行的OpenCV库中高度优化的CPU和GPU版本进行了比较。
We present a framework for representing image processing kernels based on decoupled access/execute metadata, which allow the programmer to specify both execution constraints and memory access pattern of a kernel. The framework performs source-to-source translation of kernels expressed in high-level framework-specific C++ classes into low-level CUDA or OpenCL code with effective device-dependent optimizations such as global memory padding for memory coalescing and optimal memory bandwidth utilization. We evaluate the framework on several image filters, comparing generated code against highly-optimized CPU and GPU versions in the popular OpenCV library.