Automatic Optimization of In-Flight Memory Transactions for GPU Accelerators Based on a Domain-Specific Language for Medical Imaging

Automatic Optimization of In-Flight Memory Transactions for GPU Accelerators Based on a Domain-Specific Language for Medical Imaging
复制标题

基于医学成像领域特定语言的 GPU 加速器运行中内存事务的自动优化

DOI:
--
复制
发表时间:
2012
期刊:
International Symposium on Parallel and Distributed Computing
影响因子:
--
通讯作者:
Wieland Eckert
Wieland Eckert
中科院分区:
--
文献类型:
--
作者:
Richard Membarth;Frank Hannig;J. Teich;M. Körner;Wieland Eckert

文献摘要

被引文献

相似文献

GPU加速器的有效内存带宽利用率对于内存绑定的应用程序至关重要。在医学成像中,许多内核的性能受到可用内存带宽的限制,因为每个像素仅执行少量操作。对于此类内核,只能利用GPU加速器提供的计算功率的一小部分,并且由内存带宽预先确定性能。作为一种补救措施,本文通过增加机上存储器交易来研究可用内存带宽的最佳利用。所需的CUDA和OPENCL代码不是为不同的GPU加速器手动执行此操作,而是从特定于域的语言(DSL)中为所考虑的应用程序域中自动生成的。此外,DSL的扩展还可以支持全球还原运营商。我们表明,生成的特定目标代码可显着改善用于内存的内核的带宽利用率。此外,与广泛使用的图像处理库OpenCV的GPU后端相比,竞争性能可以实现。
An efficient memory bandwidth utilization for GPU accelerators is crucial for memory bound applications. In medical imaging, the performance of many kernels is limited by the available memory bandwidth since only a few operations are performed per pixel. For such kernels only a fraction of the compute power provided by GPU accelerators can be exploited and performance is predetermined by memory bandwidth. As a remedy, this paper investigates the optimal utilization of available memory bandwidth by means of increasing in-flight memory transactions. Instead of doing this manually for different GPU accelerators, the required CUDA and OpenCL code is automatically generated from descriptions in a Domain-Specific Language (DSL) for the considered application domain. Moreover, the DSL is extended to also support global reduction operators. We show that the generated target-specific code improves bandwidth utilization for memory-bound kernels significantly. Moreover, competitive performance compared to the GPU back end of the widely used image processing library OpenCV can be achieved.