Automatic optimization of thread-coarsening for graphics processors

Automatic optimization of thread-coarsening for graphics processors
复制标题

DOI:
10.1145/2628071.2628087
复制
发表时间:
2014-08
期刊:
2014 23rd International Conference on Parallel Architecture and Compilation (PACT)
影响因子:
--
通讯作者:
A. Magni;Christophe Dubach;M. O’Boyle
A. Magni;Christophe Dubach;M. O’Boyle
中科院分区:
其他
文献类型:
--
作者:
A. Magni;Christophe Dubach;M. O’Boyle

文献摘要

被引文献

相似文献

OpenCL 旨在实现不同供应商的多核设备之间的功能可移植性。然而,缺乏单一的跨目标优化编译器严重限制了OpenCL程序的性能可移植性。程序员需要为每个特定设备手动调整应用程序,从而阻碍了有效的可移植性。我们针对特定于数据并行语言的编译器转换:线程粗化,并证明它可以提高不同 GPU 设备的性能。然后,我们解决为粗化因子参数选择最佳值的问题,即决定将多少个线程合并在一起。我们通过实验表明,这是一个很难解决的问题:很难找到好的配置,而简单的粗化实际上会导致速度大幅下降。我们提出了一种基于机器学习模型的解决方案,该模型使用核函数静态特征来预测最佳粗化因子。该模型自动专门针对所考虑的不同架构。我们在四种设备上的 17 个基准测试中评估了我们的方法:两种 Nvidia GPU 和两种不同代的 AMD GPU。使用我们的技术,我们平均实现了 1.11 倍到 1.33 倍的加速。
OpenCL has been designed to achieve functional portability across multi-core devices from different vendors. However, the lack of a single cross-target optimizing compiler severely limits performance portability of OpenCL programs. Programmers need to manually tune applications for each specific device, preventing effective portability. We target a compiler transformation specific for data-parallel languages: thread-coarsening and show it can improve performance across different GPU devices. We then address the problem of selecting the best value for the coarsening factor parameter, i.e., deciding how many threads to merge together. We experimentally show that this is a hard problem to solve: good configurations are difficult to find and naive coarsening in fact leads to substantial slowdowns. We propose a solution based on a machine-learning model that predicts the best coarsening factor using kernel-function static features. The model automatically specializes to the different architectures considered. We evaluate our approach on 17 benchmarks on four devices: two Nvidia GPUs and two different generations of AMD GPUs. Using our technique, we achieve speedups between 1.11× and 1.33× on average.