Automatic Performance Tuning of Stencil Computations on GPUs
Automatic Performance Tuning of Stencil Computations on GPUs
复制标题
GPU 上模板计算的自动性能调整
DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
T. Abdelrahman
中科院分区:
文献类型:
--
作者:
Joseph Garvey;T. Abdelrahman
We consider automatic performance tuning of stencil computations on Graphics Processing Units. We present a strategy that uses machine learning to determine the best way to use memory followed by a heuristic that divides the remaining optimizations into groups and exhaustively explores one group at a time. We evaluate our strategy using 102 synthetically generated OpenCL stencil kernels on an Nvidia GTX Titan GPU. We assess our strategy both in terms of the number of configurations explored during auto-tuning and the quality of the best configuration obtained. We explore two alternative heuristics that use different groupings of the optimizations. We show that, relative to a random sampling of the space and an expert search, our strategy achieves a reduction in the number of configurations explored of up to 80% and 84% respectively while also finding better performing configurations.