Optimising purely functional GPU programs
Optimising purely functional GPU programs
复制标题
优化纯函数式 GPU 程序
DOI:
10.1145/2500365.2500595
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
T. L. McDonell
中科院分区:
文献类型:
--
作者:
T. L. McDonell
Purely functional, embedded array programs are a good match for SIMD hardware, such as GPUs. However, the naive compilation of such programs quickly leads to both code explosion and an excessive use of intermediate data structures. The resulting slow-down is not acceptable on target hardware that is usually chosen to achieve high performance. In this paper, we discuss two optimisation techniques, sharing recovery and array fusion, that tackle code explosion and eliminate superfluous intermediate structures. Both techniques are well known from other contexts, but they present unique challenges for an embedded language compiled for execution on a GPU. We present novel methods for implementing sharing recovery and array fusion, and demonstrate their effectiveness on a set of benchmarks.