Runtime dependence computation and execution of loops on heterogeneous systems

Runtime dependence computation and execution of loops on heterogeneous systems
复制标题

异构系统上的运行时依赖计算和循环执行

DOI:
--
复制
发表时间:
2013
期刊:
IEEE/ACM International Symposium on Code Generation and Optimization
影响因子:
--
通讯作者:
R. Govindarajan
R. Govindarajan
中科院分区:
--
文献类型:
--
作者:
Jayvant Anantpur;R. Govindarajan

文献摘要

被引文献

相似文献

GPU已用于DOALL循环的并行执行。然而,带有间接数组引用的循环可能会潜在地导致交叉迭代依赖,而使用现有编译技术很难检测到这种依赖。具有此类循环的应用程序不能轻松使用GPU,因此无法从GPU的巨大计算能力中获益。在本文中,我们提出了一种在运行时计算循环中交叉迭代依赖关系的算法。该算法同时使用CPU和GPU来计算依赖关系。具体地说,它有效地利用了GPU的计算能力,通过执行为间接数组访问生成的切片函数来快速收集迭代执行的内存访问。使用依赖信息,循环迭代被均衡化,使得每个级别包含可并行执行的独立迭代。提出的解决方案的另一个有趣的方面是,它将未来级别的依赖计算与当前级别的实际计算流水线连接起来,以有效地利用GPU中的可用资源。我们使用NVIDIA Tesla C2070来评估我们的实施,使用来自Polybench Suite的基准和一些合成基准。我们的实验表明,在合理的交叉迭代依赖次数下,该技术在循环上的平均加速比可以达到6.4倍。
GPUs have been used for parallel execution of DOALL loops. However, loops with indirect array references can potentially cause cross iteration dependences which are hard to detect using existing compilation techniques. Applications with such loops cannot easily use the GPU and hence do not benefit from the tremendous compute capabilities of GPUs. In this paper, we present an algorithm to compute at runtime the cross iteration dependences in such loops. The algorithm uses both the CPU and the GPU to compute the dependences. Specifically, it effectively uses the compute capabilities of the GPU to quickly collect the memory accesses performed by the iterations by executing the slice functions generated for the indirect array accesses. Using the dependence information, the loop iterations are levelized such that each level contains independent iterations which can be executed in parallel. Another interesting aspect of the proposed solution is that it pipelines the dependence computation of the future level with the actual computation of the current level to effectively utilize the resources available in the GPU. We use NVIDIA Tesla C2070 to evaluate our implementation using benchmarks from Polybench suite and some synthetic benchmarks. Our experiments show that the proposed technique can achieve an average speedup of 6.4x on loops with a reasonable number of cross iteration dependences.