Twinned buffering: A simple and highly effective scheme for parallelization of Successive Over-Relaxation on GPUs and other accelerators

Twinned buffering: A simple and highly effective scheme for parallelization of Successive Over-Relaxation on GPUs and other accelerators
复制标题

DOI:
10.1109/hpcsim.2015.7237073
复制
发表时间:
2015-07
期刊:
2015 International Conference on High Performance Computing & Simulation (HPCS)
影响因子:
--
通讯作者:
W. Vanderbauwhede;T. Takemi
W. Vanderbauwhede;T. Takemi
中科院分区:
其他
文献类型:
--
作者:
W. Vanderbauwhede;T. Takemi

文献摘要

相似文献

在本文中,我们提出了一个新的计划,并行化的逐次超松弛方法求解泊松方程的三维体积。我们的新方案既简单又有效,在NVIDIA GeForce GTX 590 GPU上的性能比传统的红黑方案高出16倍,在NVIDIA GeForce TITAN Black GPU上高出11倍,在Intel Xeon Phi上高出5倍。与在英特尔至强CPU上运行的完全优化的参考实现相比,在GTX 590上的速度提升了16倍,在TITAN上提升了22倍,在Xeon Phi上提升了5倍。我们解释的理由和OpenCL的实施,并提出了性能评估结果。
In this paper we present a new scheme for parallelization of the Successive Over-Relaxation method for solving the Poisson equation over a 3-D volume. Our new scheme is both simple and effective, outperforming the conventional Red-Black scheme by a factor of 16 on an NVIDIA GeForce GTX 590 GPU, a factor of 11 on an NVIDIA GeForce TITAN Black GPU and a factor of 5 on an Intel Xeon Phi. The speed-up compared to the fully optimised reference implementation running on an Intel Xeon CPU is 16 times on the GTX 590, 22 times on the TITAN and 5 times on the Xeon Phi. We explain the rationale and the implementation in OpenCL and present the performance evaluation results.