Multi-GPU-based Swendsen-Wang multi-cluster algorithm with reduced data traffic

Multi-GPU-based Swendsen-Wang multi-cluster algorithm with reduced data traffic
复制标题

基于多 GPU 的 Swendsen-Wang 多集群算法,可减少数据流量

DOI:
10.1016/j.cpc.2015.04.025
复制
发表时间:
2015
影响因子:
6.3
通讯作者:
Yukihiro Komura
Yukihiro Komura
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
Hirasawa T;Kuratani S;Hirasawa T;平沢達矢;Y. Komura and Y. Okabe;Yukihiro Komura

文献摘要

相似文献

多GPU应用的计算性能会因为GPU之间的数据通信而降低。为了实现多gpu的高速计算,我们应该最小化这种数据通信的成本。在本文中,我为Swendsen-Wang (SW)多集群算法提出了一种多GPU计算方法,减少了每个GPU之间的数据流量。我通过提前调整每个GPU之间的连接信息来实现数据流量的减少。在大型开放科学超级计算机TSUBAME 2.5上实现了该代码,并通过临界温度下三维Ising模型的模拟对其性能进行了评估。结果表明,每个GPU之间的数据通信减少了90%,每个GPU之间的通信次数减少了约一半。使用512个gpu,在总系统大小为N= 4096 3的临界温度下,每次自旋更新的计算时间为0.005 ns。
The computational performance of multi-GPU applications can be degraded by the data communication between each GPU. To realize high-speed computation with multiple GPUs, we should minimize the cost of this data communication. In this paper, I propose a multiple GPU computing method for the Swendsen–Wang (SW) multi-cluster algorithm that reduces the data traffic between each GPU. I realize this reduction in data traffic by adjusting the connection information between each GPU in advance. The code is implemented on the large-scale open science TSUBAME 2.5 supercomputer, and its performance is evaluated using a simulation of the three-dimensional Ising model at the critical temperature. The results show that the data communication between each GPU is reduced by 90%, and the number of communications between each GPU decreases by about half. Using 512 GPUs, the computation time is 0.005 ns per spin update at the critical temperature for a total system size of N= 4096 3.