Scalable multi-GPU 3-D FFT for TSUBAME 2.0 Supercomputer

Scalable multi-GPU 3-D FFT for TSUBAME 2.0 Supercomputer
复制标题

DOI:
10.1109/sc.2012.100
复制
发表时间:
2012-11
期刊:
2012 International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Akira Nukada;Kento Sato;S. Matsuoka
Akira Nukada;Kento Sato;S. Matsuoka
中科院分区:
其他
文献类型:
--
作者:
Akira Nukada;Kento Sato;S. Matsuoka

文献摘要

被引文献

相似文献

对于使用多个GPU的可扩展3-D FFT计算,GPU之间有效的全能通信是良好性能的最重要因素。具有点对点MPI库功能和CUDA内存复制API的实现通常显示出非常大的间接开销,尤其是在许多节点之间的全部通信中的小消息大小。我们提出了几个方案,以最大程度地减少间接费用,包括使用Infiniband的低级API有效地重叠和节点间交流和自动调整策略来控制调度调度并确定铁路分配。结果,我们可以使用256个Tsubame 2.0超级计算机(768 GPU)的256个节点以双重精度实现高达4.8 tflops的良好性能和良好性能。
For scalable 3-D FFT computation using multiple GPUs, efficient all-to-all communication between GPUs is the most important factor in good performance. Implementations with point-to-point MPI library functions and CUDA memory copy APIs typically exhibit very large overheads especially for small message sizes in all-to-all communications between many nodes. We propose several schemes to minimize the overheads, including employment of lower-level API of InfiniBand to effectively overlap intra- and inter-node communication, as well as auto-tuning strategies to control scheduling and determine rail assignments. As a result we achieve very good strong scalability as well as good performance, up to 4.8TFLOPS using 256 nodes of TSUBAME 2.0 Supercomputer (768 GPUs) in double precision.