Hybrid MPI and CUDA paralleled finite volume unstructured CFD simulations on a multi-GPU system

Hybrid MPI and CUDA paralleled finite volume unstructured CFD simulations on a multi-GPU system
复制标题

DOI:
10.1016/j.future.2022.09.005
复制
发表时间:
2022-09
期刊:
Future Gener. Comput. Syst.
影响因子:
--
通讯作者:
Xi Zhang;Xiaohu Guo;Yue Weng;Xianwei Zhang;Yutong Lu;Zhong Zhao
Xi Zhang;Xiaohu Guo;Yue Weng;Xianwei Zhang;Yutong Lu;Zhong Zhao
中科院分区:
其他
文献类型:
--
作者:
Xi Zhang;Xiaohu Guo;Yue Weng;Xianwei Zhang;Yutong Lu;Zhong Zhao

文献摘要

被引文献

相似文献

将可压缩流的非结构计算流体动力学(CFD)分析移植到图形处理单元(GPU)面临两个困难。首先,对GPU的全局存储器的非合并访问由间接数据访问引起,从而导致性能损失。其次,由于进程之间的数据通信和主机与设备之间的数据传输,多GPU之间的数据交换复杂,这降低了可扩展性。为了提高非结构化有限体积GPU模拟可压缩流的数据局部性,我们进行了一些优化,包括单元和面重新编号,数据依赖解决,嵌套循环分裂,循环模式调整。在此基础上,提出了一种基于GPU的MPI-CUDA混合并行计算框架,该框架支持数据的打包和解包交换。最后,经过优化,整个应用程序在GPU上的性能提高了50%左右。在单个GPU(Nvidia Tesla V100)上模拟ONERA M6案例,与28个CPU核心(Intel Xeon Gold 6132)相比,平均可以实现13.4的加速比。在2个GPU的基线上,强扩展测试结果显示200个GPU上的并行效率为42%,而弱扩展测试显示200个GPU上的并行效率为82.4%。
Porting unstructured Computational Fluid Dynamics (CFD) analysis of compressible flow to Graphics Processing Units (GPUs) confronts two difficulties. Firstly, non-coalescing access to the GPU’s global memory is induced by indirect data access leading to performance loss. Secondly, data exchange among multi-GPU is complex due to data communication between processes and transfer between host and device, which degrades scalability. For increasing data locality on unstructured finite volume GPU simulations for compressible flow, we perform some optimizations, including cell and face renumbering, data dependence resolving, nested loops split, and loop mode adjustment. Then, a hybrid MPI-CUDA parallel framework with packing and unpacking exchange data on GPU is established for multi-GPU computing. Finally, after optimizations, the performance of the whole application on a GPU is increased by around 50%. Simulations of ONERA M6 cases on a single GPU (Nvidia Tesla V100) can achieve an average of 13.4 speedup compared to those on 28 CPU cores (Intel Xeon Gold 6132). On the baseline of 2 GPUs, strong scaling results show a parallel efficiency of 42% on 200 GPUs, while weak scaling tests give a parallel efficiency of 82.4% up to 200 GPUs.