Implementation of Multi-GPU Based Lattice Boltzmann Method for Flow Through Porous Media

Implementation of Multi-GPU Based Lattice Boltzmann Method for Flow Through Porous Media
复制标题

DOI:
10.4208/aamm.2014.m468
复制
发表时间:
2015-02
影响因子:
1.4
通讯作者:
Changsheng Huang;B. Shi;N. He;Z. Chai
Changsheng Huang;B. Shi;N. He;Z. Chai
中科院分区:
工程技术3区
文献类型:
--
作者:
Changsheng Huang;B. Shi;N. He;Z. Chai

文献摘要

被引文献

相似文献

格子Boltzmann方法(LBM)可以利用图形处理器(GPU)的计算优势获得大量的性能优势,因此,GPU或基于多GPU的LBM可以被认为是研究大规模流体流动的有前途和有竞争力的候选者。然而,基于多GPU的格子Boltzmann算法还没有得到广泛的研究,特别是在复杂几何形状的流动模拟。本文结合消息传递接口(MPI)技术,提出了一种基于多GPU的多孔介质渗流LBM的实现方法,并给出了基于数据结构和布局的优化策略,该方法可以显著减少内存访问,完全隐藏通信时间消耗。然后在一个配备4块Tesla C1060 GPU的单节点集群上测试了算法的性能,其中Poietille流的计算速度达到了1732 MFLUPS,并且随着GPU数量的增加,算法的加速比接近线性。
The lattice Boltzmann method (LBM) can gain a great amount of performance benefit by taking advantage of graphics processing unit (GPU) computing, and thus, the GPU, or multi-GPU based LBM can be considered as a promising and competent candidate in the study of large-scale fluid flows. However, the multi-GPU based lattice Boltzmann algorithm has not been studied extensively, especially for simulations of flow in complex geometries. In this paper, through coupling with the message passing interface (MPI) technique, we present an implementation of multi-GPU based LBM for fluid flow through porous media as well as some optimization strategies based on the data structure and layout, which can apparently reduce memory access and completely hide the communication time consumption. Then the performance of the algorithm is tested on a one-node cluster equipped with four Tesla C1060 GPU cards where up to 1732 MFLUPS is achieved for the Poiseuille flow and a nearly linear speedup with the number of GPUs is also observed.