Lattice Boltzmann for Large-Scale GPU Systems

Lattice Boltzmann for Large-Scale GPU Systems
复制标题

用于大规模 GPU 系统的格子玻尔兹曼

DOI:
--
复制
发表时间:
2011
期刊:
International Conference on Parallel Computing
影响因子:
--
通讯作者:
K. Stratford
K. Stratford
中科院分区:
--
文献类型:
--
作者:
A. Gray;Alistair Hart;A. Richardson;K. Stratford

文献摘要

被引文献

相似文献

我们描述了启用路德维希格子玻尔兹曼并行流体动力学应用程序,专为复杂的问题,大规模并行GPU加速架构。NVIDIA CUDA被引入到现有的C/MPI框架中,除了性能之外,我们还仔细考虑了可维护性。通过重组数据布局以允许内存合并和调整关键循环以减少片外内存访问,在每个GPU上实现了显着的性能提升。Halo-swap通信阶段旨在有效地并行利用多个GPU:包括使用CUDA流功能的多个阶段的重叠。新的GPU适配被认为保留了原始CPU代码的良好扩展行为,并可扩展到256个NVIDIA Fermi GPU(测试的最大资源)。NVIDIA Fermi GPU的性能被观察到比(12核)AMD Magny-Cours CPU(使用所有内核)高出4倍,用于二进制流体基准测试。
We describe the enablement of the Ludwig lattice Boltzmann parallel fluid dynamics application, designed specifically for complex problems, for massively parallel GPU-accelerated architectures. NVIDIA CUDA is introduced into the existing C/MPI framework, and we have given careful consideration to maintainability in addition to performance. Significant performance gains are realised on each GPU through restructuring of the data layout to allow memory coalescing and the adaptation of key loops to reduce off-chip memory accesses. The halo-swap communication phase has been designed to efficiently utilise many GPUs in parallel: included is the overlapping of several stages using CUDA stream functionality. The new GPU adaptation is seen to retain the good scaling behaviour of the original CPU code, and scales well up to 256 NVIDIA Fermi GPUs (the largest resource tested). The performance on the NVIDIA Fermi GPU is observed to be up to a factor of 4 greater than the (12-core) AMD Magny-Cours CPU (with all cores utilised) for a binary fluid benchmark.