Strong scaling of general-purpose molecular dynamics simulations on GPUs

Strong scaling of general-purpose molecular dynamics simulations on GPUs
复制标题

DOI:
10.1016/j.cpc.2015.02.028
复制
发表时间:
2015-07-01
影响因子:
6.3
通讯作者:
Glotzer, Sharon C.
Glotzer, Sharon C.
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
Glaser, Jens;Trung Dac Nguyen;Glotzer, Sharon C.

文献摘要

被引文献

相似文献

我们在一个支持GPU的通用分子动力学代码HOOMD-BLUE(Anderson和Glotzer,2013)中描述了MPI区域分解的高度优化实现。我们的方法的灵感来自传统的基于CPU的代码LAMMPS(Plimpton,1995),但从一开始就在为在GPU上执行而设计的代码中实现(Anderson等人,2008)。该软件支持短距离成对力场和键合力场,并使用自动调整算法实现最佳的GPU性能。我们能够在Lennard-Jones和耗散粒子动力学(DPD)模拟中展示高达3375个GPU的同等或更高的伸缩性,最高可达1.08亿个粒子。最近几代GPU的GPUDirect RDMA功能在全双精度计算中提供了更好的性能。对于典型的聚合物物理应用,HOOMD-BLUE 1.0提供了有效的GPU与CPU节点加速12.5倍。(C)2015爱思唯尔B.V.保留所有权利。
We describe a highly optimized implementation of MPI domain decomposition in a GPU-enabled, general-purpose molecular dynamics code, HOOMD-blue (Anderson and Glotzer, 2013). Our approach is inspired by a traditional CPU-based code, LAMMPS (Plimpton, 1995), but is implemented within a code that was designed for execution on GPUs from the start (Anderson et al., 2008). The software supports short-ranged pair force and bond force fields and achieves optimal GPU performance using an autotuning algorithm. We are able to demonstrate equivalent or superior scaling on up to 3375 GPUs in Lennard-Jones and dissipative particle dynamics (DPD) simulations of up to 108 million particles. GPUDirect RDMA capabilities in recent GPU generations provide better performance in full double precision calculations. For a representative polymer physics application, HOOMD-blue 1.0 provides an effective GPU vs. CPU node speed-up of 12.5x. (C) 2015 Elsevier B.V. All rights reserved.