Implementing molecular dynamics on hybrid high performance computers - Particle-particle particle-mesh

Implementing molecular dynamics on hybrid high performance computers - Particle-particle particle-mesh
复制标题

DOI:
10.1016/j.cpc.2011.10.012
复制
发表时间:
2012-03-01
影响因子:
6.3
通讯作者:
Tharrington, Arnold N.
Tharrington, Arnold N.
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
Brown, W. Michael;Kohlmeyer, Axel;Tharrington, Arnold N.

文献摘要

被引文献

相似文献

图形处理单元(gpu)等加速器的使用由于其低成本、令人印象深刻的浮点能力、高内存带宽和低电力需求而在科学计算应用程序中变得流行。由于这些优势,混合高性能计算机,即节点包含多种类型浮点处理器(例如CPU和GPU)的机器,现在变得越来越普遍。在本文中,我们提出了在分布式内存并行混合机器的LAMMPS分子动力学软件中使用加速器的算法的延续。在我们之前的工作中,我们专注于短程模型的加速,其方法旨在利用加速器和(多核)cpu的处理能力。为了加强现有的实现,我们提出了一个有效的远程静电力计算分子动力学。具体来说,我们提出了一种基于Harvey和De Fabritiis工作的粒子-粒子-粒子网格方法的实现。我们给出了Keeneland InfiniBand GPU集群上的基准测试结果。我们提供了用CUDA和OpenCL编译的相同内核的性能比较。我们讨论了并行效率的限制以及在混合或异构计算机上提高性能的未来方向。(C) 2011 Elsevier B.V.版权所有
The use of accelerators such as graphics processing units (GPUs) has become popular in scientific computing applications due to their low cost, impressive floating-point capabilities, high memory bandwidth, and low electrical power requirements. Hybrid high-performance computers, machines with nodes containing more than one type of floating-point processor (e.g. CPU and GPU), are now becoming more prevalent due to these advantages. In this paper, we present a continuation of previous work implementing algorithms for using accelerators into the LAMMPS molecular dynamics software for distributed memory parallel hybrid machines. In our previous work, we focused on acceleration for short-range models with an approach intended to harness the processing power of both the accelerator and (multi-core) CPUs. To augment the existing implementations, we present an efficient implementation of long-range electrostatic force calculation for molecular dynamics. Specifically, we present an implementation of the particle-particle particle-mesh method based on the work by Harvey and De Fabritiis. We present benchmark results on the Keeneland InfiniBand GPU cluster. We provide a performance comparison of the same kernels compiled with both CUDA and OpenCL. We discuss limitations to parallel efficiency and future directions for improving performance on hybrid or heterogeneous computers. (C) 2011 Elsevier B.V. All rights reserved.