Large-scale GW calculations on pre-exascale HPC systems

Large-scale GW calculations on pre-exascale HPC systems
复制标题

DOI:
10.1016/j.cpc.2018.09.003
复制
发表时间:
2019-02-01
影响因子:
6.3
通讯作者:
Deslippe, Jack
Deslippe, Jack
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
Del Ben, Mauro;da Jornada, Felipe H.;Deslippe, Jack

文献摘要

被引文献

相似文献

从头算引力波方法是一种严格的基于格林函数的框架,可用于计算各种材料的电子激发特性,例如扩展系统、分子以及受限和纳米结构材料,并且具有非常令人满意的精度。然而,复杂系统上的引力波计算常常受到与该方法相关的高计算成本的阻碍。在这里,我们演示了如何使用基于非阻塞分块循环通信方案的计算密集型内核的新颖算法显着加速 GW 计算,该方案最大限度地减少 MPI 消息中的延迟,允许通信和计算重叠,并提高矩阵乘法运算中的缓存使用率。在 BerkeleyGW 软件包中实现的代码优化版本能够很好地扩展到 NERSC (Cray XC40) 的完整 Cori 计算机,并实现超过 11 Peta FLOP/s 的持续性能。我们通过对硅上的缺陷结构进行大规模引力场计算来展示我们的工作,这需要包含超过 1700 个原子的模拟单元,现在可以在大型预兆级高性能计算系统上在短短几分钟内高效执行。 (C) 2018 Elsevier B.V. 保留所有权利。
The ab initio GW approach is a rigorous Green's-function-based framework that can be employed to compute electronic excitation properties of a wide variety of materials such as extended systems, molecules, as well as confined and nanostructured materials with a very satisfactory accuracy. However, GW calculations on complex systems are often hindered by the high computational cost associated with the method. Here, we demonstrate how to significantly speedup GW calculations with a novel algorithm for the computationally intense kernel based on a non-blocking chunked cyclic communication scheme that minimizes latency in MPI messages, allows for overlapping of communication and computation, and improves cache usage in matrix multiplication operations. The optimized version of the code, implemented in the BerkeleyGW software package, is capable of scaling well to the full Cori computer at NERSC (Cray XC40) and achieves over 11 Peta FLOP/s of sustained performance. We showcase our work by performing large-scale GW calculations of defect structures on silicon, which require simulation cells containing over 1700 atoms, and which can now be efficiently executed in just a few minutes on large pre-exascale high-performance computing systems. (C) 2018 Elsevier B.V. All rights reserved.