Parallel Implementation of Density Functional Theory Methods in the Quantum Interaction Computational Kernel Program

Parallel Implementation of Density Functional Theory Methods in the Quantum Interaction Computational Kernel Program
复制标题

DOI:
10.1021/acs.jctc.0c00290
复制
发表时间:
2020-07-14
影响因子:
5.5
通讯作者:
Merz, Kenneth M., Jr.
Merz, Kenneth M., Jr.
中科院分区:
化学1区
文献类型:
--
作者:
Manathunga, Madushanka;Miao, Yipu;Merz, Kenneth M., Jr.

文献摘要

被引文献

相似文献

我们给出了一个支持图形处理单元(GPU)的交换相关(XC)方案的细节,该方案集成到开源量子相互作用计算内核(QUICK)程序中。我们的实现具有基于八叉树的数值网格点划分方案,支持GPU的网格修剪和基元函数预筛选,以及完全支持GPU的XC能量和梯度算法。与CPU版本的基准测试表明,GPU实现能够提供令人印象深刻的性能,同时保持出色的准确性。对于中小型蛋白质/有机分子体系,在NVIDIA V100图形处理器上实现的双精度XC能量和梯度计算的加速比在串口CPU实现上分别提高了60-80倍和140-500倍。在密度泛函理论计算中,单个V100 GPU获得的加速显著超过了40核并行运行的现代CPU。
We present the details of a graphics processing unit (GPU) capable exchange correlation (XC) scheme integrated into the open source QUantum Interaction Computational Kernel (QUICK) program. Our implementation features an octree based numerical grid point partitioning scheme, GPU enabled grid pruning and basis and primitive function prescreening, and fully GPU capable XC energy and gradient algorithms. Benchmarking against the CPU version demonstrated that the GPU implementation is capable of delivering an impressive performance while retaining excellent accuracy. For small to medium size protein/organic molecular systems, the realized speedups in double precision XC energy and gradient computation on a NVIDIA V100 GPU were 60-80-fold and 140-500-fold, respectively, as compared to the serial CPU implementation. The acceleration gained in density functional theory calculations from a single V100 GPU significantly exceeds that of a modern CPU with 40 cores running in parallel.