Porting ONETEP to graphical processing unit-based coprocessors. 1. FFT box operations.

Porting ONETEP to graphical processing unit-based coprocessors. 1. FFT box operations.
复制标题

将 ONETEP 移植到基于图形处理单元的协处理器。

DOI:
10.1002/jcc.23410
复制
发表时间:
2013
影响因子:
3
通讯作者:
Wilkinson K
Wilkinson K
中科院分区:
化学3区
文献类型:
--
作者:
Wilkinson K

文献摘要

相似文献

我们提出了第一个支持图形处理单元 (GPU) 协处理器的 N 阶电子总能量包 (ONETEP) 代码版本,用于材料的线性缩放第一原理量子力学计算。这项工作的重点是将涉及原子局域快速傅立叶变换 (FFT) 运算的代码部分移植到 GPU。这些是代码中计算量最大的部分,用于核心算法,例如电荷密度、局部势积分、动能积分和非正交广义 Wannier 函数梯度的计算。我们发现直接移植孤立的 FFT 运算并没有带来任何好处。相反,有必要针对上述每种算法定制端口,以优化进出 GPU 的数据传输。详细讨论了所使用的方法以及对所得性能的测试,结果表明相关算法中的各个步骤都得到了显着的加速。然而,GPU 和主机之间的数据传输是所报告的代码版本中的一个重大瓶颈。此外,还对 ONETEP 能量计算的动态精度方案进行了初步研究,以利用 GPU 增强的单精度功能。这里使用的方法不会破坏现有的代码库。此外,由于此处报告的进展涉及核心算法,因此它们将使 ONETEP 的全部功能受益。我们使用基于指令的编程模型确保了对其他形式协处理器的可移植性,并使这项工作成为未来开发旨在支持新兴高性能计算平台的代码的基础。版权所有 © 2013 Wiley periodicals, Inc.
We present the first graphical processing unit (GPU) coprocessor‐enabled version of the Order‐N Electronic Total Energy Package (ONETEP) code for linear‐scaling first principles quantum mechanical calculations on materials. This work focuses on porting to the GPU the parts of the code that involve atom‐localized fast Fourier transform (FFT) operations. These are among the most computationally intensive parts of the code and are used in core algorithms such as the calculation of the charge density, the local potential integrals, the kinetic energy integrals, and the nonorthogonal generalized Wannier function gradient. We have found that direct porting of the isolated FFT operations did not provide any benefit. Instead, it was necessary to tailor the port to each of the aforementioned algorithms to optimize data transfer to and from the GPU. A detailed discussion of the methods used and tests of the resulting performance are presented, which show that individual steps in the relevant algorithms are accelerated by a significant amount. However, the transfer of data between the GPU and host machine is a significant bottleneck in the reported version of the code. In addition, an initial investigation into a dynamic precision scheme for the ONETEP energy calculation has been performed to take advantage of the enhanced single precision capabilities of GPUs. The methods used here result in no disruption to the existing code base. Furthermore, as the developments reported here concern the core algorithms, they will benefit the full range of ONETEP functionality. Our use of a directive‐based programming model ensures portability to other forms of coprocessors and will allow this work to form the basis of future developments to the code designed to support emerging high‐performance computing platforms.Copyright © 2013 Wiley Periodicals, Inc.