A Framework for Lattice QCD Calculations on GPUs

A Framework for Lattice QCD Calculations on GPUs
复制标题

GPU 上的格子 QCD 计算框架

DOI:
10.1109/ipdps.2014.112
复制
发表时间:
2014
期刊:
2014 IEEE 28th International Parallel and Distributed Processing Symposium
影响因子:
--
通讯作者:
B. Joó
B. Joó
中科院分区:
--
文献类型:
--
作者:
F. Winter;M. Clark;R. Edwards;B. Joó

文献摘要

被引文献

相似文献

配备有加速器(如GPU)的计算平台已被证明可提供强大的计算能力。然而,利用这些平台用于现有的科学应用并不是一项微不足道的任务。目前的GPU编程框架,如CUDA C/C++,需要开发人员进行低级编程,以实现高性能代码。因此,将应用程序移植到GPU上通常仅限于时间主导的算法和例程,而其余的则没有加速,这可能会引发严重的Amdahl定律问题。Lattice QCD应用程序Chroma允许我们探索不同的移植策略。软件体系结构的分层结构在逻辑上将数据并行与应用层分离。QCD数据并行软件层提供了数据类型和表达式,具有适合格点场论的模板式操作。Chroma根据这个高级接口实现算法。因此,通过移植底层,可以一次性有效地移植整个应用层。QDP-JIT/PTX库,我们重新实现的低级别层,提供了一个框架格QCD计算的CUDA架构。支持完整的软件接口,因此应用程序可以在基于GPU的并行计算机上运行。由于JIT编译器的可用性,这种重新实现是可能的,该编译器将汇编语言(PTX)转换为GPU代码。现有的表达式模板使我们能够使用编译时计算来构建代码生成器并自动化CUDA的内存管理。我们的实现使我们能够在大规模基于GPU的机器(如Titan和Blue沃茨)上部署完整的色度计生成程序,并将计算速度加快一个数量级以上。
Computing platforms equipped with accelerators like GPUs have proven to provide great computational power. However, exploiting such platforms for existing scientific applications is not a trivial task. Current GPU programming frameworks such as CUDA C/C++ require low-level programming from the developer in order to achieve high performance code. As a result porting of applications to GPUs is typically limited to time-dominant algorithms and routines, leaving the remainder not accelerated which can open a serious Amdahl's law issue. The Lattice QCD application Chroma allows us to explore a different porting strategy. The layered structure of the software architecture logically separates the data-parallel from the application layer. The QCD Data-Parallel software layer provides data types and expressions with stencil-like operations suitable for lattice field theory. Chroma implements algorithms in terms of this high-level interface. Thus by porting the low-level layer one effectively ports the whole application layer in one swing. The QDP-JIT/PTX library, our reimplementation of the low-level layer, provides a framework for Lattice QCD calculations for the CUDA architecture. The complete software interface is supported and thus applications can be run unaltered on GPU-based parallel computers. This reimplementation was possible due to the availability of a JIT compiler which translates an assembly language (PTX) to GPU code. The existing expression templates enabled us to employ compile-time computations in order to build code generators and to automate the memory management for CUDA. Our implementation has allowed us to deploy the full Chroma gauge-generation program on large scale GPU-based machines such as Titan and Blue Waters and accelerate the calculation by more than an order of magnitude.