Evolving MPI+X Toward Exascale

Evolving MPI+X Toward Exascale
复制标题

MPI X 向百亿亿次演进

DOI:
10.1109/mc.2016.232
复制
发表时间:
2016
期刊:
影响因子:
2.2
通讯作者:
David A. Bader
David A. Bader
中科院分区:
计算机科学4区
文献类型:
--
作者:
David A. Bader

文献摘要

被引文献

相似文献

最近高性能计算(HPC)采用诸如GPU、可编程门阵列和协处理器等加速器的趋势导致了计算和存储子系统的显著异构性。应用程序开发者通常跨集群的计算节点使用分层消息传递接口(MPI)编程模型,以及用于每个计算节点内的CPU和加速器设备的内部节点模型(例如OpenMP)或加速器专用库(例如计算统一设备体系结构(CUDA)或开放计算语言(OpenCL))。要达到可接受的性能水平,应用程序程序员必须深入了解机器拓扑、计算能力、内存层次结构、计算内存同步语义和其他系统特征。然而,对计算和内存资源的显式管理以及脱节的编程模型意味着程序员必须在性能和生产率之间做出折衷的S。在“MPI-ACC:用于科学应用的加速器感知MPI”(IEEE译文)中。并行和分布式系统,第27卷,第5期,2016年,第1401-1414页),来自弗吉尼亚理工大学、阿贡国家实验室、北卡罗来纳州立大学和莱斯大学的Ashwin Aji和他的同事们提出了一种用于具有异类计算设备的HPC集群的统一编程模型和运行时系统。具体地说,他们引入了MPI-ACC,这是MPI+X编程模型的一个演进步骤,MPI+X编程模型是分布式内存集群的事实标准。通过发展一种已经很流行的编程模型,作者使现有的基于MPI的应用程序的代码现代化变得更容易。Aji和他的团队注意到,当调用MPI-ACC中的数据移动例程时,程序员可以简单地描述节点内元素的额外数据属性规范-例如GPU命令队列、执行流或设备上下文-而不需要更改MPI标准。MPI-ACC的运行时系统使用特定于用户的数据属性,不仅在网络上执行端到端的数据移动,而且还与实时GPU内核同步,以实现通信与计算的明显重叠。作者将他们的简单描述方法与现有支持GPU的MPI实现的复杂规范方法进行了对比。他们认为,尽管其他方法在GPU之间提供了端到端的数据移动支持,但它们没有一种机制来表达数据的执行属性,这给最终用户带来了重叠通信和计算的负担。研究人员还深入分析了MPI-ACC如何用于规模化生产中的科学应用,如流行病传播模拟和地震学模拟。它们进一步表明,MPI-ACC的流水线端到端数据移动、可扩展的中间资源管理技术和增强的执行进度引擎的性能优于单独使用MPI和CUDA的基线实现。
The recent trend in highperformance computing (HPC) to adopt accelerators such as GPUs, eld-programmable gate arrays, and coprocessors has led to signi cant heterogeneity in computation and memory subsystems. Application developers typically employ a hierarchical message passing interface (MPI) programming model across the cluster’s compute nodes, and an intranode model such as OpenMP or an accelerator-speci c library such as compute uni ed device architecture (CUDA) or open computing language (OpenCL) for the CPUs and accelerator devices within each compute node. To achieve acceptable performance levels, application programmers must have in-depth knowledge of machine topology, compute capability, memory hierarchy, compute-memory synchronization semantics, and other system characteristics. However, explicit management of computation and memory resources along with a disjointed programming model mean that programmers must make tradeo s between performance and productivity. In “MPI-ACC: Accelerator-Aware MPI for Scienti c Applications” (IEEE Trans. Parallel and Distributed Systems, vol. 27, no. 5, 2016, pp. 1401–1414), Ashwin Aji and his colleagues from Virginia Tech, Argonne National Laboratory, North Carolina State University, and Rice University present a uni ed programming model and runtime system for HPC clusters with heterogeneous computing devices. Speci cally, they introduce MPI-ACC, an evolutionary step in the MPI+X programming model, which is the de facto standard for distributed memory clusters. By evolving an already popular programming model, the authors make it easier to modernize the code of existing MPI-based applications. Aji and his team note that when invoking a data-movement routine in MPI-ACC, programmers can simply describe additional data attributes speci c to the within-node elements— such as the GPU command queue, execution stream, or device context— without changing the MPI standard. MPI-ACC’s runtime system employs user-speci ed data attributes to not only perform end-to-end data movement across the network but also synchronize with inight GPU kernels to achieve e cient overlap of communication with computation. The authors contrast their simple descriptive approach with the complex prescriptive approach of existing GPU-aware MPI implementations. They argue that although other approaches provide end-to-end data movement support between GPUs, they don’t have a mechanism to express the data’s execution attributes, which puts the burden of overlapping communication with computation on end users. The investigators also performed in-depth analysis of how MPI-ACC can be used to scale in-production scienti c applications such as an epidemic spread simulation and a seismology simulation. They further show that the MPI-ACC’s pipelined end-to-end data movement, scalable intermediate resource-management techniques, and enhanced execution progress engine outperform baseline implementations that use MPI and CUDA separately.