High Performance Computing Systems. Performance Modeling, Benchmarking, and Simulation - 5th International Workshop, PMBS 2014, New Orleans, LA, USA, November 16, 2014. Revised Selected Papers

High Performance Computing Systems. Performance Modeling, Benchmarking, and Simulation - 5th International Workshop, PMBS 2014, New Orleans, LA, USA, November 16, 2014. Revised Selected Papers
复制标题

高性能计算系统。

DOI:
10.1007/978-3-319-17248-4_5
复制
发表时间:
2015
期刊:
--
影响因子:
--
通讯作者:
Mudalige G
Mudalige G
中科院分区:
--
文献类型:
--
作者:
Mudalige G

文献摘要

相似文献

在这篇文章中,我们提出了一种应用领域特定的高级抽象(Hla)开发策略的研究,其目的是在AWE plc模拟流体力学计算的一类关键的高性能计算(Hpc)应用程序中实现面向未来的应用。我们建立在现有的高级抽象框架OPS的基础上,该框架正在为牛津大学基于网格的多块结构化应用程序的解决方案而开发。OPS使用“活动库”方法,其中使用OPS API编写的单个应用程序代码可以转换为不同的高度优化的并行实现,然后这些实现可以链接到适当的并行库,从而能够在不同的后端硬件平台上执行。这项工作中的目标应用程序是来自Sandia国家实验室的Mantevo代码套件的三叶草迷你应用程序,该程序包含来自流体力学工作负载的感兴趣的算法。具体地说,我们介绍了(1)重新设计一个具有工业代表性的流体力学应用程序以利用OPS高级框架和后续代码生成以获得一系列并行实现的经验教训,以及(2)自动生成的三叶草的OPS版本与手动编码的原始三叶草实现在一系列平台上的性能比较。基准系统包括英特尔多核CPU和NVIDIA GPU、Archer(Cray XC30)CPU集群和具有不同并行化(OpenMP、Openacc、CUDA、OpenCL和MPI)的Titan(Cray XK7)GPU集群。我们的结果表明,使用高级框架(如OPS)开发并行HPC应用程序并不比仅针对单个并行实现编写一次性并行程序更耗时,也不困难。然而,OPS策略的回报是高度可维护的单一应用程序源,通过它可以实现多个并行化,而不会影响一系列并行系统上的性能可移植性。
In this paper we present research on applying a domain specific high-level abstractions (HLA) development strategy with the aim to “future-proof” a key class of high performance computing (HPC) applications that simulate hydrodynamics computations at AWE plc. We build on an existing high-level abstraction framework, OPS, that is being developed for the solution of multi-block structured mesh-based applications at the University of Oxford. OPS uses an “active library” approach where a single application code written using the OPS API can be transformed into different highly optimized parallel implementations which can then be linked against the appropriate parallel library enabling execution on different back-end hardware platforms. The target application in this work is the CloverLeaf mini-app from Sandia National Laboratory’s Mantevo suite of codes that consists of algorithms of interest from hydrodynamics workloads. Specifically, we present (1) the lessons learnt in re-engineering an industrial representative hydro-dynamics application to utilize the OPS high-level framework and subsequent code generation to obtain a range of parallel implementations, and (2) the performance of the auto-generated OPS versions of CloverLeaf compared to that of the performance of the hand-coded original CloverLeaf implementations on a range of platforms. Benchmarked systems include Intel multi-core CPUs and NVIDIA GPUs, the Archer (Cray XC30) CPU cluster and the Titan (Cray XK7) GPU cluster with different parallelizations (OpenMP, OpenACC, CUDA, OpenCL and MPI). Our results show that the development of parallel HPC applications using a high-level framework such as OPS is no more time consuming nor difficult than writing a one-off parallel program targeting only a single parallel implementation. However the OPS strategy pays off with a highly maintainable single application source, through which multiple parallelizations can be realized, without compromising performance portability on a range of parallel systems.