OP2: An active library framework for solving unstructured mesh-based applications on multi-core and many-core architectures

OP2: An active library framework for solving unstructured mesh-based applications on multi-core and many-core architectures
复制标题

OP2:一个主动库框架,用于解决多核和众核架构上的非结构化基于网格的应用程序

DOI:
10.1109/inpar.2012.6339594
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
Mudalige G
Mudalige G
中科院分区:
--
文献类型:
--
作者:
Mudalige G

文献摘要

参考文献

被引文献

相似文献

OP 2是一个“主动”库框架,用于解决非结构化网格应用程序。它利用源到源的转换和编译,因此使用OP 2 API编写的单个应用程序代码可以转换为不同的并行实现,以便在不同的后端硬件平台上执行。在本文中,我们提出了目前的OP 2库的设计,并探讨其在实现性能的可移植性,接近最佳的性能,并在现代多核和众核处理器为基础的系统扩展的能力。这项工作的一个关键特点是OP 2的最新扩展,促进分布式内存集群的GPU上的应用程序的开发和执行。我们讨论了在异构平台上并行化非结构化网格应用程序的主要设计问题。这些包括在访问间接引用数据时处理数据依赖关系,非结构化网格数据布局的影响(结构数组与数组结构)以及生成在GPU集群上执行的代码时的设计考虑因素。使用OP 2框架编写的代表性CFD应用程序用于提供一系列多核/众核系统的对比基准测试和性能分析研究。其中包括英特尔(韦斯特米尔和桑迪桥)和AMD(Magny-Cours)的多核CPU,NVIDIA(GTX 560 Ti,Tesla C2070)的GPU,分布式内存CPU集群(Cray XE 6)和分布式内存GPU集群(带InfiniBand的Tesla C2050 GPU)。OP 2的设计选择进行了探讨,定量的见解,他们的贡献性能。我们证明,一个应用程序编写一次在高层次上使用OP 2 API可以很容易地跨各种对比平台的便携式,是能够实现接近最佳性能的域应用程序员的干预。
OP2 is an “active” library framework for the solution of unstructured mesh-based applications. It utilizes source-to-source translation and compilation so that a single application code written using the OP2 API can be transformed into different parallel implementations for execution on different back-end hardware platforms. In this paper we present the design of the current OP2 library, and investigate its capabilities in achieving performance portability, near-optimal performance, and scaling on modern multi-core and many-core processor based systems. A key feature of this work is OP2's recent extension facilitating the development and execution of applications on a distributed memory cluster of GPUs. We discuss the main design issues in parallelizing unstructured mesh based applications on heterogeneous platforms. These include handling data dependencies in accessing indirectly referenced data, the impact of unstructured mesh data layouts (array of structs vs. struct of arrays) and design considerations in generating code for execution on a cluster of GPUs. A representative CFD application written using the OP2 framework is utilized to provide a contrasting benchmarking and performance analysis study on a range of multi-core/many-core systems. These include multi-core CPUs from Intel (Westmere and Sandy Bridge) and AMD (Magny-Cours), GPUs from NVIDIA (GTX560Ti, Tesla C2070), a distributed memory CPU cluster (Cray XE6) and a distributed memory GPU cluster (Tesla C2050 GPUs with InfiniBand). OP2's design choices are explored with quantitative insights into their contributions to performance. We demonstrate that an application written once at a high-level using the OP2 API can be easily portable across a wide range of contrasting platforms and is capable of achieving near-optimal performance without the intervention of the domain application programmer.
适用于非结构化网格应用的 OP2 库的设计和性能
DOI: 10.1007/978-3-642-29737-3_22
发表时间: 2011
期刊: bioRxiv
影响因子: --
作者:
C. Bertolli;A. Betts;G. Mudalige;M. Giles;P. Kelly
通讯作者: P. Kelly
DOI: 10.1016/j.procs.2010.04.203
发表时间: 2010-05
期刊: --
影响因子: --
作者:
G. Markall;D. Ham;P. Kelly
通讯作者: G. Markall;D. Ham;P. Kelly
不连续伽辽金方法的自动代码生成
DOI: 10.1137/070710032
发表时间: 2008
期刊: SIAM J. Sci. Comput.
影响因子: --
作者:
Kristian B. Ølgaard;A. Logg;G. N. Wells
通讯作者: G. N. Wells
DOI: 10.1016/j.finel.2003.10.006
发表时间: 2004-07
影响因子: 3.1
作者:
James R Stewart;H.Carter Edwards
通讯作者: James R Stewart;H.Carter Edwards
混合并行 CFD 求解器的新型共享内存线程池实现
DOI: 10.1007/978-3-642-23397-5_18
发表时间: 2011
期刊: SIGMETRICS Perform. Evaluation Rev.
影响因子: --
作者:
J. Jägersküpper;C. Simmendinger
通讯作者: C. Simmendinger