Mesh independent loop fusion for unstructured mesh applications

Mesh independent loop fusion for unstructured mesh applications
复制标题

适用于非结构化网格应用的网格独立循环融合

DOI:
10.1145/2212908.2212917
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
Bertolli C
Bertolli C
中科院分区:
--
文献类型:
--
作者:
Bertolli C

文献摘要

参考文献

被引文献

相似文献

基于非结构化网格的应用程序通常是计算密集型的,导致运行时间较长。原则上,最先进的硬件,如多核CPU和众核GPU,可以用于加速,但这些深奥的架构需要专业知识才能实现最佳性能。OP 2是一个并行编程层,它试图通过允许程序员通过API调用(所谓的OP 2循环)在非结构化网格中的元素上表达并行迭代来减轻这种编程负担。OP 2编译器的基础设施,然后使用源到源的转换,以实现每个OP 2循环的并行实现,并发现optimization.In本文中的机会,我们描述了几个编译器技术可以有效地利用串联,以提高非结构化网格应用程序的性能。特别是,我们展示了如何整个程序分析--这往往是由于控制流图的大小被抑制-往往成为可行的OP 2编程模型的结果,促进积极的优化。随后,我们展示了如何整个程序分析,然后成为一个使OP 2循环优化。在此基础上,我们展示了如何一个经典的技术,即循环融合,这是通常难以适用于非结构化网格应用程序,可以在编译时定义。我们研究其应用的局限性,并显示实验结果的计算流体动力学应用基准,评估由于循环融合的性能增益。
Applications based on unstructured meshes are typically compute intensive, leading to long running times. In principle, state-of-the-art hardware, such as multi-core CPUs and many-core GPUs, could be used for their acceleration but these esoteric architectures require specialised knowledge to achieve optimal performance. OP2 is a parallel programming layer which attempts to ease this programming burden by allowing programmers to express parallel iterations over elements in the unstructured mesh through an API call, a so-called OP2-loop. The OP2 compiler infrastructure then uses source-to-source transformations to realise a parallel implementation of each OP2-loop and discover opportunities for optimisation.In this paper, we describe how several compiler techniques can be effectively utilised in tandem to increase the performance of unstructured mesh applications. In particular, we show how whole-program analysis --- which is often inhibited due to the size of the control flow graph - often becomes feasible as a result of the OP2 programming model, facilitating aggressive optimisation. We subsequently show how whole-program analysis then becomes an enabler to OP2-loop optimisations. Based on this, we show how a classical technique, namely loop fusion, which is typically difficult to apply to unstructured mesh applications, can be defined at compile-time. We examine the limits of its application and show experimental results on a computational fluid dynamic application benchmark, assessing the performance gains due to loop fusion.
适用于非结构化网格应用的 OP2 库的设计和性能
DOI: 10.1007/978-3-642-29737-3_22
发表时间: 2011
期刊: bioRxiv
影响因子: --
作者:
C. Bertolli;A. Betts;G. Mudalige;M. Giles;P. Kelly
通讯作者: P. Kelly
非结构化网格求解器的并行框架
DOI: 10.1007/978-3-0348-8534-8_10
发表时间: 1994
期刊: 2012 IEEE 18th International Conference on Parallel and Distributed Systems
影响因子: --
作者:
D. A. Burgess;P. Crumpton;M. Giles
通讯作者: M. Giles
DOI: 10.1093/comjnl/bxr062
发表时间: 2011
期刊: The Computer Journal
影响因子: --
作者:
Giles M
通讯作者: Giles M
从 WCET 的程序跟踪中保证循环界限识别
DOI: --
发表时间: 2009
期刊: 2009 15th IEEE Real-Time and Embedded Technology and Applications Symposium
影响因子: --
作者:
M. Bartlett;I. Bate;D. Kazakov
通讯作者: D. Kazakov
DOI: --
发表时间: 1996
期刊: ACM-SIGPLAN Symposium on Programming Language Design and Implementation
影响因子: --
作者:
G. Bilardi;K. Pingali
通讯作者: K. Pingali