Auto-vectorizing a large-scale production unstructured-mesh CFD application

Auto-vectorizing a large-scale production unstructured-mesh CFD application
复制标题

自动矢量化大规模生产非结构化网格 CFD 应用程序

DOI:
10.1145/2870650.2870651
复制
发表时间:
2016
期刊:
--
影响因子:
--
通讯作者:
Mudalige G
Mudalige G
中科院分区:
--
文献类型:
--
作者:
Mudalige G

文献摘要

参考文献

被引文献

相似文献

对于矢量长度越来越长的现代x86 CPU,实现良好的矢量化对于获得更高的性能变得非常重要。使用非常显式的SIMD向量编程技术已被证明可以提供接近最佳的性能,但是它们很难实现所有类别的应用程序,特别是那些具有非常不规则的内存访问的应用程序,并且通常需要对代码进行大量的重构。向量内部函数也不适用于诸如Fortran之类的语言,Fortran仍然在大型生产应用程序中大量使用。另一种方法是依赖于编译器自动向量化,这通常在向量化具有不规则内存访问模式的代码时效果较差。在本文中,我们提出了最近的研究探索技术,以获得编译器自动矢量化的非结构化网格应用程序。一个关键的贡献是软件技术的细节,实现自动矢量化的大型生产级非结构化网格应用程序从计算流体力学领域,以便受益于最新的英特尔处理器上的矢量单元,而无需大量的代码重写。我们使用OP2领域特定库中的代码生成工具,将自动向量化优化自动应用于生产代码库,并与其他并行化(如最新的NVIDIA GPU)的性能相比,进一步探索应用程序的性能。我们看到自动向量化有相当大的性能改进。大型CFD应用程序中计算最密集的并行循环在20核Intel Haswell系统上的加速比非矢量化版本快近40%。然而,并非所有循环都因向量化而增益,其中具有较小计算强度的循环因相关开销而损失性能。
For modern x86 based CPUs with increasingly longer vector lengths, achieving good vectorization has become very important for gaining higher performance. Using very explicit SIMD vector programming techniques has been shown to give near optimal performance, however they are difficult to implement for all classes of applications particularly ones with very irregular memory accesses and usually require considerable re-factorisation of the code. Vector intrinsics are also not available for languages such as Fortran which is still heavily used in large production applications. The alternative is to depend on compiler auto-vectorization which usually have been less effective in vectorizing codes with irregular memory access patterns. In this paper we present recent research exploring techniques to gain compiler auto-vectorization for unstructured mesh applications. A key contribution is details on software techniques that achieve auto-vectorisation for a large production grade unstructured mesh application from the CFD domain so as to benefit from the vector units on the latest Intel processors without a significant code re-write. We use code generation tools in the OP2 domain specific library to apply the auto-vectorising optimisations automatically to the production code base and further explore the performance of the application compared to the performance with other parallelisations such as on the latest NVIDIA GPUs. We see that there is considerable performance improvements with autovectorization. The most compute intensive parallel loops in the large CFD application shows speedups of nearly 40% on a 20 core Intel Haswell system compared to their non-vectorized versions. However not all loops gain due to vectorization where loops with less computational intensity lose performance due to the associated overheads.
使用 OP2 加速全面的工业 CFD 应用
DOI: 10.1109/tpds.2015.2453972
发表时间: 2014
影响因子: 5.3
作者:
I. Reguly;G. Mudalige;C. Bertolli;M. Giles;A. Betts;P. Kelly;David Radford
通讯作者: David Radford
异构并行系统上的高级非结构化网格框架的设计和初始性能
DOI: 10.1016/j.parco.2013.09.004
发表时间: 2013
期刊: Parallel Computing
影响因子: 1.4
作者:
Mudalige G
通讯作者: Mudalige G
DOI: 10.1093/comjnl/bxr062
发表时间: 2011
期刊: The Computer Journal
影响因子: --
作者:
Giles M
通讯作者: Giles M