Some useful optimisations for unstructured computational fluid dynamics codes on multicore and manycore architectures

Some useful optimisations for unstructured computational fluid dynamics codes on multicore and manycore architectures
复制标题

对多核和众核架构上的非结构化计算流体动力学代码进行一些有用的优化

DOI:
10.1016/j.cpc.2018.07.001
复制
发表时间:
2019
期刊:
Comput. Phys. Commun.
影响因子:
--
通讯作者:
L. Mare
L. Mare
中科院分区:
--
文献类型:
--
作者:
I. Hadade;Feng Wang;M. Carnevale;L. Mare

文献摘要

被引文献

相似文献

本文提出了一些优化,以提高非结构化计算流体动力学代码在多核和多核架构上的性能,如英特尔Sandy Bridge, Broadwell和Skylake cpu以及英特尔Xeon Phi Knights Corner和Knights Landing多核处理器。我们讨论并演示了它们在两种不同类型的计算核中的实现:由通量计算表示的基于人脸的循环和表示状态向量更新的基于细胞的循环。我们提出了在这两类计算核中有效利用底层向量单位的重要性,特别强调了向量化基于人脸的循环及其固有的间接和不规则访问模式所需的变化。我们展示了以细胞为中心的不同数据布局的优势,以及面部数据结构和架构特定优化,以提高非结构化网格应用中普遍存在的收集和分散操作的性能。基于自动调优的软件预取策略的实现也显示了多线程对顺序架构(如Knights Corner)的重要性的经验评估。我们探索了英特尔至强Phi骑士登陆架构上可用的各种存储模式,并提出了一种方法,可以利用传统DRAM和MCDRAM接口来实现最大性能。我们在双插槽节点配置的多核cpu上获得了2.8到3倍的显著应用程序加速,在Intel Xeon Phi Knights Corner协处理器上获得了8.6倍的速度,在Intel Xeon Phi Knights Landing处理器上获得了5.6倍的速度,在一个非结构化的有限体积CFD代码中,其大小和复杂性对于工业应用具有代表性。程序摘要程序标题:some_opt_for_unstructured_cfdProgram文件doi:http://dx.doi.org/10.17632/zyh2zkf3jw.1Licensing条款:GNU通用公共许可证3 (GPL)编程语言:C/ c++问题的性质:流体流动问题的解决方案在复杂的几何图形附近强制使用非结构化网格。然而,这种非结构化网格方法在处理复杂几何形状时的灵活性是以从现代处理器中提取高性能的难度增加为代价的。我们提供了许多优化的实现,这些优化有助于提高现代多核和多核架构上非结构化CFD代码的性能。解决方法:通过Reverse Cuthill-Mckee进行网格重新编号,实现向量化所需的代码转换,在积累残差时消除面端点依赖的面部着色/重新排序,减少缓存缺失的数据布局转换,手动调整寄存器内转置的聚集和分散原语,通过自动调整和多线程利用现代处理器的SMT功能的软件预取。
This paper presents a number of optimisations for improving the performance of unstructured computational fluid dynamics codes on multicore and manycore architectures such as the Intel Sandy Bridge, Broadwell and Skylake CPUs and the Intel Xeon Phi Knights Corner and Knights Landing manycore processors. We discuss and demonstrate their implementation in two distinct classes of computational kernels: face-based loops represented by the computation of fluxes and cell-based loops representing updates to state vectors. We present the importance of making efficient use of the underlying vector units in both classes of computational kernels with special emphasis on the changes required for vectorising face-based loops and their intrinsic indirect and irregular access patterns. We demonstrate the advantage of different data layouts for cell-centred as well as face data structures and architectural specific optimisations for improving the performance of gather and scatter operations which are prevalent in unstructured mesh applications. The implementation of a software prefetching strategy based on auto-tuning is also shown along with an empirical evaluation on the importance of multithreading for in-order architectures such as Knights Corner. We explore the various memory modes available on the Intel Xeon Phi Knights Landing architecture and present an approach whereby both traditional DRAM as well as MCDRAM interfaces are exploited for maximum performance. We obtain significant full application speed-ups between 2.8 and 3X across the multicore CPUs in two-socket node configurations, 8.6X on the Intel Xeon Phi Knights Corner coprocessor and 5.6X on the Intel Xeon Phi Knights Landing processor in an unstructured finite volume CFD code representative in size and complexity to an industrial application.Program summaryProgram Title:some_opt_for_unstructured_cfdProgram Files doi:http://dx.doi.org/10.17632/zyh2zkf3jw.1Licensing provisions:GNU General Public License 3 (GPL)Programming language:C/C++Nature of problem:The solution of fluid flow problems in the vicinity of complex geometries mandates the utilisation of unstructured grids. However, this flexibility of unstructured mesh methods in dealing with complicated geometries comes at a cost of increased difficulty in extracting high performance out of modern processors. We provide implementations for a number of optimisations useful for improving the performance of unstructured CFD codes on modern multicore and manycore architectures.Solution method:grid renumbering via Reverse Cuthill–Mckee, code transformations necessary for enabling vectorisation, face colouring/reordering for removing dependencies at the face end-points when accumulating residuals, data layout transformations for reducing cache misses, hand-tuned gather and scatter primitives for in-register transpositions, software prefetching via auto-tuning and multithreading for exploiting SMT features of modern processors.