Practical applicability of optimizations and performance models to complex stencil-based loop kernels in CFD

Practical applicability of optimizations and performance models to complex stencil-based loop kernels in CFD
复制标题

优化和性能模型对 CFD 中基于模板的复杂循环内核的实际适用性

DOI:
10.1177/1094342018774126
复制
发表时间:
2019
期刊:
The International Journal of High Performance Computing Applications
影响因子:
--
通讯作者:
W. Wall
W. Wall
中科院分区:
--
文献类型:
--
作者:
K. Wichmann;M. Kronbichler;R. Löhner;W. Wall

文献摘要

被引文献

相似文献

这项工作研究了计算流体动力学 (CFD) 方法中优化技术和性能模型的应用和交互,该方法采用 OpenMP 并行、显式、弱可压缩、基于有限差分的求解器,使用五点宽模板求解不可压缩的纳维-斯托克斯方程。所提出的循环和模板优化使每核吞吐量增加了 6.8 倍。为了验证最佳的 CPU 利用率,将性能模型应用于调整后的代码。考虑了三种不同的性能模型:基于屋顶线的模型,利用纯理论数据,通过测量增强的模型,以及执行高速缓存模型。结果表明,这些模型为简单的基准提供了可靠的估计,例如标量拉普拉斯的七点模板,但对于复杂和调整的模板,估计质量明显较差。虽然可以在模型中包含更多细节,但它最终会导致一种状态,在这种状态下,它纯粹再现了派生它的基准。因此,发现所应用的通用性能模型不能准确地预测实际性能。对于高度调优的代码,他们高估了可实现的性能约 97% 以上。通过进一步的代码调优,可以实现 66% 的预测性能。
This work investigates the application and interaction of optimization techniques and performance models in a computational fluid dynamics (CFD) approach employing an OpenMP parallelized, explicit, weakly compressible, finite difference–based solver for the incompressible Navier–Stokes equations using a five-point wide stencil. The presented loop and stencil optimizations lead to a 6.8× increase in per core throughput. In order to verify optimal CPU utilization, performance models are applied to the tuned code. Three different performance models are considered: a roofline-based model, utilizing purely theoretical figures, one which is enhanced by measurements, and the execution cache memory model. It is shown that the models provide reliable estimates for simple benchmarks, such as seven-point stencils for scalar Laplacians, but the estimate quality is significantly worse for the complex and tuned stencil. While it is possible to include even more details in the model, it eventually leads to a state in which it purely reproduces the benchmarks from which it was derived. Thus, the applied general-purpose performance models are found to inaccurately predict the actual performance. They overestimate the achievable performance by more than about 97% for highly tuned code. Through further code tuning, 66% of the predicted performance could be achieved.