Practical applicability of optimizations and performance models to complex stencil-based loop kernels in CFD
Practical applicability of optimizations and performance models to complex stencil-based loop kernels in CFD
复制标题
优化和性能模型对 CFD 中基于模板的复杂循环内核的实际适用性
DOI:
10.1177/1094342018774126
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
W. Wall
中科院分区:
文献类型:
--
作者:
K. Wichmann;M. Kronbichler;R. Löhner;W. Wall
This work investigates the application and interaction of optimization techniques and performance models in a computational fluid dynamics (CFD) approach employing an OpenMP parallelized, explicit, weakly compressible, finite difference–based solver for the incompressible Navier–Stokes equations using a five-point wide stencil. The presented loop and stencil optimizations lead to a 6.8× increase in per core throughput. In order to verify optimal CPU utilization, performance models are applied to the tuned code. Three different performance models are considered: a roofline-based model, utilizing purely theoretical figures, one which is enhanced by measurements, and the execution cache memory model. It is shown that the models provide reliable estimates for simple benchmarks, such as seven-point stencils for scalar Laplacians, but the estimate quality is significantly worse for the complex and tuned stencil. While it is possible to include even more details in the model, it eventually leads to a state in which it purely reproduces the benchmarks from which it was derived. Thus, the applied general-purpose performance models are found to inaccurately predict the actual performance. They overestimate the achievable performance by more than about 97% for highly tuned code. Through further code tuning, 66% of the predicted performance could be achieved.